privacy

When Your AI Assistant Can Read All Your Files

AI agents in Google Workspace, Microsoft 365, and Apple Intelligence now access your photos, documents, and email. What data they use, retain, and train on.

The pitch is irresistible: an AI that knows everything you’ve ever saved, everything you’ve written, every photo you’ve taken — and can surface exactly the right thing when you need it. Ask it for “that receipt from the dentist in March,” and it finds it. Ask it to draft a message to your lawyer referencing last year’s emails, and it does it.

This capability is being rolled out across the major platforms right now. Google’s Gemini runs inside Google Workspace. Microsoft’s Copilot runs inside Microsoft 365. Apple Intelligence runs on-device and draws on iCloud content. All of them offer some version of “your AI, with access to all your stuff.”

Most users accept the permission request and move on. Very few read what they’ve agreed to.

What AI Agents Are Actually Accessing

The term “AI agent” refers to an AI system that can take actions on your behalf, not just answer questions. An agent can read files, send emails, search calendars, and in some implementations, create or delete documents.

To do any of this, the agent needs access to your data. The scope of that access varies by platform, but it is typically broad.

Google Gemini in Workspace can access Gmail messages, Google Drive files, Google Calendar events, Google Docs, Sheets, Slides, and Google Photos. The scope of access is comprehensive across all data the Google account has permission to reach. When you ask Gemini a question that touches email or documents, it is actively reading that content to formulate a response.

Microsoft Copilot in Microsoft 365 accesses OneDrive files, SharePoint documents, Outlook email, Teams messages, and calendar data. Copilot can search across all of this and use the content in its responses. Microsoft has implemented Graph API access controls that allow organizations to restrict what Copilot can reach, but individual consumer accounts have more limited control options.

Apple Intelligence takes a hybrid approach. Much of the processing happens on-device using local models, and Apple’s Private Cloud Compute routes more sensitive queries to Apple’s servers under technical attestation that the data isn’t logged. Apple Intelligence can access Photos, Notes, Messages, Mail, and Files on the device and in iCloud.

The Training Data Question

The question most users ask first is: “Is my data being used to train the AI?”

The answer depends heavily on which product you’re using, which tier you’re on, and how carefully you’ve read the terms.

Google Workspace: Google’s terms for Workspace allow the company to use content to operate and improve its services. Google has added specific commitments that Workspace customer data is not used to train generalist AI models without an enterprise agreement opting in. Consumer Google accounts — Gmail, Google Photos, standard Drive — have weaker protections. Google’s privacy policy for consumer products allows broader use of content to “improve services.”

Microsoft 365: Microsoft differentiates between Microsoft 365 commercial customers (where content is not used to train foundation models without consent) and consumer services. The distinction in practice means that a Microsoft 365 Personal or Family subscriber exists in a middle category where terms are less explicit than enterprise agreements.

Apple Intelligence: Apple has stated that queries processed through Private Cloud Compute are not stored or used for training. On-device processing never leaves the device. Apple’s privacy stance here is among the strongest of the three, though it is harder to independently verify than on-device-only claims.

The honest summary is: for enterprise customers on explicit contractual terms, the boundaries are clearer. For individual consumers using free or standard tiers, the protections are often expressed in terms that allow for significant flexibility in how data is “used to improve services.”

What Happens When an Agent Makes a Mistake

Beyond training, there’s a second risk that gets less attention: what happens when the AI agent acts incorrectly on your data.

Agents that can send emails, delete files, or modify documents introduce the possibility of irreversible mistakes at scale. An agent asked to “clean up your inbox” might archive or delete emails you wanted to keep. An agent asked to “organize your drive” might move files to locations that break links or folder structures.

Most implementations include confirmation prompts before destructive actions. But as AI agents become more autonomous — particularly in “autopilot” or “background” modes — the oversight loop shrinks.

Some platforms now offer agents that can run continuously and take actions proactively, without requiring you to initiate each task. The privacy surface of these systems extends beyond what you explicitly ask — it includes what the agent decides to read in the background as part of its ongoing access.

How Third-Party AI Apps Plug Into Your Cloud

The situation becomes more complex when you consider third-party AI applications that request access to your cloud storage.

Many productivity apps, note-taking tools, and AI assistants request OAuth access to Google Drive, iCloud, or OneDrive. When you grant this access, the third-party app gains the ability to read (and sometimes write or delete) files in your cloud storage.

The permissions granted during OAuth authorization are often broader than necessary for the stated use case. An AI writing assistant that says it needs to “access your Drive to import documents” may request permission to read all files, not just the ones you intend to import.

Third-party apps are bound by the platform’s developer policies, which restrict certain uses. But enforcement of these policies is imperfect, and the data has left your storage and entered a third party’s systems.

The chain of custody for your personal files in this ecosystem can be difficult to trace: your files are in Google Drive, which Gemini reads, and which a third-party AI app also has OAuth access to, and that third-party app uses a different AI backend with its own terms of service.

The On-Device Alternative

There is a growing category of AI tools that process your data locally, without sending it to cloud servers.

Obsidian with local AI plugins runs models entirely on-device. Jan.ai is an open-source AI assistant that operates without any cloud connectivity. Tools built on locally-run large language models using frameworks like Ollama keep all data on the machine.

The privacy guarantee of on-device AI is straightforward: data that never leaves the device cannot be accessed by third-party servers, trained on by cloud systems, or exposed in a server-side breach. The tradeoff is capability — local models are generally less capable than large cloud-hosted models, though the gap is narrowing.

For personal data — private documents, sensitive photos, personal correspondence — the on-device tradeoff is often worth making. For professional tasks where raw capability matters more, cloud-hosted agents remain more practical.

What to Actually Check Before Giving Access

Before granting any AI assistant access to your files and cloud storage, these are the specific questions worth answering:

What data does it access? Read the permission request carefully. “Access your files” can mean read-only on files you explicitly share, or it can mean read/write access to everything in your storage.

Where is processing happening? On-device, in the provider’s own cloud, or routed through a third-party AI backend? This determines who has possession of your data and under what terms.

Is your content used for model training? Check the specific terms for your account tier, not the general marketing claims. “We don’t train on your data” sometimes has a carve-out for “aggregate” or “de-identified” data derived from your content.

What is the data retention policy? If the AI processes a document from your drive, how long is that document (or a representation of it) retained in the AI system’s logs or memory?

Can you revoke access? OAuth access to your storage should be revocable from your account settings at any time. Confirm that revocation is complete — that it removes not just future access but also any previously processed data the service has retained.

The Principle at Stake

The core privacy principle that AI agents challenge is purpose limitation — the idea that data collected for one purpose shouldn’t be repurposed for something else without your informed consent.

You created your documents and photos for personal purposes. You stored them in cloud services for convenience and access. Granting an AI agent access to that content for productivity purposes introduces a new principal into the relationship between you and your data.

The AI agent has interests — or more accurately, its developers have interests — that may not align with yours. A calendar integration that reads your appointments to help you schedule also reveals your meeting habits, relationships, and travel patterns to a third party.

Privacy-conscious storage means keeping control over who can access your files and for what purpose. It means choosing services that have clear policies about not training AI models on your personal content, and not granting access to that content to agents whose terms are opaque or permissive.

The convenience of “your AI knows all your stuff” has a real price. Whether that price is worth paying depends on what’s in your files, how much you trust the platforms involved, and how carefully you’ve read the agreement you’re accepting.

Your memories deserve better than an ad platform.

Try daftei free →
← All posts