privacy

Your AI Assistant Can Read Your Files. Know What It Sees.

Copilot, Gemini, and ChatGPT have different levels of access to your personal files. Here's what each actually sees, stores, and potentially uses.

The AI assistant bundled with your cloud storage is no longer optional. It’s the feature that every major platform is building toward: an intelligent interface that can answer questions about your files, summarize documents, find information across your stored content, and work with what you’ve saved. Microsoft has Copilot integrated with OneDrive. Google has Gemini integrated with Drive and Photos. OpenAI’s ChatGPT allows file uploads and maintains memory of past conversations.

These are genuinely useful features. They’re also features that, by design, require the AI to process the content of your personal files. Understanding what each assistant actually has access to, what happens to the content it processes, and what you can control is worth knowing before you rely on these features for anything sensitive.


How AI File Access Actually Works

Before comparing the specific platforms, it’s worth understanding the basic mechanism of AI file access.

Modern AI assistants operate by sending file content to a large language model for processing. When you ask Copilot to summarize a document in your OneDrive, the document’s text is retrieved from your storage, packaged in a prompt, and sent to the model. The model processes the text and returns a response. What happens to that content after the response — whether it’s retained, whether it influences future model training, whether it’s logged — depends on the platform’s specific policies.

There’s an important distinction between two types of AI interaction with your files:

Session-based access. The AI reads your file to answer a question in a specific session. The content is used for the duration of that conversation and then discarded from the AI’s active context. This is the stated model for many AI-assisted file features.

Training data retention. The AI platform retains some or all of the content processed through user interactions to improve future versions of the model. This is governed by each platform’s privacy policy and opt-out mechanisms, and the policies differ significantly between providers.

Understanding which of these is happening — and under what conditions — requires reading the fine print, which is what this post attempts to summarize.


Microsoft Copilot and OneDrive

Microsoft Copilot is deeply integrated with the Microsoft 365 ecosystem, including OneDrive. For personal Microsoft accounts (as opposed to work or school accounts managed by an organization), Copilot can access files stored in OneDrive when you explicitly invoke it to do so.

What Copilot can access: Files in your OneDrive that you explicitly share with Copilot in a conversation. When you use the “Work with your files” feature or upload a file directly to a Copilot conversation, Copilot processes that file’s content.

Training on personal content: Microsoft’s consumer privacy policy states that for personal accounts, data submitted through Copilot may be used to improve Microsoft’s AI services. Microsoft has an opt-out for this: in your Microsoft account privacy settings, under Bing search settings (which governs consumer Copilot), you can disable the setting that allows your conversations to improve AI services.

Work vs. personal accounts: The distinction between personal and work Microsoft accounts matters significantly. Microsoft 365 accounts managed by an organization (the M365 Copilot for business product) have a separate policy: Microsoft commits not to use that content to train general-purpose AI models. Personal consumer accounts have different — and weaker — commitments.

What to watch for: Microsoft’s Recall feature, which captures screenshots of your computer activity at regular intervals and makes them searchable through AI, is available on Copilot+ PCs running Windows 11. If you have a Copilot+ PC, Recall may be capturing screenshots of files you open, documents you read, and content you view — not just files you explicitly upload. Recall’s privacy settings and controls are separate from OneDrive and Copilot integration.


Google Gemini and Drive / Google Photos

Google Gemini is integrated with Google Workspace and with personal Google accounts. Its access to Google Drive and Google Photos is part of Google’s broader effort to make its AI assistant able to answer questions about your stored content.

What Gemini can access: When you use Gemini in the context of Google Drive or Google Workspace, Gemini can access and process files in your Drive when you explicitly direct it to. Gemini can also access Google Photos through Lens and related features, which are governed by the separate Search Services History settings discussed elsewhere.

Training on personal content: Google’s privacy practices for Gemini have evolved. Under Google’s updated terms of service effective July 30, conversations with Gemini may be reviewed by human reviewers and may be used to improve Google products. Google offers a pause-history feature within the Gemini settings that prevents conversations from being saved — but this applies to Gemini conversations, not necessarily to the underlying file access.

The Google ecosystem compounding effect: Gemini’s access is not limited to a single Google product. Because Gemini integrates across Google Drive, Gmail, Google Calendar, Google Docs, and Google Photos, a Gemini assistant that can answer questions about your files has access to a cross-product view of your digital life. A question that references a meeting on your calendar, a document in your Drive, and a photo from your Photos library pulls together data from multiple Google services simultaneously.

Workspace vs. personal accounts: Similar to Microsoft, Google Workspace accounts managed through an organization have different AI training commitments than personal Gmail accounts. Google has committed that data from paid Workspace business accounts is not used to train general-purpose AI models. Personal Google accounts have weaker protections.


ChatGPT and File Uploads

OpenAI’s ChatGPT handles file access differently from Copilot and Gemini, primarily because it doesn’t have persistent integration with a specific cloud storage service. Instead, you explicitly upload files to individual conversations.

What ChatGPT can access: Whatever you explicitly upload to a conversation. ChatGPT can process PDFs, text files, spreadsheets, images, and other file types that you add to a chat. The platform doesn’t proactively access your OneDrive, Google Drive, or any other cloud storage.

Training on personal content: OpenAI’s default policy for personal accounts is that conversations may be used to improve models, including the content of uploaded files. OpenAI offers an opt-out: in your ChatGPT settings, you can disable “Improve the model for everyone,” which prevents your conversations (including file uploads) from being used for training. Business and API accounts have a separate data policy that excludes user content from training by default.

Memory feature: ChatGPT’s Memory feature, when enabled, allows the AI to retain facts and preferences from previous conversations and apply them to future ones. This creates a persistent profile that accumulates across sessions. Memory can be disabled in settings, and you can view and delete what’s been retained. If you’ve uploaded sensitive files in past conversations and had memory enabled, reviewing what’s stored in your ChatGPT memory is worth doing.

The opt-out state matters: For all ChatGPT personal accounts, the default state is opt-in to training. Unless you have specifically navigated to settings and disabled training, your file uploads are currently eligible to be included in model training data.


The Files You Should Keep Out of AI Reach

Regardless of which AI assistant you use and what your settings are, certain categories of files warrant extra care.

Financial documents. Tax returns, bank statements, investment account details, and pay stubs contain sensitive financial information that has direct utility for identity theft and fraud. Even a brief processing of a tax return’s content by an AI system creates an exposure that the convenience of AI summarization probably doesn’t justify.

Legal documents. Contracts, settlement agreements, legal correspondence, and wills contain information that may be subject to attorney-client privilege or other confidentiality obligations. Uploading these to a cloud AI assistant may create complications depending on the sensitivity of the content and your jurisdiction.

Medical and health records. Health information is among the most sensitive personal data. Documents from medical providers, insurance explanation-of-benefits statements, prescription records, and test results deserve deliberate handling. See also the considerations for health journey photos discussed in the photo privacy context.

Personal identification. Passports, driver’s licenses, Social Security cards, and similar documents are high-value targets. Uploading images of these to any AI assistant — even for the purpose of extracting information — creates a record of that document on the platform’s servers.

Personal correspondence. Private messages, emails, and journal entries that you wouldn’t share publicly deserve the same standard in storage as you’d apply offline. An AI assistant that can “read” your personal letters to help you draft responses is doing so by processing the content of those letters on a third-party server.


What Controls You Actually Have

Each platform provides some controls, and the state of those controls matters more than the controls themselves.

Check your training opt-out status: For all three platforms — Microsoft Copilot, Google Gemini, and ChatGPT — there is a setting to opt out of having your content used for training. This setting is not always enabled by default. If you haven’t checked recently, check now.

For Microsoft: My Account → Privacy → Privacy Dashboard → Microsoft AI settings.

For Google: myaccount.google.com → Data & Privacy → My Activity → Gemini Apps Activity. Pause activity to prevent Gemini conversations from being saved.

For ChatGPT: Settings → Data controls → Improve the model for everyone (disable).

Use organization accounts where possible: If you have access to a work or school account with Microsoft 365 Copilot for Business or Google Workspace, the AI training protections for those accounts are substantially stronger than for personal consumer accounts. Where it’s practical to use an organizational account for work that involves sensitive files, that’s worth doing.

Separate sensitive files from AI-accessible storage: The most reliable protection is isolation. Files stored in a service that doesn’t have an AI assistant — or in a service whose AI you haven’t enabled — aren’t accessible to AI processing. If you use Google Drive for general file storage but don’t want certain files processed by Gemini, storing those files in a separate service creates a practical barrier.


What “Processing” Means in Practice

A reasonable concern is whether AI processing of your files means a human has read them. The answer is nuanced.

For AI model inference — the AI generating a response — automated systems process the file content without a human reading it. The file content is converted to a numerical representation, processed through the model, and a response is generated. No human reviewer typically reads the content in this workflow.

For model training review — when conversations are used to improve models — a portion of interactions are reviewed by human annotators to evaluate the quality of AI responses. These reviewers may see the content of files that were uploaded. The probability that any specific file is reviewed is low, but it’s not zero.

Opting out of training doesn’t change how the AI processes your files in real time. It changes whether the content of that processing is retained for training purposes. This distinction matters: opting out of training doesn’t prevent the AI from reading your file to answer a question. It prevents what you uploaded from being retained and potentially used to train future model versions.


The Alternative Position

There’s a version of this that’s not doom-and-gloom. AI assistants that can navigate your personal files are genuinely useful tools. The ability to ask questions about a large set of stored documents, to summarize content, to find information across thousands of files — this has real productivity value.

The question isn’t whether to use these features at all. It’s whether the current privacy conditions for using them are acceptable for the specific files you’re working with — and whether you’ve made a deliberate choice about that, rather than finding out after the fact.

For general productivity tasks with non-sensitive files, AI file assistance is a reasonable tool. For anything in the categories above — financial documents, health records, legal correspondence, personal identification — a more deliberate approach to what you upload and where those files live is worth the friction.

The best outcome is knowing what these systems actually see, deciding which files are appropriate to process through them, and keeping the most sensitive content in storage that doesn’t have an AI reading it on your behalf.

Services like daftei don’t use your stored files to train AI models, don’t integrate AI assistants that read your personal content without your explicit action, and store files with AES-256 encryption at rest under a privacy policy that commits to never selling your data or showing ads. For sensitive personal files — the ones you wouldn’t want an AI system to process — that’s the relevant differentiator.

Your memories deserve better than an ad platform.

Try daftei free →
← All posts