Every time you hit “transcribe” in an AI transcription app, your voice travels to a cloud server you don’t control, gets processed by a model you didn’t train, and enters a data pipeline governed by a privacy policy you almost certainly haven’t read in full.
For a one-off business meeting summary, the stakes feel manageable. But a growing number of people use these apps for therapy session notes, personal journals, medical appointments, and intimate conversations. The gap between what users expect and what actually happens to that audio is wider than most realize.
How AI Transcription Actually Works
On-device and cloud transcription feel identical from the user’s perspective — you speak, text appears. The critical difference is where the processing happens.
On-device transcription runs entirely on your phone’s local chip. Audio never leaves the device. Apple’s Voice Memos transcription works this way: the Speech Recognition framework processes audio locally, nothing is sent to Apple’s servers, and no third party ever touches your recording. Pixel Recorder on Android operates similarly for basic transcription.
Cloud transcription — the default for most productivity-focused apps — sends your audio to remote servers. The voice recording travels over the network, is processed by a model on company infrastructure, and is returned to you as text. The audio file may or may not be stored after transcription; the transcript almost always is.
Apps like Otter.ai and Fireflies.ai are cloud-based. That’s how they achieve accuracy across accents, noisy environments, and multiple simultaneous speakers — it requires compute that isn’t available on a phone chip alone. The price of that accuracy is that your audio leaves your device.
What Otter.ai’s Privacy Policy Actually Says
Otter.ai is one of the most widely used AI transcription services, and its privacy practices are representative of the category.
The company de-identifies user data before using it for model training, stripping individual voices and names before audio enters training pipelines. Otter also states that audio is processed automatically, without humans manually reviewing your recordings — that’s a genuine safeguard.
The more significant issue is retention. Otter’s privacy policy reserves the right to retain recordings and transcripts even after you delete them from your account, if the company determines retention serves “legitimate business purposes.” Legal analysts reviewing the policy have described this as effectively indefinite retention for training-adjacent purposes.
A class action lawsuit — In re Otter.AI Privacy Litigation — is currently working through the courts challenging this retention posture. Until the court rules, Otter operates under legal uncertainty that its competitors don’t face.
Otter also shares data with third-party data labeling providers who help create annotated training datasets. The de-identification travels with that sharing, but the data leaves Otter’s own infrastructure and enters a third party’s systems.
The practical implication: deleting a conversation from your Otter account likely removes your ability to access it. What happens to the de-identified version of that audio within the company’s training pipeline is a separate question the policy doesn’t clearly resolve.
The Voiceprint Problem
A detail that hasn’t received enough attention: some AI transcription systems don’t just convert speech to text. They analyze vocal characteristics to identify who is speaking in a multi-person recording — pitch, cadence, pacing, the acoustic fingerprint that distinguishes your voice from others in a group call. This “speaker diarization” feature is now standard in meeting transcription tools.
When a system captures and stores enough information about your voice to consistently identify you across sessions, what it has captured is legally a biometric identifier in many jurisdictions.
The Illinois Biometric Information Privacy Act defines voice characteristics as protected biometric data, and multiple companies have faced BIPA litigation for capturing voiceprints without explicit written consent. Several lawsuits filed in 2025 and 2026 specifically named AI transcription services under BIPA, arguing that speaker identification models constitute biometric collection.
If a transcription app generates a representation of your voice sufficient to re-identify you across recordings — even without storing your name — it may be handling biometric data subject to state-level requirements in the U.S. and potentially GDPR obligations in the EU, where voice characteristics can fall under the definition of biometric data.
Your Transcripts Can Be Subpoenaed
This is the angle most users don’t consider until it’s too late.
AI-generated transcripts are text files stored on cloud servers, which means they’re subject to legal discovery. If a company receives a valid subpoena or court order, it’s required to produce records it holds — and your transcripts may be in scope.
For most people, most of the time, this risk is theoretical. But for conversations that touch anything legally sensitive — a dispute with an employer, a medical situation, a family legal matter, anything financial — a full verbatim transcript stored on a third-party server is a liability that wouldn’t exist if the same conversation had been written in a private notebook.
Civil discovery in commercial disputes has expanded significantly in 2025 and 2026. Attorneys now routinely flag AI meeting notes and voice transcripts as discoverable. A transcript is a more complete record than memory, more legible than handwritten notes, and easier to subpoena than either — which is precisely why it requires more care about where it lives.
Conference Calls With AI Notetakers: A Problem for Everyone Else
Most discussions about transcription privacy focus on the person who enabled the app. The other participants in the conversation are equally affected — and in many states, equally protected by law.
Recording consent requirements vary significantly by jurisdiction. In all-party consent states — California, Illinois, Washington, Florida, and a dozen others — recording a conversation without the consent of all participants is a violation, not a technicality. An AI notetaker that silently joins a meeting without announcement may trigger these statutes.
Meeting organizers who add a transcription bot without notifying participants aren’t just being inconsiderate. They may be creating legal exposure for themselves or their organization, depending on where participants are located and which laws apply.
The emerging standard in legal guidance: inform participants at the start of any recorded call that the session is being transcribed, identify the tool being used, and give people the option to decline. That’s not a complete legal shield, but it’s the baseline that courts and regulators have begun to expect.
What On-Device Transcription Can and Can’t Do
On-device transcription solves the cloud privacy concern entirely. If audio never leaves your phone, there’s no third-party server to subpoena, no training pipeline to enter, no privacy policy to track.
The practical limitations are real. On-device models lag cloud models in accuracy for challenging audio: strong accents, overlapping speakers, technical terminology, low-quality microphones in noisy environments. They can’t update silently in the background as model capabilities improve. And transcription is computationally demanding, which affects battery life on mobile devices.
Apple’s on-device transcription in Voice Memos handles single-speaker dictation well for most everyday use cases. Multi-speaker meeting transcription with high accuracy remains a domain where cloud tools have a clear advantage.
The trade-off is explicit: better privacy in exchange for some accuracy. How much that trade-off matters depends entirely on what you’re recording. For a meeting to finalize marketing copy, it probably doesn’t matter much. For a conversation about a medical diagnosis or a legal dispute, it matters considerably.
Questions Worth Asking Before Any Transcription App
Where does processing happen? If the app doesn’t clearly state on-device processing, assume cloud.
Does the company train on user data? Most do, even with de-identification. Look for an explicit opt-out in the settings — it’s often buried.
How long is audio retained? A transcript being deleted from your account is not the same as the underlying audio being deleted from the company’s infrastructure.
What’s the third-party sharing picture? Data labeling providers are third parties; de-identified audio may leave the original company’s servers and enter another company’s systems.
Is the service subject to U.S. law enforcement requests? For cloud services operating in the U.S., valid court orders can compel data disclosure regardless of the company’s stated privacy commitments.
The Difference Between Transcription and Storage
If what you actually need is a place to keep voice notes privately — not a tool to transcribe live meetings — a storage-first approach sidesteps most of these concerns.
Recording into an app that preserves the audio file without routing it through a cloud AI means the content stays as you recorded it. It isn’t analyzed for voiceprints, isn’t run through model training pipelines, and isn’t sitting in a transcript database subject to third-party discovery.
daftei stores voice notes and audio files without processing the content for AI training — neither its own models nor any third party’s. Files are encrypted in transit with TLS 1.3 and at rest with AES-256. The app is available on iOS, Android, and the web, with 5 GB free and unlimited storage on Pro. The content of what you record is yours; it doesn’t get analyzed, indexed for inference, or shared with advertising or data partners.
For personal journaling, medical reminders, or anything where the content is sensitive enough that you’d rather not see it in a court filing: the right tool isn’t a transcription app. It’s private storage.