Effective today, Connecticut’s amended Data Privacy Act requires companies to tell you — clearly, in their privacy notices — whether they collect, use, or sell your personal data to train large language models. Connecticut is the first US state to mandate this specific disclosure.
Most people won’t hear about this. It passed as part of a broader amendment to the Connecticut Data Privacy Act, received limited press coverage, and takes effect on a date that happens to coincide with several other privacy law changes. But it matters more than its quiet arrival suggests.
For the first time, one of the most contested questions in personal data — “is my information being used to train AI?” — now has a required, legally accountable answer in at least one US jurisdiction. And answers like this tend not to stay in one state.
Why This Fills a Gap
The standard technique for obscuring AI training on user data has been broad licensing language embedded in terms of service. When you agreed to use a cloud storage service, a note-taking app, or a social platform, you likely granted them a licence to use your content for purposes including “improving services,” “developing new technologies,” and “research and development.”
Those phrases are doing substantial work. “Improving services” can mean training a recommendation system on your files. “Developing new technologies” can mean using your documents and photos to build general-purpose language models. “Research and development” can mean sharing your data with AI research teams working on models that have no relationship to the product you signed up for.
None of this was spelled out in the language most users actually read. It lived in subordinate clauses that were technically disclosed but functionally invisible. The ambiguity was, and often still is, a deliberate feature — it maximises the platform’s flexibility to use data for AI training without creating the accountability that explicit disclosure would generate.
Connecticut’s amendment collapses that ambiguity for covered companies. You’re not allowed to shelter LLM training inside generic “service improvement” language anymore. You have to say whether you’re doing it.
What the Law Actually Requires
The CTDPA amendment requires companies to include in their privacy notices a statement that is “reasonably accessible, clear and meaningful” disclosing whether they:
- Collect personal data for the purpose of training large language models
- Use personal data to train large language models
- Sell personal data for LLM training purposes
The disclosure must be kept current as practices change. This is not a one-time compliance checkbox. If a company adds LLM training to its data practices after publication, the privacy notice must be updated to reflect that.
Crucially, the law does not define “large language model.” The legislature deliberately chose the recognisable public term rather than a narrow technical definition, presumably to prevent the requirement from becoming obsolete as AI terminology evolves. Legal guidance from practitioners tracking the law suggests companies should err on the side of breadth — if a system you use or build upon might plausibly qualify as an LLM, treat the disclosure requirement as applying.
The obligation also extends to vendors and processors acting on your behalf. If your company sends user data to an analytics partner, a cloud infrastructure provider, or a data enrichment service that then uses that data to train models, your disclosure must account for that too. The chain of liability runs to the downstream use.
Who Is Covered
The CTDPA threshold was also lowered by the same amendment package: companies previously needed to process data for 100,000 or more Connecticut consumers annually to fall under the law. That number drops to 35,000 starting today. A significantly larger population of businesses now must comply.
The law applies to for-profit entities processing personal data of Connecticut residents that meet the threshold and don’t qualify for specific exemptions. Financial institutions under Gramm-Leach-Bliley, HIPAA-covered healthcare entities, government agencies, and a few other categories receive partial or complete carve-outs.
If you don’t live in Connecticut, this requirement still affects your experience. Privacy law compliance in practice tends toward the strictest applicable standard applied universally, because building parallel compliance systems for different user populations is expensive and error-prone. Many national companies will update their global privacy notice to include the LLM training disclosure for all users — not only Connecticut residents.
How to Find the Disclosure
The practical question is how this shows up in a real privacy notice, and what to look for.
Under the “reasonably accessible, clear and meaningful” standard, companies can’t satisfy this requirement with a buried footnote or a vague reference to “AI initiatives” in a 15,000-word document. But “reasonably accessible” is still a judgment call, and companies will test its limits. Here’s what to actually look for.
Search for “LLM” or “large language model”
The most direct route. After today, privacy notices from covered companies should contain one of these terms — or a reference to “training artificial intelligence models,” “training generative AI,” or equivalent language. If a notice doesn’t mention LLMs at all, the company either isn’t covered, qualifies for an exemption, or hasn’t updated its notice yet. Each of those possibilities is worth knowing.
Look under “how we use your data”
The disclosure is most likely to appear in the section of a privacy notice that lists purposes of processing. Look for language that explicitly references model training or LLM development — and read carefully. “We use your data to improve our services” without further specification does not satisfy this requirement. “We do not use your personal data to train large language models” or “we use [specific categories] of your data for the purpose of training large language models, including for [specified purpose]” does.
Check when the notice was last updated
Privacy notices that haven’t been updated since before today may simply not include this disclosure yet. If the “last updated” date at the top of a notice is months old, the policy may be non-compliant — or the company may have recently moved outside the coverage threshold. Either way, a freshly updated notice is more likely to contain accurate information about current data practices.
Pay attention to negative disclosures
This is the underreported dimension of the law. Connecticut’s requirement creates a strong commercial incentive for companies that don’t train models on user data to say so explicitly in their privacy notices, because that disclosure is now a meaningful differentiator. If a notice says “we do not collect, use, or sell your personal data for the purpose of training large language models,” that statement is on record. Violating it creates material legal exposure.
A company that can honestly make that statement, and does, is signalling something genuine about its practices — not just something about its marketing copy.
What the Law Doesn’t Do
The Connecticut LLM disclosure requirement is a transparency mandate, not a prohibition. Companies are permitted to train models on user data; they just have to say so. This makes the law most useful as an information tool rather than a protection in itself.
If you read a privacy notice, find a clear statement that the company uses your data for LLM training, and decide that’s not acceptable to you, you now have the information to act on that choice. But the law doesn’t attach an opt-out right specifically to LLM training — though Connecticut’s existing rights to opt out of sale of personal data and to opt out of targeted advertising still apply to their respective uses.
The law also doesn’t reach LLM training that happens entirely outside Connecticut’s scope. A company that doesn’t process Connecticut residents’ data, or that falls below the threshold, isn’t covered. Training conducted in other jurisdictions under other rules isn’t affected.
And the requirement only covers personal data, not all data. Aggregate or anonymised datasets — if the anonymisation is genuine and legally compliant — fall outside the definition.
The Broader Pattern
State privacy law in the United States has a consistent direction: provisions introduced in one state tend to appear in other states’ legislation within a few years. The right to opt out of the sale of personal data started in California in 2018 and has since appeared in the privacy statutes of more than a dozen other states. Sensitive data protections for biometric information and neural data have followed similar paths.
The LLM training disclosure requirement is precisely the kind of well-defined, technically coherent, politically accessible provision that travels well. It doesn’t require building new enforcement infrastructure. It doesn’t impose a ban that would generate industry opposition. It requires honesty, and honesty is difficult to argue against publicly.
The practical implication: even if you live outside Connecticut and use services not covered by the CTDPA, the privacy notices you interact with are going to start including LLM training disclosures — because companies with national user bases find it simpler to apply the strictest requirement everywhere than to maintain state-by-state variations.
How daftei Handles This
daftei doesn’t use your personal content — your photos, documents, voice notes, or any other stored files — to train large language models for third parties. There’s no AI training pipeline that extracts value from your archive for anyone’s benefit other than your own.
If Connecticut’s disclosure requirement applied to daftei, the statement would be direct: we do not collect, use, or sell your personal data for the purpose of training large language models. Your content is stored encrypted at rest with AES-256 and in transit with TLS 1.3. daftei is GDPR and CCPA compliant, and the business model — subscription-based, no advertising, no data brokerage — doesn’t create incentives to train models on what you store.
The disclosure Connecticut now requires is, for some companies, a difficult admission. For a company genuinely not engaged in that practice, it’s just the truth.
The Upshot
Starting today, covered companies must tell Connecticut residents — and in practice, often all their users — whether they’re training AI on personal data. Most privacy notices haven’t been updated yet. They will be, and when they are, you’ll have a new signal to look for.
Find the LLM training disclosure in the privacy notice of any service you use to store files, take notes, or keep records. If it’s absent from a recently updated notice, ask why. If it says yes, decide whether that’s acceptable to you. If it says no, note that the company has put itself on record.
This is one new privacy right you can use without filing a formal request. Read the notice.