In the final weeks of its legislative session, New York passed two pieces of AI legislation that drew national attention. One regulated companion chatbots. The other — the AI Training Data Transparency Act — created a new category of legal obligation for AI developers: disclosure of what was in their training data.
The bill, A 6578/S 6955, passed both chambers and now awaits the Governor’s signature. If signed into law, it will require any company that releases or substantially modifies a generative AI model or service to post a public summary of the datasets used to train it — including whether copyrighted material was used, whether personal information was included, and how that data was sourced and processed.
For the first time in US law, the answer to “was my personal data used to train this AI?” would need to be publicly answerable at a meaningful level of specificity.
What the Law Actually Requires
The AI Training Data Transparency Act targets developers of generative AI models — the systems behind text generators, image synthesis tools, voice assistants, and similar products. When a model is released or receives a substantial modification, the developer must publish:
- A high-level summary of the datasets used in training
- Whether those datasets included copyrighted material
- Whether those datasets included personal information
- How the data was collected, processed, and filtered
The disclosure must be posted on the developer’s website and updated when the model changes substantially.
This is more stringent than California’s SB 942 from 2024, which focused primarily on AI-generated content disclosure (the “is this AI-made?” question) rather than training data provenance. New York’s law specifically addresses the input — what went in to create the model — rather than the output.
Why Training Data Transparency Matters
Most major AI models in use today were trained on data scraped from the internet. Photos from public websites. Text from forums, books, and social media posts. Voice recordings from transcription services. Documents uploaded to platforms that had AI training clauses buried in their terms.
Much of this data was personal. A comment someone posted on Reddit in 2015 is personal data — it identifies a real person’s views, often alongside their username, which is often connected to their real identity. A photo published on Instagram with location metadata is personal data. An article someone wrote on Medium is copyrighted. An audio clip uploaded to a transcription service is both.
The people who created this content typically had no idea it would be used for AI training. The companies that scraped it typically did not ask. The result is a situation where billions of people’s personal content has been incorporated into AI systems without their knowledge or consent, and until now there has been no legal mechanism to even confirm that this happened with any particular model.
The New York law does not fix that retroactively. It does not give you a right to demand your data be removed from a model (a technically complex ask, since trained neural networks do not store data the way a database does). But it creates the first meaningful right to know.
What “Personal Information” Means Under This Law
The definition of personal information in the AI Training Data Transparency Act is the same as in New York’s existing privacy statute — which covers the standard categories: names, addresses, biometric data, financial information, and other identifying or linked data.
The critical phrase in the bill is “whether the training data included personal information.” Developers will need to affirmatively investigate and disclose this, not simply assert that their processes “did not intentionally include” personal data.
This matters because AI training pipelines are frequently large-scale and automated. Companies scrape data with crawlers, apply filters, and construct datasets from the results. Many companies building on top of foundation models did not build those models themselves — they fine-tuned existing models with additional data. The law will require disclosure at each stage where personal information may have entered.
The practical enforcement question — how do you verify a dataset claim a developer made on a website? — is a genuine challenge. The law does not create a federal data authority to audit compliance. But it does create a cause of action: if a company claims it did not use personal information in training and that claim turns out to be false, the company faces legal liability.
How This Compares to Other Jurisdictions
New York’s law joins a growing international framework around AI training data, though the global landscape is still fragmented.
European Union: The EU AI Act, which entered enforcement in August 2026 for high-risk systems, imposes obligations on AI providers including documentation of training data and human oversight requirements. The GDPR, which has been in effect since 2018, already subjects AI training data practices involving EU residents to consent and legitimate interest requirements — and regulators have brought enforcement actions against companies that scraped personal data for training without a valid legal basis.
California: SB 942 requires disclosure when AI-generated content is published, but does not directly regulate training data. A separate California Privacy Protection Agency rulemaking on automated decision-making is ongoing and may eventually reach training data questions.
Singapore: The Personal Data Protection Commission issued guidance in 2026 specifically addressing AI training data, signaling that “product improvement” as a catch-all consent basis will no longer be sufficient and that specific, meaningful disclosure is required.
New York’s contribution is significant because it targets training data specifically and at a level of specificity that goes beyond most existing frameworks. It also applies to companies that sell into New York’s market, not just those headquartered there — which gives it national and potentially international reach.
What This Means If Your Photos or Files Were Used in Training
If a generative AI model was trained on images scraped from the web, and some of those images were yours — from your blog, your public social media, a photo you shared on a forum — the new law would require the developer to disclose that personal information was part of the training set.
You would not necessarily find out that your specific photo was included. The disclosure is at the dataset level, not the individual level. But you would know, for the first time, that the model was trained on a dataset that included personal information collected from the web.
From there, depending on where you live, you may have additional rights. Under GDPR, if you are in the EU, you may be able to object to the processing of your personal data for AI training purposes. Under California law, you may be able to submit a deletion request — though the technical feasibility of deletion from a trained model is a matter of ongoing legal and technical debate.
What the law does not do is create a right to opt out prospectively from future training on your data. That right would require separate legislation, and several advocacy organizations are pushing for exactly that.
The Implications for Apps That Store Your Files
The AI Training Data Transparency Act has indirect implications for the apps and services where you store your personal files, photos, and documents.
If a cloud storage service uses the content you upload to fine-tune its AI features — and does not clearly disclose that in its privacy policy — the New York law creates a new standard against which that practice will be judged.
The services with the cleanest position are those that have made an affirmative, product-level commitment not to use user content for AI training at all. Not “we follow applicable law” but “your files are never used to train AI models, period.”
This is the right question to ask of any app that stores your personal files:
- Does your privacy policy permit you to use my uploaded content for AI training?
- If yes, is that an opt-in or an opt-out default?
- Have you ever used uploaded content to train or fine-tune a model?
The answers reveal a lot about how the service thinks about user data — whether your files are an asset it holds for you or a resource it is managing for its own purposes.
The Transparency Gap Still to Be Closed
Even if New York’s law is signed and enforced robustly, it will leave significant gaps.
The disclosure requirement applies at the dataset level, not the individual data point level. Knowing that “this model was trained on a dataset that included personal information from public websites” is meaningfully better than knowing nothing. It is still far short of knowing whether your specific photos, writings, or voice recordings were included.
The law applies to models released or substantially modified after it takes effect. The massive corpus of AI models already trained and deployed — including the foundation models underlying most consumer AI products — are not retroactively subject to it.
And the enforcement mechanism relies on the company’s own disclosure being truthful, backed by litigation risk rather than proactive auditing.
These are not arguments against the law. They are arguments for understanding what it does and does not accomplish, and for the larger work of building privacy infrastructure — from technical standards to regulatory frameworks to better defaults in the apps people use — that treats personal data with the seriousness it deserves.
What to Do Now
The New York AI Training Data Transparency Act has not yet been signed. If and when it is, disclosure obligations will take effect at a date specified in the final law.
In the meantime, the questions the law is designed to answer are worth asking today:
- Review the privacy policies of the AI tools you use most often, looking for training data language
- For services where you store personal files, check whether the terms permit AI training on user content
- Submit data deletion requests under CCPA (if in California) or GDPR (if in the EU) to services that may have scraped or processed your data without clear consent
- Follow the progress of the New York AI training bill and similar legislation in other states — the Transparency Coalition maintains a legislative tracker
The right to know whether your personal data trained a commercial AI system is a minimum baseline. Transparency legislation like New York’s is a first step toward making that baseline a legal reality rather than a request that companies are free to ignore.