privacydeep-dive

New EU Rules on Scraping Your Public Photos for AI

The EDPB's new guidelines on web scraping for generative AI establish when companies can legally collect your public photos — and your rights to object.

On July 8, 2026, the European Data Protection Board published Guidelines 03/2026 on web scraping in the context of generative AI — the first comprehensive regulatory framework specifically addressing the large-scale collection of publicly available personal data for AI training purposes. The guidelines are currently in public consultation until October 2026, but they represent the most detailed statement yet from Europe’s top data protection authority on what AI companies are and aren’t allowed to do with data they collect from the web.

The personal data most at stake includes photos.

If you’ve published images of yourself — on social media, on a personal website, in forums, on review platforms, in comment sections — you are, with near certainty, present in the training datasets of at least some generative AI models. The question the EDPB guidelines address is whether that collection was legal under GDPR, and what you can do about it if it wasn’t.


What Web Scraping for AI Training Looks Like

Training a generative AI model requires enormous quantities of examples. For image-generating models, training datasets typically contain billions of images collected using automated tools that crawl publicly accessible URLs and download whatever they find. These datasets are built quickly, at scale, and without the involvement of the people depicted in the images.

Researchers who have studied the composition of major training datasets have documented that photos of private individuals — people who never intended their image to be used to train AI models — appear routinely. This includes social media profile pictures, photos posted in community forums, images from personal websites, pictures of people in backgrounds of other people’s photos, and images scraped from platforms where scraping is explicitly prohibited in the terms of service.

The legal question under GDPR is whether collecting, storing, and training on these photos is lawful when they contain personal data — that is, when a person can be identified from them, or from information linked to them.


What the EDPB Guidelines Say

The EDPB’s guidelines acknowledge that the legal answer is not a simple yes or no. They work through the GDPR’s legal basis framework and apply it to web scraping at scale. The key conclusions:

The guidelines are clear on one point: consent under Article 6(1)(a) GDPR cannot serve as the legal basis for collecting personal data through web scraping. Consent requires a “freely given, specific, informed and unambiguous indication” of agreement. Collecting millions of photos from across the internet at scale makes it practically impossible to obtain meaningful consent from each person depicted.

Some AI companies have argued that making content publicly available online implies consent to its use in AI training. The EDPB explicitly rejects this framing. Making information public — posting a photo to a social platform or personal website — is not equivalent to consenting to any conceivable use of it, including aggregation into an AI training dataset.

Legitimate interest is the primary option — with strict requirements

The most viable legal basis for web scraping under these guidelines is legitimate interest under Article 6(1)(f). A company can rely on legitimate interest if it has a genuine purpose, the processing is necessary for that purpose, and the legitimate interest is not overridden by the data subject’s rights and interests.

The EDPB requires a rigorous three-part test before a controller can rely on this basis for AI training:

Purpose test. Is there a genuine, specific, and lawful purpose being pursued? General references to “improving AI” or “advancing technology” are insufficient — the purpose must be identifiable and proportionate to the data being collected. This is a meaningful hurdle: training datasets of billions of images scraped from any publicly available source are hard to justify as necessary for a specific, defined purpose.

Necessity test. Is processing personal data necessary to achieve that purpose? Could the same result be achieved with anonymized data, synthetic data, or data from consenting participants? The EDPB pushes companies to justify why identifiable personal data is required rather than alternatives that would pose lower privacy risks.

Balancing test. Do the company’s interests outweigh the data subjects’ rights? This test requires considering the nature of the data being processed, the reasonable expectations of the data subjects, and the potential harms. The EDPB’s guidance on this point is the most practically significant part of the guidelines.

On the balancing test, the EDPB observes that most people who post photos online have a reasonable expectation that those photos will be seen by other users on the platform — not that they’ll be scraped into a proprietary AI training dataset maintained by a third-party company for commercial use. This asymmetry between what users expected and what actually happened weighs against the company in the balancing analysis.

Special category data

Certain types of information require additional legal grounding even if legitimate interest would otherwise apply. Under Article 9 GDPR, photos that reveal — or could reasonably be used to infer — racial or ethnic origin, religious beliefs, health conditions, sexual orientation, or political opinions qualify as “special category data.” Processing these requires either explicit consent from the data subject or a specific, narrowly defined exception.

This is a practical problem for large-scale scraping. A dataset of billions of images will inevitably contain photos that reveal sensitive information about the people depicted — illness, religion, relationships, physical characteristics associated with ethnic background. The EDPB acknowledges that incidental and unintended processing of such images within a broader training pipeline may be permissible if robust risk-mitigation measures are applied throughout the data lifecycle, but sets a high bar for what those measures must look like.


Your Rights Under These Guidelines

If the company scraping your photos is subject to GDPR — which generally means it has an establishment in the EU, offers services to EU residents, or monitors EU residents’ behavior — you have enforceable rights that apply to AI training data specifically.

The right to object. Article 21 GDPR gives you the right to object to processing of your personal data based on legitimate interest. When you exercise this right, the company must stop processing your data unless it can demonstrate “compelling legitimate grounds” that override your interests and rights. The EDPB’s guidelines reaffirm that this right applies to personal data collected through web scraping and used for AI training — the fact that your photo is publicly available does not eliminate your right to object to this specific use of it.

The right to erasure. Article 17 GDPR — the “right to be forgotten” — allows you to request that a company delete your personal data. For AI training datasets, this obligation is clearest with respect to the raw collected data (deleting your photo from the dataset before or after training). Deleting data from an already-trained model is a technically unsolved problem that the guidelines address only in general terms, leaving the specifics for future guidance and national enforcement practice.

The right of access. Article 15 GDPR allows you to request confirmation of whether a company holds your personal data and a copy of that data. Exercising this right against an AI company to determine whether your photos are in their training dataset is theoretically possible, though challenging given the scale of these datasets and the difficulty of locating a specific person’s data within them.

The practical limitation of all of these rights is enforcement. Exercising them requires knowing which companies have scraped your data — which most people don’t know and can’t easily find out. Enforcement through national data protection authorities can be slow and resource-intensive. The guidelines create the legal framework; making it actionable at scale requires the enforcement machinery to follow through.


What This Doesn’t Cover

The guidelines apply to GDPR-regulated processing. Several important categories fall outside their scope:

US-based AI companies training exclusively on US-resident data. If a company processes only data of US residents through US infrastructure, without an EU establishment and without specifically targeting EU users, GDPR doesn’t apply. The EDPB guidelines have no reach over that processing.

Already-trained models. The guidelines address the data collection and training process. They do not directly resolve what an AI company must do about a model that was trained on personal data collected without a sufficient legal basis before these guidelines existed. Some national data protection authorities have taken positions on this; others have not. It remains an area of active regulatory development.

Genuinely anonymous data. If photos are sufficiently anonymized that individuals cannot be identified — including by combining the data with other available information — they fall outside GDPR’s scope entirely. The guidelines reaffirm that truly anonymous data is not personal data. They also note that anonymizing photos is technically difficult and that claims of anonymization deserve scrutiny, since face-blurring or metadata removal doesn’t always produce data that is truly non-identifiable.

Non-photographic personal data scraped from public sources. The guidelines apply broadly to any personal data collected through web scraping, not just photos. Text you’ve published, professional profiles, comments, and other publicly available information containing personal data is covered by the same framework. Photos are the most intuitive example, but they’re not the only category at stake.


Practical Implications for How You Manage Personal Photos

The EDPB guidelines have implications for a distinction most people navigate intuitively but rarely make explicit: the difference between private storage and public sharing.

Photos you store in a private app — one where your images are not publicly accessible via a URL and are not indexed by search engines — are not reachable by web scraping in the first place. The GDPR framework the guidelines describe doesn’t apply because the data collection step can’t happen. The issue arises when photos are published to publicly accessible locations.

This creates a real distinction between:

  • Private storage: Personal files and photos stored in a private cloud service, behind authentication, without public URLs. These are not scrapeable. The EDPB guidelines are irrelevant because the data isn’t accessible to scrapers in the first place.

  • Semi-public sharing: Photos shared with a defined group — friends on a social platform, members of a private group, recipients of a shared link — where the accessibility depends on platform settings and link circulation. Platform security controls determine whether these are scrapeable.

  • Public sharing: Photos published with publicly accessible URLs, on platforms that don’t require login, or on personal websites. These are the photos the EDPB guidelines are most directly about.

The guidelines don’t tell you not to post publicly. They establish legal standards for what can be done with publicly posted data by third parties. But the practical takeaway is the same one that broader privacy considerations have been pushing toward: keeping personal photos and files in private storage, sharing selectively and through controlled mechanisms, and being deliberate about what ends up with a public URL that anyone — including an AI training crawler — can reach.


The Consultation and What Comes Next

The guidelines are in public consultation until the end of October 2026. After that, the EDPB will review consultation responses and adopt a final version, expected before the end of the year.

That final version will be guidance rather than binding law — but EDPB guidelines carry significant weight in how national data protection authorities across Europe enforce GDPR. AI companies operating in Europe will be expected to have aligned their practices with the guidelines or to demonstrate clearly why their approach satisfies the underlying legal requirements through a different route.

The consultation period provides an opportunity for civil society organizations, privacy advocates, and industry groups to comment. Several major privacy organizations have already submitted preliminary reactions, most focusing on the practical enforceability of the right to object against large AI training pipelines, and the difficulty of verifying whether any given company has actually complied with a deletion request.


The Larger Pattern

The EDPB’s web scraping guidelines fit into a broader pattern of European regulators applying the existing GDPR framework to AI practices rather than waiting for AI-specific legislation. The same plenary session that adopted the scraping guidelines also published new guidelines on anonymization techniques — another core concept for AI data processing — and finalized separate blockchain guidance.

This approach — adapting existing privacy law to new technology through detailed regulatory guidance — creates an evolving body of requirements that changes how AI companies must operate in Europe, without requiring new legislation for each new practice. It’s slower than the pace at which AI capabilities develop, but it maintains continuous regulatory pressure on data practices that would otherwise operate in a legal vacuum.

The scraping guidelines are one piece. The EU AI Act’s transparency requirements for training data, the Digital Omnibus package’s proposed changes to the definition of personal data, and ongoing enforcement by national data protection authorities against specific companies are others. Together they form a regulatory environment that takes the use of publicly available personal data for AI training significantly more seriously than any other major jurisdiction currently does.

For individuals, the practical significance of the guidelines depends partly on jurisdiction and partly on which AI companies they want to exercise rights against. For everyone, the underlying question the guidelines address — what can be done with your publicly available photos, and by whom, without your active consent — is one worth understanding regardless of which legal framework applies to your specific situation.

Your memories deserve better than an ad platform.

Try daftei free →
← All posts