The challenge isn't whether you can rely on legitimate interests for AI web scraping. It's that many teams treat the legitimate interests assessment like a checkbox exercise rather than a defensible compliance framework.
When the EDPB published Guidelines 03/2026 on July 8, 2026, they made one thing clear: legitimate interests will often be the most viable lawful basis for scraping personal data to train generative AI models. But viable doesn't mean easy. The Guidelines outline compliance requirements that expose a fundamental gap in how many organizations approach AI data collection. They're building massive datasets without the documented risk analysis that Article 6(1)(f) demands.
The Challenge
Your AI team needs web-scraped data at scale. Your legal team knows consent is impractical for billions of data points collected from public sources. So you land on legitimate interests as your lawful basis.
Here's where the wheels come off: legitimate interests isn't a declaration, it's a three-part test you must document before you start processing. You need to identify a legitimate purpose, demonstrate that scraping is necessary for that purpose, and balance the processing against the rights and interests of the individuals whose data you're collecting. The EDPB's guidance makes clear that "we're training an AI model" doesn't satisfy any of these elements on its own.
The regulatory risk compounds when you consider special category personal data. Under Article 9, processing health information, biometric data, or other sensitive categories is prohibited unless you meet a specific condition, even if you collected that data accidentally. The EDPB, citing the Court of Justice ruling in GC & Others (C-136/17), indicates you might justify incidental SCPD processing if your activity resembles a search engine and you implement measures to prevent dissemination of that data. But "might" and "resembles" don't give you much to work with when a supervisory authority asks why your training dataset contains medical records.
Regulatory Constraints
The Guidelines operate in a regulatory context where deployers face enforcement action if they fail to assess whether upstream AI models were developed through unlawful processing. This shifts the compliance burden beyond the companies doing the scraping. If you're integrating a third-party model into your EU-facing products, you're expected to verify its provenance.
That verification requirement creates a documentation problem. If you're the controller scraping data, you need records that demonstrate:
- Precise collection criteria that limit data to what's necessary
- Exclusion of websites likely to contain data about vulnerable individuals
- Respect for technical signals like robots.txt, ai.txt files, and CAPTCHAs
- Filters to prevent collection of SCPD
- A complete list of data sources, or where that's impossible, identification of source types and an explanation of omissions
These aren't optional governance practices. They're transparency obligations under Articles 13 and 14. The EDPB states that a generic AI privacy notice won't satisfy these requirements if it doesn't explain the characteristics of your crawlers and provide source documentation to the greatest extent possible.
Structured Approach
Organizations that treat the legitimate interests assessment as a living compliance tool rather than a one-time form tend to structure their approach around three operational layers.
First, they scope the purpose with specificity. "Training a generative AI model" is too broad to support a necessity argument. You need to articulate what the model does, why web-scraped personal data is required for that function, and why alternative data sources won't work. This level of detail forces you to confront whether you're actually applying data minimization or just collecting everything available.
Second, they implement technical controls before scraping begins. This means defining collection criteria that exclude categories of websites, not just individual URLs, where personal data about children, health conditions, or private communications is likely to appear. It means configuring crawlers to honor robots.txt and ai.txt files as a baseline, not as an optional courtesy. And it means building filters that can identify and delete SCPD immediately upon collection, rather than hoping to clean datasets later.
Third, they extend compliance beyond model training into deployment. The Guidelines emphasize continuous monitoring of system outputs and measures to prevent generation of SCPD. In practice, this requires output filters, prompt restrictions, and regular verification that your controls actually work. If your model can be prompted to produce health information about identifiable individuals, your upstream filtering failed.
Results and Metrics
The EDPB's consultation period runs until October 30, 2026, so the final Guidelines may shift based on industry feedback. But the core compliance framework is unlikely to change: legitimate interests requires documented assessment, data minimization applies regardless of dataset size, and incidental SCPD creates regulatory risk even when unintended.
For organizations that built their scraping operations before these Guidelines, the documentation gap is the most immediate issue. If you can't produce a legitimate interests assessment that addresses the EDPB's considerations, balancing factors, mitigating measures, impact on individuals, you don't have a defensible lawful basis. That's not a theoretical problem. It's the first thing a supervisory authority will request during an investigation.
The transparency obligation creates a second pressure point. If your privacy notice doesn't identify scraping sources with the level of detail the Guidelines require, you're in breach of Articles 13 and 14 regardless of whether your underlying processing is lawful. This is particularly acute for companies that scraped data years ago and now need to reconstruct source lists from incomplete records.
Strategic Adjustments
The pattern that emerges from the Guidelines is that compliance decisions made at the data collection stage determine your regulatory exposure at deployment. Organizations that deferred SCPD filtering to post-training cleanup now face a choice: retrain models with proper controls, or accept ongoing risk that outputs will violate Article 9.
If you're building a new AI capability, the decision tree is clearer. Before you scrape:
- Document your legitimate interests assessment with the specificity the EDPB expects, including identification of mitigating measures
- Configure exclusions for website categories likely to contain SCPD or data about vulnerable individuals
- Implement collection filters, not just post-processing cleanup
- Build source documentation into your crawling workflow, not as an afterthought
If you're already operating a model trained on scraped data, your priority is evidentiary. Can you demonstrate that your original processing met GDPR requirements? If not, can you implement controls now that reduce risk to an acceptable level? The EDPB's emphasis on "responsibilities, powers, and capabilities" suggests that supervisory authorities will evaluate your compliance efforts against what's technically feasible for an organization of your resources, which means larger AI developers face higher expectations.
Takeaways for Your Team
The legitimate interests assessment is not a legal opinion you commission once and file away. It's a compliance artifact that must reflect the actual technical measures you've implemented and the ongoing monitoring you perform. If your assessment says you're excluding health-related websites but your crawler logs show you're not, the document becomes evidence against you.
Data minimization in AI contexts means defining what you won't collect, not just what you will. The EDPB's guidance on excluding vulnerable populations, respecting technical signals, and filtering SCPD creates a floor for responsible scraping. Your job is to document how you meet that floor and what additional measures you've implemented based on your specific use case.
Transparency obligations extend beyond a privacy notice URL in your footer. You need to explain the characteristics of your crawlers and provide source documentation that allows individuals to understand whether their data was processed. Where complete source lists aren't feasible, you must explain the omission, "we scraped millions of websites" doesn't satisfy Article 14.
For deployers and enterprise customers, the Guidelines confirm that you can't outsource GDPR compliance to your AI processor. If you're integrating a third-party model into EU-facing operations, you're expected to assess whether that model was developed lawfully. This means asking vendors for evidence of their legitimate interests assessments, data minimization measures, and SCPD controls. If they can't produce that documentation, you're assuming their regulatory risk.
The consultation period closing October 30, 2026 offers a narrow window to shape the final Guidelines. If your current scraping practices conflict with the EDPB's expectations, submit feedback that explains the operational challenges. But don't mistake consultation for negotiation, the core principles of lawful basis, data minimization, and SCPD protection aren't optional under the GDPR. The Guidelines just make explicit what was always required.



