What happens when you scrape 20 years of internet data to train a generative AI model and then discover you can't identify a lawful basis that holds up under scrutiny? This isn't hypothetical. The European Data Protection Board's draft guidelines on web scraping for generative AI, published on July 7, 2026, read more like a forensic analysis of where AI developers are already failing than forward guidance.
Treat these guidelines as an incident report written before the enforcement actions arrive.
What Happened
The EDPB released draft guidelines addressing GDPR compliance for organizations scraping internet data to train generative AI systems. The guidance applies to entities that scrape data directly or via third parties, and those acquiring pre-scraped datasets from data brokers. The guidelines highlight six compliance failure points: data protection role allocation, lawful basis selection, data minimization, transparency obligations, accuracy requirements, and special category personal data handling.
The EDPB's consultation period remains open until October 30, 2026.
Timeline
- July 7, 2026: EDPB publishes draft guidelines on web scraping for generative AI
- October 30, 2026: Consultation period closes
This timeline is crucial. If you're building or procuring AI systems trained on internet-sourced data, you have three months to assess whether your current practices align with the EDPB's interpretation of controller obligations.
Which Controls Failed or Were Missing
Role allocation confusion: Organizations treated data scraping relationships as simple processor arrangements without analyzing whether they qualified as processor relationships, joint controllerships, or separate controller scenarios. The EDPB confirms that an AI developer acquiring a pre-scraped dataset bears independent controller responsibility for its own use of that data, regardless of what the scraper did upstream.
Consent as default legal fiction: Teams assumed that publicly available data meant consent to process it. The EDPB explicitly rejects this interpretation. Where personal data are collected indirectly and at scale, with no direct relationship to data subjects who are rarely identifiable, freely given informed consent is unworkable. Simply making data publicly available does not constitute consent to scraping.
Legitimate interests without documentation: Organizations relied on legitimate interests as a lawful basis but failed to document all three required elements: identifying a legitimate interest, demonstrating processing necessity, and conducting a balancing exercise showing that controller interests don't override data subject interests. The EDPB notes that growing public awareness of online data use may weigh favorably in a legitimate interests assessment, but this doesn't provide blanket authorization.
Data minimization as an afterthought: Collection architectures lacked filters to exclude sensitive data categories, high-risk websites (such as those directed at minors), or sites that prohibit scraping. Post-collection cleansing wasn't implemented. The EDPB calls this a "major challenge" but confirms that data minimization applies from pre-collection through output.
Transparency exemption as default: Teams invoked the disproportionate effort exemption for individual notifications without case-by-case justification. The EDPB provides a clear threshold: large-scale collection from varied sources spanning 20 years, with no direct identifiers, covering thousands or millions of individuals, where the controller published its privacy notice and excluded directly identifiable data, likely qualifies. Targeted collection from a closed group of 5,000 identifiable individuals spanning two years doesn't.
Special category personal data without safeguards: Organizations failed to implement measures preventing special category personal data collection or couldn't demonstrate that collection was genuinely incidental rather than deliberate. The EDPB references the GC & Others (C-136/17) search-engine case to provide a framework, but only where processing is analogous to a search engine, collection is incidental, assessing presence is genuinely difficult, and robust technical and organizational measures span the full lifecycle.
What the Relevant Standard Requires
Article 5(1)(a) (lawfulness): Personal data must be processed lawfully. This requires identifying and documenting a valid lawful basis before processing begins. The EDPB considers legitimate interests more viable than consent for web scraping but requires genuine engagement with all three limbs of the legitimate interests assessment.
Article 5(1)(c) (data minimization): Personal data must be adequate, relevant, and limited to what's necessary. This applies from pre-collection through output. Collection criteria must be defined in advance, with filters excluding categories that aren't necessary for the stated purpose.
Articles 13 and 14 (transparency): Controllers must provide information to data subjects. Where individual notification involves disproportionate effort, Article 14(5)(b) permits alternative measures, but this requires case-by-case justification considering the number of data subjects, age of data, and safeguards adopted. A publicly accessible privacy notice is mandatory and should include sources to the greatest extent possible, whether sources are publicly accessible, and crawler characteristics.
Article 5(1)(d) (accuracy): Personal data must be accurate and, where necessary, kept up to date. This applies to both collected data and model outputs. Inaccurate training data increases the likelihood of factually wrong or harmful outputs.
Article 9 (special categories): Processing special category personal data requires an Article 9(2) condition assessed on a case-by-case basis. Controllers should implement measures preventing collection where technically feasible.
Lessons and Action Items for Your Team
Map your role before you scrape: Assess whether you're acting as processor, joint controller, or separate controller. If you're acquiring third-party datasets, conduct upstream due diligence. You bear independent controller responsibility for your use of that data regardless of what happened before you acquired it.
Document your legitimate interests assessment now: If you're relying on legitimate interests, create a written assessment covering all three elements. Include the specific legitimate interest, why processing is necessary to achieve it, and a balancing exercise weighing your interests against data subject rights. Generic templates won't survive scrutiny.
Build minimization into collection architecture: Define collection criteria in advance. Implement filters excluding sensitive data categories, high-risk websites, and sites prohibiting scraping. Consider post-collection cleansing and synthetic data as a substitute for real personal data where viable.
Justify your transparency approach: If you're claiming disproportionate effort for individual notifications, document why your specific scenario qualifies. Publish a privacy notice that includes sources, whether they're publicly accessible, crawler characteristics, domain names and URLs in searchable format, date ranges, and contact details for originating controllers where you purchased data.
Timestamp and validate before training: Timestamp data at collection. Restrict scraping to credible sources where possible. Validate data before it enters training pipelines to meet accuracy obligations.
Prevent special category data collection: Implement filtering at collection, deletion post-collection, resistance to privacy attacks during development, and output filtering post-deployment with ongoing monitoring. If you're claiming the GC & Others framework applies, document how you meet each criterion.
The EDPB's guidelines don't resolve the tension between large-scale AI training and GDPR's individual-centric protections. They do, however, provide a compliance map. Publicly available data remains subject to GDPR. The compliance burden falls on the controller, and it arises before scraping begins.



