Skip to main content
EDPB Anonymisation FAQ: What ChangedSupervisory Authorities & Enforcement
4 min readFor Data Protection Officers (DPOs)

EDPB Anonymisation FAQ: What Changed

The recent EDPB plenary introduced new guidelines that could reshape how you handle anonymisation, web scraping, and blockchain. If you're responsible for data privacy, you're likely wondering how these changes impact your AI training pipelines, analytics platforms, and data retention strategies.

Here's what you need to know.

Understanding the Source of These Questions

The EDPB released draft guidelines on anonymisation and web scraping for generative AI on July 8, along with final blockchain guidance. These drafts are open for public consultation until October 30, 2026. The anonymisation guidelines now incorporate the CJEU's September 2025 ruling in C-413/23 P EDPS v SRB, reflecting updated case law.

If you're involved in building AI models, conducting web scraping, or managing analytics datasets, these guidelines are crucial for your compliance strategy.

Should You Use the Contextual or Simplified Approach?

You have two options: the contextual approach or the simplified approach.

The contextual approach evaluates whether data is anonymous based on the actual capabilities of entities that might try to re-identify individuals. This means data could be considered anonymous for your purposes even if a more sophisticated actor could potentially re-identify it.

The simplified approach treats data as personal if any relevant entity could use available means to re-identify someone. This approach is stricter but offers operational simplicity, as you don't need to assess capabilities for every potential recipient or threat actor.

For smaller teams without dedicated anonymisation expertise, the simplified approach is often safer. Larger enterprises with mature data governance might prefer the contextual approach for its flexibility, but it requires thorough documentation.

What Are the Three Anonymisation Criteria?

The EDPB framework uses three criteria to test anonymisation: no record isolation, no linkage, and no inference.

No record isolation means you can't single out an individual's record from the dataset. For example, identifying "the person who earns €200,000 and lives in a specific small town" would isolate that record.

No linkage prevents connecting records about the same person across different datasets. Matching anonymised purchase data to browsing data using shared attributes would create linkage.

No inference ensures you can't deduce new information about an individual with significant probability. If your dataset reveals that "people with these three attributes have a 95% likelihood of a specific diagnosis," you've enabled inference.

All three criteria must be met. If any fail, further analysis is needed to determine if the data is truly anonymous.

Lawful Basis for Scraping Public Websites for AI Training

This is a common sticking point. The guidelines clarify that legitimate interests (Article 6(1)(f)) can apply to web scraping for AI training, but the bar is high.

You'll need a legitimate interests assessment covering:

  • Purpose limitation: You can't scrape data for one purpose and then use it for unrelated AI training.
  • Transparency: Even if informing individuals personally is impractical, you must provide clear documentation on your website about your scraping practices.
  • legitimate interests assessment: Consider whether individuals reasonably expect their publicly posted data to be scraped for AI models. Public data doesn't automatically mean free use.

The guidelines also recommend scraping from reliable sources, recording timestamps, and validating data to meet the accuracy principle. For example, scraping outdated social media posts requires assessing their current accuracy and relevance.

Handling Accidental Collection of Special Category Data

This is the most challenging part of the guidance. Processing special categories of personal data is prohibited under Article 9 unless you have both a lawful basis under Article 6 and an exception under Article 9(2).

The EDPB references the GC & Others ruling (C-136/17) for incidental collection of special category data. You must act "within the framework of their responsibilities, powers, and capabilities" and implement measures to prevent collection and dissemination of such data.

If your scraper collects profile photos (biometric data) or health-related posts, you can't claim it's incidental without proper assessment:

  • Did you implement filters to exclude special category data?
  • Are you scraping from sources where such data is likely present?
  • Can you show you've minimized collection through technical design?

If you're scraping broadly, the "incidental collection" argument weakens. You need documented measures showing efforts to prevent it.

Blockchain and Controller/Processor Roles

The final blockchain guidelines clarify that GDPR roles don't always align with blockchain architectures.

In a permissioned blockchain where your organization controls transaction validation, you're likely a controller or joint controller. You'll need a joint controllership arrangement under Article 26 to specify responsibilities for transparency, data subject rights, and security measures.

In a public blockchain, the situation is less clear. The core principle is that whoever determines the purposes and means of processing is a controller, regardless of the technology.

The right to erasure remains a challenge. You must design your implementation so personal data can be deleted or made inaccessible, often by storing only hashes or encrypted pointers on-chain.

Next Steps

The anonymisation and web scraping guidelines are open for public consultation until October 30, 2026. If your organization has specific use cases not addressed, submit comments. The EDPB has incorporated stakeholder input from their Helsinki statement commitment to collaborative dialogue.

The final blockchain guidelines include a track changes version showing what changed after consultation. Review both versions to understand where the Board shifted based on feedback.

If you need to make decisions now, the draft versions reflect the Board's current thinking. Supervisory authorities will likely reference this framework in enforcement, even during consultation.

You Might Also Like