Many believe that if you're publishing open datasets, running them through a privacy-by-design framework at the start is enough to ensure compliance. The idea is that building privacy in from day one will prevent the compliance failures that reactive organizations face. It's a comforting notion, but it's incomplete.
Why Privacy-by-Design Alone Isn't Enough
Privacy-by-design is often seen as a one-time decision rather than what Article 25 actually requires: ongoing data protection measures that adapt throughout processing. The real issue with open data isn't that teams skip privacy-by-design; it's that they apply it once at the project's start and assume the work is done.
Open datasets evolve. You might publish a transport usage dataset with aggregated counts. Six months later, someone could link it with another public dataset and re-identify individuals. Adding new data sources or changing your retention schedule creates fresh privacy risks, but your initial privacy-by-design assessment remains untouched.
The truth is, privacy-by-design principles don't exempt you from Article 5’s accountability obligation. You need to demonstrate compliance at every stage of processing. A solid design phase is important, but it won't catch the risks that emerge when your open data meets real-world usage patterns.
The Evidence
Article 25 mandates implementing appropriate technical and organizational measures "at the time of the determination of the means for processing and at the time of the processing itself." That second phrase is crucial. The regulation anticipates that design decisions alone won't suffice.
Your transparency obligations under Articles 13 and 14 don't stop after initial publication. Material changes in how you process open data trigger new transparency obligations. If you're relying on legitimate interests under Article 6(1)(f), your legitimate interests assessment must reflect current processing, not your original intent.
Consider this: you publish an anonymized health dataset. Initially, the anonymization technique was deemed sufficient. Two years later, computational advances make re-identification feasible with publicly available tools. Did your privacy-by-design protect you? No. Article 32 requires security measures "taking into account the state of the art." Your measures must evolve with technical capability.
Supervisory authorities have been clear about this in enforcement. When assessing accountability under Article 5(2), they look for evidence of ongoing compliance verification, not just initial design documentation. Your records management system needs to capture this continuous assessment, not just the kickoff workshop notes.
What to Do Instead
Treat privacy-by-design as the floor, not the ceiling. Build ongoing mechanisms into your records management workflow:
Establish review triggers tied to dataset changes. Every time you add a data source, modify aggregation methods, or extend retention periods, trigger a documented reassessment. Don't wait for annual compliance reviews. The change itself is the trigger.
Create monitoring protocols for linkage risk. Your open datasets exist in an ecosystem. Set up quarterly reviews to actively test whether your published data can be combined with other public sources to enable re-identification. Document what you find and what measures you adjust.
Version your privacy assessments with the data itself. If you're publishing dataset v3.2, you should have a corresponding privacy assessment v3.2. Your records management system should enforce this pairing. When someone requests historical data, they should see the privacy analysis that governed that specific version.
Build feedback loops from data users. Your Article 25 obligations include considering "the nature, scope, context and purposes of processing." That context changes as users apply your data in unexpected ways. Create formal channels for users to report potential privacy issues and document how you responded.
Measure your effectiveness through incident near-misses. Track how often your ongoing reviews catch risks that your initial design missed. If that number is zero, your review process isn't working. A healthy privacy program catches problems before they become personal data breaches.
When Privacy-by-Design Is Right
Privacy-by-design principles are essential for open data. If you're waiting until after publication to think about privacy, you've already failed. The upfront work prevents entire categories of risk that no amount of monitoring can fix later.
When determining aggregation thresholds, choosing what fields to publish, or setting access controls, those design decisions create constraints that protect privacy by default. You can't bolt on proper anonymization after publishing granular records. The initial architecture defines what's possible.
For records managers, privacy-by-design helps build the classification and retention schedules that open data depends on. If you haven't embedded privacy considerations into how you categorize and lifecycle records from creation, your open data program will constantly struggle with scope decisions.
The conventional wisdom also holds when dealing with Article 89 processing for archiving, research, or statistical purposes. Those derogations assume you've implemented appropriate safeguards "by design and by default." Without that foundation, you can't rely on the flexibility Article 89 provides.
But here's the distinction: privacy-by-design is necessary, not sufficient. Your records management practice needs both the initial rigor and ongoing verification. Treat the design phase as the beginning of your accountability obligation, not the end of it.



