When ShinyHunters dumped what it claimed was 50GB of Carhartt customer data in August, the headline number looked catastrophic: twenty-five million email addresses. Except it wasn't true. Troy Hunt's forensic review for Have I Been Pwned revealed the actual count: 12,933,413 genuine accounts, with the rest padded out by synthetic data designed to inflate the breach's apparent scale.
If you're a DPO facing a similar breach notification from a threat actor, here's what that gap between claimed and verified numbers means for your Article 33 and Article 34 obligations.
The Challenge
ShinyHunters claimed to have exfiltrated data on roughly 25 million individuals after Carhartt allegedly hired what the group called "a very unskilled and incompetent negotiator" to handle a $3.3 million extortion demand. The dataset appeared massive. Your first instinct in a similar scenario: assume the worst, notify your supervisory authority within 72 hours, and start drafting individual notifications under Article 34 if the breach poses a high risk to rights and freedoms.
But what if half the "affected individuals" don't exist?
The Carhartt incident wasn't just inflated by duplicates or formatting errors. The dataset contained millions of email addresses following TPC-DS synthetic data patterns: random domain strings like lkvb06fkzsjv.org, names paired with .edu addresses that didn't match real institutions, and customer records claiming residence in Montenegro at rates higher than the United States, where Carhartt is headquartered. Birth dates clustered in the early 1900s. Geographic distributions that made no commercial sense for a workwear retailer.
Without verification, you'd be reporting a breach twice the actual size and notifying millions of people who were never customers.
The Environment and Constraints
Article 33(1) gives you 72 hours to notify your supervisory authority once you become aware of a personal data breach. That clock doesn't pause while you verify attacker claims. Article 34 requires direct notification to affected individuals when the breach is likely to result in a high risk to their rights and freedoms, unless you've applied encryption or taken measures to mitigate that risk.
You're working against two competing pressures: the regulatory obligation to act quickly and the operational need to understand what actually happened. Overreporting wastes resources and erodes trust. Underreporting violates the GDPR and exposes you to administrative fines under Article 83.
In Carhartt's case, the challenge was technical and analytical. The initial extraction using an open-source email address parser returned nearly 25 million addresses. Manual review would take weeks. Automated deduplication catches obvious duplicates but won't flag sophisticated synthetic injection. You need a methodology that scales to millions of records while maintaining accuracy.
The Approach Taken
Hunt used a two-layer verification process combining AI analysis and manual review. First, he ran the dataset through OpenClaw, an AI tool designed to detect anomalies in breach data. The AI flagged patterns inconsistent with a legitimate customer database: the disproportionate number of .edu and .org domains for a retail brand, geographic distributions that didn't match Carhartt's market footprint, and demographic data (birth dates in the early 1900s) that didn't align with the company's customer base.
OpenClaw's first pass reduced the estimated genuine records from 24.8 million to 13.6 million by filtering out obvious synthetic entries.
Hunt then applied manual review to edge cases: Microsoft 365 duplicate addresses, accounts marked for deactivation, and records with suspicious metadata. This iterative process, moving between algorithmic filtering and human judgment, arrived at the final verified count of 12,933,413 accounts.
Critically, 83 percent of those verified accounts had already appeared in previous breaches, which affects your risk assessment under Articles 33 and 34. If the exposed data is already publicly available from prior incidents, the marginal risk to individuals may be lower, though this doesn't eliminate your notification obligations.
Results and Metrics
The verified dataset was roughly half the size of the attacker's claim. That's not a rounding error; it's the difference between notifying 12.9 million individuals and 25 million, between calculating breach-related costs on one scale versus another, and between presenting your supervisory authority with accurate versus inflated risk assessments.
The 83 percent overlap with prior breaches also matters for your Article 34 analysis. While previously compromised data doesn't excuse you from notification, it does inform how you describe the risk to individuals and what protective measures you recommend. Someone whose email and physical address have been circulating since 2019 faces different exposure than someone appearing in a breach dataset for the first time.
What They Would Do Differently
Hunt's analysis was retrospective, conducted after the data was already public. In a live incident, you won't have the luxury of a multi-day forensic review before your 72-hour Article 33 clock expires.
If you're facing a similar scenario, your initial notification to the supervisory authority should acknowledge uncertainty about the scope. Article 33(4) allows you to provide information in phases if you don't have complete details at the outset. Report what you know, flag what you're still verifying, and commit to a timeline for updates.
Invest in tooling before the breach happens. Open-source parsers and anomaly detection models exist; you don't need Hunt's specific stack, but you need something that scales beyond manual spreadsheet review. Test those tools against sample datasets so you understand their false positive rates and limitations.
And never treat attacker claims as ground truth for regulatory reporting. Threat actors have every incentive to exaggerate scope, whether to increase pressure during extortion negotiations or to amplify reputational damage. Your Article 33 notification should be based on your investigation, not their press release.
Takeaways for Your Team
Verify before you scale your response. Use AI-assisted tools to filter obvious synthetic data, but pair them with manual review of edge cases and statistical anomalies. A two-layer approach catches what automation alone misses.
Understand what "affected individuals" means in your context. If 83 percent of the records were already in prior breaches, that doesn't reduce your notification obligation, but it does change how you frame risk and what mitigation steps you recommend.
Phase your Article 33 reporting when scope is uncertain. The GDPR anticipates incomplete information at the 72-hour mark. Report your initial findings, explain what you're still verifying, and provide updates as your investigation progresses.
Don't let attacker inflation drive your compliance decisions. ShinyHunters claimed 25 million; the real number was half that. Your supervisory authority expects accuracy, not amplified threat actor narratives.
The Carhartt breach shows that the headline number is often wrong. Your job isn't to repeat it; it's to verify it, report it accurately, and respond proportionately to the actual risk.



