Quasi-Identifier
A quasi-identifier is a piece of information that cannot single out a specific person on its own, but can identify someone when combined with other available information. For example, attributes such as a postal code, date of birth, or gender may each be shared by many people, yet together they may point to one individual. Because quasi-identifiers can contribute to re-identifying a person, data containing them is generally still treated as personal data and remains subject to the GDPR.
A quasi-identifier is an attribute that is not uniquely identifying in isolation but can identify an individual, or narrow identification to a specific record or subset of records, through association or combination with other attributes or auxiliary data. Quasi-identifiers are central to re-identification risk assessment and to statistical disclosure control techniques (for example, k-anonymity and related models), where the combination of quasi-identifiers across a dataset determines whether individuals can be distinguished. Under the GDPR framework, the presence of quasi-identifiers means data cannot be treated as anonymous where re-identification remains reasonably likely; per Recital 26 and the identifiability test, data that can be attributed to an individual by means reasonably likely to be used generally remains personal data. Note in particular that pseudonymized data, which typically retains quasi-identifiers, is expressly defined as personal data under Article 4(5) GDPR and does not qualify as anonymous. Whether a given attribute functions as a quasi-identifier is context-dependent, turning on the auxiliary information available and the dataset in question, and should be determined by assessment rather than treated as a fixed classification. This definition addresses the general concept; the threshold for effective anonymization and the treatment of re-identification risk can vary between regulators and remains subject to evolving guidance.
Why it matters
Quasi-identifiers matter because they determine whether a dataset is genuinely anonymous or remains personal data subject to the GDPR. Organizations frequently assume that removing direct identifiers such as names or ID numbers renders data anonymous, but attributes that seem innocuous in isolation, such as a postal code, date of birth, or gender, can combine to single out an individual. Where re-identification remains reasonably likely by means reasonably likely to be used, the data generally continues to be personal data under the identifiability test reflected in Recital 26, and the full range of GDPR obligations continues to apply.
This has direct consequences for how data is classified and handled. Pseudonymized data, which typically retains quasi-identifiers, is expressly defined as personal data under Article 4(5) GDPR and does not qualify as anonymous. Treating such data as outside the scope of the Regulation is a common and significant misclassification, because it can lead to inadequate safeguards, unlawful further processing, or unlawful transfers on the mistaken basis that no personal data is involved.
Because whether an attribute functions as a quasi-identifier depends on the auxiliary information available and the specific dataset, this is a matter for assessment rather than a fixed label. The threshold for effective anonymization and the treatment of re-identification risk can vary between regulators and remains subject to evolving guidance, so practitioners should document their reasoning and revisit it as external data availability and analytical techniques change.
Who it's relevant to
Inside Quasi-Identifier
Common questions
Answers to the questions practitioners most commonly ask about Quasi-Identifier.