Skip to main content
Category: Data Classification & Identifiers

Quasi-Identifier

Also known as: Indirect Identifier
Simply put

A quasi-identifier is a piece of information that cannot single out a specific person on its own, but can identify someone when combined with other available information. For example, attributes such as a postal code, date of birth, or gender may each be shared by many people, yet together they may point to one individual. Because quasi-identifiers can contribute to re-identifying a person, data containing them is generally still treated as personal data and remains subject to the GDPR.

Formal definition

A quasi-identifier is an attribute that is not uniquely identifying in isolation but can identify an individual, or narrow identification to a specific record or subset of records, through association or combination with other attributes or auxiliary data. Quasi-identifiers are central to re-identification risk assessment and to statistical disclosure control techniques (for example, k-anonymity and related models), where the combination of quasi-identifiers across a dataset determines whether individuals can be distinguished. Under the GDPR framework, the presence of quasi-identifiers means data cannot be treated as anonymous where re-identification remains reasonably likely; per Recital 26 and the identifiability test, data that can be attributed to an individual by means reasonably likely to be used generally remains personal data. Note in particular that pseudonymized data, which typically retains quasi-identifiers, is expressly defined as personal data under Article 4(5) GDPR and does not qualify as anonymous. Whether a given attribute functions as a quasi-identifier is context-dependent, turning on the auxiliary information available and the dataset in question, and should be determined by assessment rather than treated as a fixed classification. This definition addresses the general concept; the threshold for effective anonymization and the treatment of re-identification risk can vary between regulators and remains subject to evolving guidance.

Why it matters

Quasi-identifiers matter because they determine whether a dataset is genuinely anonymous or remains personal data subject to the GDPR. Organizations frequently assume that removing direct identifiers such as names or ID numbers renders data anonymous, but attributes that seem innocuous in isolation, such as a postal code, date of birth, or gender, can combine to single out an individual. Where re-identification remains reasonably likely by means reasonably likely to be used, the data generally continues to be personal data under the identifiability test reflected in Recital 26, and the full range of GDPR obligations continues to apply.

This has direct consequences for how data is classified and handled. Pseudonymized data, which typically retains quasi-identifiers, is expressly defined as personal data under Article 4(5) GDPR and does not qualify as anonymous. Treating such data as outside the scope of the Regulation is a common and significant misclassification, because it can lead to inadequate safeguards, unlawful further processing, or unlawful transfers on the mistaken basis that no personal data is involved.

Because whether an attribute functions as a quasi-identifier depends on the auxiliary information available and the specific dataset, this is a matter for assessment rather than a fixed label. The threshold for effective anonymization and the treatment of re-identification risk can vary between regulators and remains subject to evolving guidance, so practitioners should document their reasoning and revisit it as external data availability and analytical techniques change.

Who it's relevant to

Data protection officers and compliance leads
DPOs and compliance leads need to assess whether datasets described internally as anonymous or de-identified actually contain quasi-identifiers that keep the data within GDPR scope. This affects records of processing, data classification, and decisions about whether obligations such as data subject rights and transfer controls apply. Because pseudonymized data retaining quasi-identifiers remains personal data under Article 4(5), a documented identifiability assessment is generally advisable rather than relying on a label.
Privacy engineers and data scientists
Those building or operating anonymization and pseudonymization pipelines work directly with quasi-identifiers when applying techniques such as k-anonymity and related statistical disclosure control models. Their assessments of which attributes act as quasi-identifiers, and what auxiliary data is reasonably available, shape the residual re-identification risk and therefore the legal status of the output.
Lawyers advising on data sharing and transfers
Legal advisers assessing data sharing, publication, or cross-border transfers should consider whether datasets contain quasi-identifiers, because this bears on whether the data is personal data at all. Given that the anonymization threshold and treatment of re-identification risk can vary between regulators and continue to evolve, advice should be framed around the specific dataset and current guidance rather than presented as a settled outcome.

Inside Quasi-Identifier

Definition of a quasi-identifier
A quasi-identifier is an attribute or a combination of attributes that does not directly identify an individual on its own but can, when linked with other available information, single out or re-identify a data subject. Common examples typically discussed include date of birth, postal code, and gender, which in combination can narrow a population to a single person.
Role in re-identification risk
Quasi-identifiers are central to assessing re-identification risk because their combination can uniquely or near-uniquely distinguish individuals. Assessing this risk generally requires considering the auxiliary datasets and means reasonably likely to be used to identify a person, consistent with the identifiability test in Recital 26.
Relationship to direct identifiers
Unlike direct identifiers (such as a name or national identification number), which point to an individual on their own, quasi-identifiers typically require linkage or combination to achieve identification. Both categories are relevant when evaluating whether data relate to an identified or identifiable natural person under Article 4(1).
Interaction with pseudonymization and anonymization
Quasi-identifiers are a key consideration when deciding whether a dataset is anonymous or merely pseudonymized. Pseudonymized data remain personal data under the GDPR (see Recital 26 and the definition in Article 4(5)), because re-identification remains possible using additional information or quasi-identifier combinations. Anonymization, by contrast, is generally understood to require that identification is no longer reasonably likely, which involves addressing quasi-identifiers; whether a given dataset is truly anonymous is subject to assessment and can be contested.

Common questions

Answers to the questions practitioners most commonly ask about Quasi-Identifier.

Is data safe to share once direct identifiers like names have been removed but quasi-identifiers remain?
Not necessarily. Removing direct identifiers does not by itself render data anonymous. Quasi-identifiers, such as date of birth, postcode, or gender, can be combined with each other or with external datasets to single out an individual. Where re-identification remains reasonably likely, the data typically continue to constitute personal data and remain within the scope of the GDPR. A risk assessment considering the means reasonably likely to be used to identify a person is generally required before treating such data as outside the Regulation.
Does stripping direct identifiers turn a dataset into anonymous rather than pseudonymous data?
Generally no. Where quasi-identifiers persist and re-identification is still possible using additional information, the data are typically pseudonymized rather than anonymized. Under the GDPR, pseudonymized data are expressly treated as personal data (see Recital 26 and the definition of pseudonymisation in Article 4(5)) and remain subject to the Regulation. True anonymization, by contrast, requires that identification is no longer reasonably possible, which is a high and context-dependent threshold. Practitioners should verify their assessment against the current official text and applicable regulatory guidance.
How can quasi-identifiers be identified within a dataset?
Identifying quasi-identifiers typically involves reviewing which attributes, alone or in combination, could distinguish or single out an individual, and considering what external or auxiliary datasets might realistically be linked. Common examples include demographic and geographic attributes, but the assessment is context specific and depends on the population, the granularity of the values, and the data reasonably available to potential re-identifiers. This is a matter of case-by-case judgment rather than a fixed list.
What techniques are commonly used to reduce re-identification risk from quasi-identifiers?
Techniques frequently discussed include generalization, suppression, aggregation, and applying formal models such as k-anonymity and related approaches. These measures aim to reduce the likelihood that combinations of quasi-identifiers single out an individual. The appropriate technique and its parameters depend on the use case and the acceptable residual risk, and effectiveness should be tested rather than assumed. No single technique guarantees anonymization in all circumstances.
Does treating quasi-identifiers reduce GDPR obligations?
It depends on the outcome. Where treatment results in pseudonymized data, GDPR obligations generally continue to apply, though pseudonymisation is recognized as a security and data-protection-by-design measure that can support compliance. Only where the data are genuinely anonymized such that identification is no longer reasonably possible would they typically fall outside the Regulation. Because that threshold is high and context dependent, the reduction in obligations should be assessed rather than presumed.
How should re-identification risk from quasi-identifiers be documented in a compliance program?
Organizations generally document the quasi-identifiers considered, the external datasets assessed as reasonably available, the techniques applied, the residual risk determination, and the rationale for classifying the data as pseudonymized or anonymized. Where processing is likely to result in high risk, this analysis may form part of a Data Protection Impact Assessment. Documentation should be revisited over time, as advances in available data and techniques can change the re-identification risk, and approaches may differ between regulators.

Common misconceptions

Removing direct identifiers such as names makes a dataset anonymous.
Stripping names or other direct identifiers does not necessarily render data anonymous. Remaining quasi-identifiers can, in combination, still allow re-identification. Where re-identification remains reasonably likely, the data are generally treated as personal data (or as pseudonymized data), not anonymous, subject to a case-by-case assessment.
Pseudonymized data fall outside the GDPR.
Pseudonymized data are still personal data under the GDPR, as reflected in Recital 26 and the definition in Article 4(5), because the data subject can be re-identified using additional information, which may include quasi-identifier combinations. Pseudonymization is a risk-reducing safeguard, not a route out of the Regulation's scope.
An attribute is either always a quasi-identifier or never one.
Whether an attribute functions as a quasi-identifier depends on context, including the population, the granularity of the values, and the auxiliary data reasonably likely to be available. The same attribute may pose negligible risk in one dataset and meaningful re-identification risk in another, so classification is context-dependent and should be assessed rather than fixed.

Best practices

Identify and document candidate quasi-identifiers in each dataset, considering combinations of attributes rather than individual fields in isolation.
Assess re-identification risk in light of auxiliary data and the means reasonably likely to be used to identify individuals, consistent with the identifiability test in Recital 26, and record the assessment.
Treat pseudonymized data as personal data subject to the GDPR (Recital 26, Article 4(5)), and do not assume that removing direct identifiers alone achieves anonymization.
Apply and evaluate techniques that reduce quasi-identifier risk, such as generalization or reducing data granularity, and reassess whether identification remains reasonably likely after applying them.
Re-evaluate quasi-identifier and re-identification risk periodically and when datasets are combined or new auxiliary information becomes available, since context can change the classification.
Where a dataset's status as anonymous is uncertain or contested, document the reasoning and consult current regulatory guidance, recognizing that assessments may diverge and should be verified against the applicable official text.