Skip to main content
Category: Data Classification & Identifiers

Linkability

Simply put

Linkability refers to the possibility that separate pieces of information about a person can be connected or logically associated with one another. When data is linkable, records that might seem unrelated on their own can be tied together to build a fuller picture of an individual. Reducing linkability is generally one goal of privacy-protective techniques, though whether data is linkable depends on the specific context and the other information available.

Formal definition

Linkability describes the property whereby information about or related to an individual can be logically associated with other information about that same individual, based on the evidence provided (see NIST glossary on 'linkable information'). In privacy engineering, it is typically treated as a risk dimension: the greater the capacity to correlate distinct data items or records to a single data subject, the higher the potential re-identification and profiling risk. Assessment of linkability is context-dependent and turns on the auxiliary information reasonably available to a given party; the evidence supplied here does not itself establish thresholds, quantitative measures, or a definitive taxonomy, so practitioners should validate scoping and measurement approaches against current authoritative guidance. Note that the concept as defined here derives from technical/glossary sources rather than a specific GDPR article, and the term should not be confused with the unrelated organizations named 'LinkAbility' appearing in the evidence.

Why it matters

Linkability sits at the heart of re-identification and profiling risk. Individual data items may appear innocuous in isolation, but when they can be logically associated with other information about the same individual, they can be combined to build a richer and potentially identifying picture. This is why linkability is generally treated as a risk dimension in privacy engineering rather than a binary state: the concern is not only what a single dataset reveals, but what it reveals once connected to auxiliary information reasonably available to a given party.

Because linkability is context-dependent, its significance shifts with the data environment. Data that is difficult to link in the hands of one party may become readily linkable to another party who holds additional information. This has direct implications for whether data can be considered anonymous or merely pseudonymous, for the scope of data minimisation efforts, and for how organisations assess residual risk when sharing or publishing datasets.

The evidence supplied here does not establish quantitative thresholds, measurement methods, or a definitive taxonomy for linkability, so organisations should not treat any single characterisation as settled. Practitioners should validate scoping and measurement approaches against current authoritative guidance and assess linkability in light of the specific auxiliary information available in their particular context.

Who it's relevant to

Privacy Engineers and Data Architects
Privacy engineers treat linkability as a design-time risk dimension. When structuring datasets, pipelines, and identifiers, they aim to limit the extent to which distinct data items can be correlated to a single individual, and they assess residual linkability against the auxiliary information a given party might hold. Because the evidence here does not define thresholds or measurement methods, engineers should validate their scoping and measurement approaches against current authoritative guidance.
Data Protection Officers and Compliance Leads
For DPOs and compliance leads, linkability is central to judging whether data is anonymous, pseudonymous, or identifiable, and to evaluating re-identification and profiling risk. It informs data minimisation decisions and risk assessments for data sharing or publication. Because linkability is context-dependent and the concept derives from technical sources rather than a specific GDPR article, these assessments should be documented with reference to the reasonably available auxiliary information in the relevant context.
Data Scientists and Analysts Handling Shared or Published Datasets
Those preparing datasets for internal analysis, sharing, or release should consider that records which appear unrelated may be linkable when combined with other information. Understanding linkability helps them evaluate whether de-identification is adequate for the intended recipient and use, recognising that data difficult to link for one party may be linkable to another with additional data.

Inside Linkability

Linkability (core concept)
The degree to which two or more records, data items, or observations within a dataset (or across datasets) can be related to one another, whether or not the linked information can be attributed to a specific identified individual. Linkability is one of the privacy risks addressed in data protection guidance on anonymisation and pseudonymisation.
Relationship to identifiability
Linkability is distinct from, but related to, the broader question of whether data qualifies as personal data. Even where direct identifiers are removed, the ability to link records can contribute to a risk of re-identification when combined with other available information, and should therefore be assessed as part of an overall identifiability analysis.
Pseudonymisation and linkability
Pseudonymisation typically reduces, but does not eliminate, linkability, because records tied to the same pseudonym remain linkable and the mapping to identifiers may still exist. As a result, pseudonymised data generally remains personal data subject to the GDPR.
Singling out and inference
Linkability is commonly discussed alongside two other re-identification risks referenced in regulatory guidance: singling out (isolating an individual within a dataset) and inference (deducing information about an individual). These three risks are typically evaluated together when assessing whether data is genuinely anonymous.
Cross-dataset dimension
Linkability can arise internally, between records in a single dataset, or externally, by combining a dataset with other data sources. Assessment should generally take account of information reasonably available to parties who might attempt linkage, subject to the specific circumstances.

Common questions

Answers to the questions practitioners most commonly ask about Linkability.

Does removing direct identifiers like names eliminate linkability?
Not necessarily. Linkability concerns the ability to connect two or more records or data points relating to the same individual or group, even without direct identifiers. Combinations of quasi-identifiers or behavioural attributes can still enable records to be linked, so stripping names alone does not generally remove linkability. Whether residual linkability renders data personal data, pseudonymous, or genuinely anonymous is subject to assessment of the means reasonably likely to be used.
Are pseudonymised data and unlinkable data the same thing?
No. Pseudonymisation, as referenced in the GDPR, typically reduces but does not eliminate linkability, because records remain attributable to an individual through the use of additional information kept separately. Data can therefore be linkable while pseudonymised. Reducing linkability is one factor considered when assessing whether data has moved toward anonymisation, but pseudonymous data generally remains personal data subject to the Regulation.
How can linkability be assessed within a data set?
Assessment generally involves examining whether records can be connected as belonging to the same data subject or group, both within a single data set and across data sets. This typically considers quasi-identifiers, unique attribute combinations, and the means reasonably likely to be used to link data, including auxiliary information available to the controller or third parties. Because outcomes are context and risk dependent, such assessment is usually documented and revisited as data sets and available information change.
What techniques are commonly used to reduce linkability?
Techniques that may reduce linkability include generalisation, aggregation, noise addition, suppression of unique or rare attribute combinations, and separating datasets so that linking keys are not readily available. The appropriate combination depends on the data, the use case, and the residual risk that is acceptable following assessment. No single technique should be treated as guaranteeing that data is unlinkable in all circumstances.
How does linkability relate to a decision about whether data is anonymous?
Linkability is one of the risks typically evaluated when determining whether data can be regarded as anonymous, alongside considerations such as singling out and inference. Where records remain linkable to an individual by means reasonably likely to be used, the data will generally still be treated as personal data. The threshold and methodology for such determinations can be the subject of regulator guidance and case law, and interpretations may diverge, so the position should be verified against current official sources.
How should linkability be documented for accountability purposes?
Documentation typically records the linkability risks identified, the assumptions about available auxiliary information, the techniques applied to reduce linkability, and the residual risk accepted. This supports the accountability principle and can inform related instruments such as a Data Protection Impact Assessment where high risk is indicated. Because the available means to link data can change over time, such documentation is generally kept under review rather than treated as a one-off exercise.

Common misconceptions

Removing names and direct identifiers eliminates linkability and makes data anonymous.
Stripping direct identifiers does not by itself remove linkability. Records may still be linked through combinations of remaining attributes or by reference to other data sources. Whether the result is anonymous depends on a context-specific assessment of the means reasonably likely to be used to link and re-identify, and such data may still constitute personal data.
Pseudonymised data is not linkable and therefore falls outside the GDPR.
Pseudonymisation is a risk-reduction measure, not anonymisation. Records sharing a pseudonym remain linkable, and additional information can often restore attribution. Pseudonymised data is generally treated as personal data and remains within the scope of the GDPR.
Linkability is the same thing as identifiability.
Linkability concerns whether records can be related to each other; identifiability concerns whether data can be attributed to a specific individual. Linkability is one factor that can increase re-identification risk, but it is analytically distinct and is typically assessed together with singling out and inference.

Best practices

Assess linkability explicitly as part of any anonymisation or pseudonymisation exercise, evaluating it alongside the singling-out and inference risks rather than in isolation.
Consider both internal linkability (between records in the same dataset) and external linkability (through combination with other data sources reasonably available), rather than examining a dataset in isolation.
Treat pseudonymised data as personal data by default, applying appropriate safeguards to the pseudonym mapping and to any additional information that could restore linkage and attribution.
Document the linkability assessment, including the assumptions made about who might attempt linkage and what auxiliary information they could plausibly access, so the analysis can be revisited if circumstances change.
Re-evaluate linkability periodically, since the availability of external datasets and re-identification techniques evolves over time and can affect whether previously low-risk data becomes linkable.
Where the assessment indicates residual linkability, verify the appropriate legal basis and any additional conditions applicable to the data, and avoid asserting that data is fully anonymous without a supporting, context-specific evaluation.