Skip to main content
Category: Data Classification & Identifiers

Reasonably Likely Means of Identification

Also known as: Means reasonably likely to be used, Reasonably likely means
Simply put

This is a test used to decide whether a person can be identified from data, and therefore whether that data counts as personal data. Instead of asking if identification is theoretically possible, it asks whether identification is realistically likely using methods that could reasonably be used. If the only ways to identify someone are far-fetched or remote, the data may fall outside this threshold, though this always depends on the specific circumstances.

Formal definition

A concept used when assessing identifiability of a natural person, drawn from Recital 26 of the GDPR, which states that account should be taken of all the means reasonably likely to be used to identify an individual, whether directly or indirectly, including techniques such as singling out. It sets a risk-based rather than an absolute standard: rather than considering every conceivable method, the assessment weighs the means that could realistically be deployed, taking into account factors such as available technology, cost, and effort. In practice, the assessment is contextual and forward-looking, and ICO guidance treats it as central to determining whether data has been effectively anonymised or remains personal data (including pseudonymised data that can still be attributed to an individual). Note that Recital 26 is interpretive rather than an operative article, that the threshold of 'reasonably likely' is qualitative and subject to assessment on the facts, and that regulator guidance and case law on where the boundary lies continue to evolve; practitioners should verify the current position against official sources.

Why it matters

The 'reasonably likely means of identification' test sits at the heart of one of the most consequential decisions in data protection: whether a given dataset is personal data at all. If identification is reasonably likely, the data falls within the scope of the GDPR and the full range of obligations, from lawful basis to data subject rights, applies. If the only routes to identification are far-fetched or remote, the data may fall outside that threshold and, where genuinely anonymised, sit outside the Regulation. Because the entire compliance perimeter can turn on this assessment, getting it wrong in either direction carries real risk: treating personal data as anonymous can leave individuals unprotected and organisations exposed, while over-classifying anonymous data can impose unnecessary obligations.

The test also matters because it is deliberately risk-based rather than absolute. Recital 26 of the GDPR directs assessors to consider all the means reasonably likely to be used, including techniques such as singling out, taking into account factors like available technology, cost, and effort. This means the analysis is not a one-off checkbox but a contextual and forward-looking judgement that can shift as technology and re-identification techniques evolve. What was reasonably difficult to identify at one point may become straightforward later, so an assessment made today may not hold indefinitely.

Because the standard is qualitative and applied on the facts, there is recognised uncertainty about precisely where the boundary lies. Recital 26 is interpretive rather than an operative article, and regulator guidance and case law continue to develop. Practitioners should therefore treat any identifiability determination as provisional, document their reasoning, and revisit it against the current official position rather than assuming a past conclusion remains valid.

Who it's relevant to

Data Protection Officers and Compliance Leads
DPOs and compliance leads rely on this test to determine whether a dataset falls within the scope of data protection obligations at all. It shapes decisions about whether anonymisation claims are defensible and whether pseudonymised data must still be treated as personal data. Because the assessment is contextual and forward-looking, these roles should document the reasoning and revisit conclusions as circumstances and available technology change.
Engineers and Data Architects
Those designing datasets, analytics pipelines, and anonymisation or pseudonymisation processes need to understand that theoretical possibility of identification is not the standard; the question is whether identification is reasonably likely using means that could realistically be deployed, including singling out. This affects how data is transformed, aggregated, and released, and requires attention to whether evolving techniques could later make individuals identifiable.
Privacy Lawyers and Advisors
Legal advisors assessing whether the GDPR applies to a given processing activity depend on this qualitative threshold to frame their analysis. Given that Recital 26 is interpretive rather than operative, and that regulator guidance and case law on where the boundary lies continue to evolve, advisors should qualify their conclusions and verify the current position against official sources rather than relying on a fixed interpretation.
Organisations Sharing or Publishing Data
Bodies releasing datasets, whether for research, open data, or to third parties, use this test to judge whether the data they share remains personal data or has been effectively anonymised. Because identification risk can shift with new technology and depends heavily on context, these organisations should treat any anonymisation determination as provisional and subject to reassessment over time.

Inside Reasonably Likely Means of Identification

Means Reasonably Likely to Be Used Test
A standard, drawn from Recital 26 of the GDPR, for determining whether a natural person is identifiable. It requires accounting for all the means reasonably likely to be used to identify the individual, either directly or indirectly, by the controller or by another person.
Objective Factors Assessment
Recital 26 indicates that ascertaining what is reasonably likely should take account of objective factors, such as the costs of and the amount of time required for identification, considering available technology at the time of processing and technological developments.
Direct and Indirect Identification
The concept covers identification not only through direct identifiers but also indirectly, for example by singling out an individual or by combining a dataset with other information reasonably available.
Boundary With Anonymous Data
Where identification is not reasonably likely by any means, data may be regarded as anonymous and generally falls outside the material scope of the GDPR. The test therefore marks the line between personal data and anonymous information, and this boundary is subject to interpretation and evolving guidance and case law.
Contextual and Dynamic Nature
Whether means are reasonably likely to be used depends on context and can change over time, since technological developments and newly available data can alter the outcome of the assessment.

Common questions

Answers to the questions practitioners most commonly ask about Reasonably Likely Means of Identification.

Does data become anonymous simply because we have removed direct identifiers like names?
Generally, no. Removing direct identifiers such as names does not by itself render data anonymous. The relevant question is whether an individual remains identifiable by any means reasonably likely to be used, which can include indirect identifiers, combinations of attributes, or linkage with other available datasets. Where re-identification remains reasonably likely, the data is typically still personal data and remains within the scope of data protection law. This is a context-dependent assessment rather than a fixed technical checklist.
If only a third party, and not us, could re-identify the individuals, does that mean the data is not personal data in our hands?
Not necessarily. The assessment of reasonably likely means of identification is not limited strictly to the means available to you alone; it can take account of means reasonably likely to be used by others, including third parties, subject to the analysis of what is realistically likely rather than merely theoretically possible. Regulatory guidance and case law have addressed how far means available to other parties should be considered, and the position can turn on factors such as the realistic prospect of linkage and legal or practical barriers to obtaining the additional information. This area involves recognized nuance, and the reader should assess the specific facts and consult current guidance.
How do we practically assess whether a means of identification is reasonably likely to be used?
In most cases the assessment considers factors such as the cost of and time required for identification, the technology available at the time of processing and expected technological developments, and the realistic likelihood that such means would actually be used. It is generally treated as an objective, context-dependent test rather than a purely theoretical one. Documenting the factors considered, including relevant technical and organizational barriers, helps support the conclusion, though the analysis should be revisited as circumstances change.
Should we revisit an identifiability assessment over time?
Typically, yes. Because the analysis depends on available technology, cost, and the datasets that could be combined, a conclusion reached at one point may not hold indefinitely. Developments in technology or the emergence of new linkable data sources can make identification more likely over time. It is generally advisable to treat identifiability as subject to periodic review rather than a one-time determination, and to record the basis and date of each assessment.
What kinds of factors weigh in favor of treating data as still identifiable?
Factors that generally weigh toward continued identifiability include the presence of rich or unique indirect identifiers, the availability of auxiliary datasets that could be linked, low cost or effort required to re-identify, and weak technical or organizational barriers to combining information. Conversely, robust barriers, significant cost or effort, and the absence of realistically available linking data may weigh against a finding of reasonably likely identification. These factors are assessed together rather than in isolation, and the outcome remains fact-specific.
How should the outcome of this assessment be documented?
It is generally good practice to record the datasets considered, the means of identification assessed, the technical and organizational safeguards relied upon, and the reasoning supporting the conclusion, along with the date of the assessment. Clear documentation supports accountability and helps demonstrate why data was treated as personal data or, where applicable, as falling outside that scope. Because regulators may take differing views on borderline cases, the documentation should reflect the specific context and note where the conclusion depends on assumptions that may need to be revisited.

Common misconceptions

If a dataset does not contain names or obvious identifiers, it is automatically anonymous and outside the GDPR.
Removing direct identifiers does not by itself make data anonymous. The assessment must consider all means reasonably likely to be used, including indirect identification through singling out or combination with other reasonably available data, so pseudonymised or de-identified data may still be personal data.
The test only concerns means available to the controller holding the data.
Recital 26 refers to means reasonably likely to be used by the controller or by another person. The analysis can extend beyond the controller, though how far to account for third-party means is a matter of interpretation and remains subject to guidance and case law.
Once data is assessed as anonymous, that conclusion is permanent.
The assessment is dynamic and should account for available technology at the time of processing and for technological developments. Data considered non-identifiable at one point may become identifiable as technology and available datasets evolve, so the position should be kept under review.

Best practices

Document a reasoned identifiability assessment for each dataset, referencing the objective factors noted in Recital 26 such as cost, time required, and available technology.
Assess both direct and indirect identification risks, including the potential to single out an individual or to re-identify by combining the data with other reasonably available information.
Consider means reasonably likely to be used not only by your organisation but, where relevant, by other persons, and record the reasoning where you conclude such means are or are not likely.
Treat anonymisation conclusions as time-bound and periodically re-evaluate them in light of technological developments and newly available datasets.
Where data is only pseudonymised rather than anonymised, continue to treat it as personal data and apply appropriate safeguards, avoiding reliance on the removal of direct identifiers alone.
Verify your approach against the current official text of the GDPR and applicable regulatory guidance, and note that interpretation of this boundary can vary and continues to evolve.