Matching or Combining Datasets
Matching or combining datasets is the process of comparing, linking, or merging personal data drawn from two or more sources to identify records that refer to the same person or thing. It is often used for purposes such as fraud prevention, direct marketing, or building a more complete picture of an individual. Because it brings together information from different places, it can create new insights about people that were not apparent from any single source.
Matching or combining datasets refers to processing operations that compare, link, or merge personal data obtained from multiple sources to determine or establish that records relate to the same data subject or entity (sometimes described as entity resolution or record linkage). Under the UK GDPR framework, the ICO lists 'data matching', combining, comparing or matching personal data obtained from multiple sources, within its examples of processing likely to result in a high risk to individuals; where a controller intends to carry out such processing, a Data Protection Impact Assessment (DPIA) is required under this guidance rather than merely being a discretionary consideration. Controllers should note that the ICO's high-risk list operates alongside the general criteria for identifying high-risk processing, and that the position under the EU GDPR is informed by national supervisory authority lists and EDPB guidance, which may diverge; readers should verify the applicable high-risk lists and current guidance for their jurisdiction. Independently of the DPIA question, combining datasets must have a valid Article 6 lawful basis (and, where special category data is involved, an additional Article 9 condition), and can raise purpose limitation and further-processing compatibility considerations that require separate assessment.
Why it matters
Matching or combining datasets is powerful precisely because it produces information that did not exist in any single source. Linking records that refer to the same person across separate systems can reveal patterns, correlations, and inferences that individuals never disclosed and may not anticipate, which is why the process carries distinct privacy risks beyond those of the underlying datasets. These risks include re-identification of data that seemed low-risk in isolation, use of data for purposes incompatible with why it was originally collected, and the construction of detailed profiles used for decisions affecting individuals.
For these reasons, the UK ICO includes 'data matching', described as combining, comparing or matching personal data obtained from multiple sources, within its examples of processing likely to result in a high risk to individuals. Under the ICO's guidance, where a controller intends to carry out this type of processing, a Data Protection Impact Assessment (DPIA) is required rather than merely being a discretionary or optional step. Controllers should treat this as a mandatory starting point under the UK framework and not as a factor to weigh at their own discretion.
The DPIA obligation is separate from, and additional to, the requirement to establish a valid lawful basis. Combining datasets still needs an Article 6 lawful basis, and, where special category data is involved, an additional Article 9 condition. It also frequently raises purpose limitation and further-processing compatibility questions, because data assembled for one purpose may be repurposed when linked. The position under the EU GDPR is shaped by national supervisory authority high-risk lists and EDPB guidance, which may diverge from the ICO's approach, so readers should verify the applicable high-risk lists and current guidance for their jurisdiction against the official text.
Who it's relevant to
Inside Matching or Combining Datasets
Common questions
Answers to the questions practitioners most commonly ask about Matching or Combining Datasets.