Skip to main content
Category: Data Classification & Identifiers

Derived Data

Also known as: derived information, computed data
Simply put

Derived data is new information that is created by processing, transforming, or combining existing datasets rather than being collected directly. For example, calculating a value (such as a file size or a score) from other stored data produces derived data. Because it is generated from other data, it must be kept consistent with the base data it depends on.

Formal definition

Derived data refers to data values that are generated computationally from one or more existing (base) datasets through extraction, transformation, cross-referencing, or aggregation, rather than being captured at source. In data engineering contexts, it encompasses properties automatically calculated during processing (e.g., a computed size property set on save) as well as new information produced by combining raw datasets. A recognized operational risk is that derived data can diverge from or introduce errors relative to its base data if the derivation logic or synchronization is flawed. Note that this evidence base addresses derived data as a general data-engineering concept; it does not establish how a given item of derived data is classified under data protection law. Whether specific derived data constitutes personal data, and any resulting legal-basis, special-category, or transfer considerations, requires separate assessment against the applicable framework and is out of scope for this definition.

Why it matters

Derived data matters because it expands the amount of information an organization holds beyond what it directly collected. New values produced by processing, transforming, or combining existing datasets can carry meaning and sensitivity that the original base data did not obviously present. For privacy practitioners, this means that a data inventory built only around directly collected fields may understate what an organization actually processes, since derived values are generated computationally rather than captured at source.

A further concern is reliability. As the evidence digest notes, derived data can diverge from or introduce errors relative to its base data if the derivation logic or synchronization is flawed. Where derived values inform decisions about individuals, faulty derivation logic can propagate inaccuracies. This intersects with data-quality obligations under many privacy frameworks, though the specific applicability depends on whether the derived data in question qualifies as personal data and on the framework at issue.

Importantly, the evidence base here treats derived data as a general data-engineering concept and does not establish how any particular item is classified under data protection law. Whether specific derived data constitutes personal data, and any resulting legal-basis, special-category, or cross-border transfer considerations, requires separate assessment against the applicable framework. Practitioners should therefore treat the presence of derived data as a trigger for analysis rather than assume a fixed legal characterization.

Who it's relevant to

Data Protection Officers and Privacy Leads
DPOs and privacy leads should account for derived data when mapping processing activities, since values generated from base datasets may not appear in inventories focused only on directly collected fields. Whether any given derived value is personal data, and how it should be treated, requires separate assessment against the applicable framework; the presence of derived data is a prompt for that analysis rather than a settled classification.
Data Engineers and Architects
Engineers and architects design the extraction, transformation, cross-referencing, and aggregation processes that create derived data, including properties automatically calculated during operations such as a document save. They are responsible for ensuring derivation logic is sound and that derived values stay synchronized with their base data, since flaws in either can introduce errors relative to the source.
Compliance and Data Quality Teams
Compliance and data-quality teams have an interest in the reliability of derived data, given the recognized risk that it can diverge from base data if derivation or synchronization is flawed. Where such data may inform decisions about individuals, the accuracy of derived values can be relevant to data-quality expectations, subject to assessment of whether and how a given framework applies.

Inside Derived Data

Source personal data
The original personal data from which new information is generated. Derived data typically originates from data provided by or observed about a data subject, and the underlying source data usually remains personal data subject to the GDPR.
Derivation logic or processing operation
The calculation, inference, aggregation, or algorithmic process applied to source data to create new data points, such as a credit score, a risk category, or a behavioural profile. This processing is itself an activity requiring a lawful basis under Article 6.
Resulting derived data point
The output produced from the derivation, for example an inferred preference, a segment label, or a predicted attribute. Where this output relates to an identified or identifiable individual, it generally constitutes personal data in its own right.
Potential special category character
Derived data can, through inference, reveal special category information (for example health, ethnicity, or beliefs) even where the source data was not itself sensitive. In such cases an additional Article 9 condition is generally required, subject to assessment of whether the inference falls within that category.
Link to identifiability
Whether derived data is in scope depends on whether it relates to an identifiable person. Genuinely anonymous derived outputs generally fall outside the GDPR, but the threshold for anonymisation is high and should be assessed on the facts.

Common questions

Answers to the questions practitioners most commonly ask about Derived Data.

Is derived data exempt from the GDPR because the organisation created it rather than collecting it directly?
No. The fact that data is generated internally rather than collected from the individual does not remove it from scope. If derived data relates to an identified or identifiable natural person, it is generally personal data under the GDPR regardless of how it came into existence. The origin of the data affects transparency and lawful basis considerations, not whether the Regulation applies. Only genuinely anonymous data falls outside scope, and derived data is frequently still linkable to an individual.
Is derived data the same thing as anonymised or aggregated data?
Not necessarily, and the two should not be conflated. Derived data is produced by processing or inference from existing data, but it can still be attributable to an individual and therefore remain personal data. Aggregation or anonymisation is a separate question that turns on whether re-identification is reasonably possible. Derived data may be anonymous in some cases and identifiable in others; each output should be assessed on its own facts rather than assumed to be anonymous because it was calculated or inferred.
What lawful basis under Article 6 typically applies when generating derived data?
There is no single answer; the appropriate Article 6 basis depends on the purpose and context of the derivation. Organisations should identify a basis for the generating activity as a distinct processing operation, considering whether it is compatible with the purpose for which the source data was originally obtained. Legitimate interests, contract, or consent may each be relevant in different scenarios, subject to assessment. Where the derived output concerns special category data under Article 9, an additional condition must also be satisfied.
How should derived data be reflected in a record of processing and in transparency notices?
Because derived data is often not obvious to the individual, transparency is a particular focus. Where personal data is inferred or generated rather than obtained from the data subject, information obligations generally still apply, and privacy notices should describe that such processing occurs and its purposes. Records of processing activities should typically capture derivation as a processing operation, including its purpose and legal basis. The precise disclosures required should be verified against the current text of the applicable transparency provisions.
Do data subject rights such as access, rectification, and erasure extend to derived data?
In general, where derived data is personal data, data subject rights apply to it as they would to other personal data, subject to the usual conditions and exemptions. Access requests may extend to inferred information about the individual, and rectification and erasure may be relevant depending on the circumstances. Practical handling of inferences and profiling can be complex, and organisations should assess each right against the specific derived output and any applicable limitations rather than assuming a blanket position.
Should a Data Protection Impact Assessment be conducted before generating derived data?
It depends on the nature and risk of the derivation. A DPIA under Article 35 is generally required where processing is likely to result in a high risk to individuals, and certain derivation activities, such as extensive profiling or inference of sensitive characteristics, may meet that threshold. Organisations should assess whether the generating activity triggers the DPIA requirement and consult applicable regulator guidance and lists, which can vary between member states, to determine when an assessment is expected.

Common misconceptions

Derived data is not personal data because the organisation created it rather than collecting it from the individual.
The origin of the data does not remove it from scope. Where derived data relates to an identified or identifiable individual, it is generally treated as personal data regardless of whether it was inferred or generated internally.
Because it is created from non-sensitive inputs, derived data can never be special category data.
Inferences can reveal special category information even from ordinary inputs. Where a derived attribute reveals, for example, health or ethnicity, an additional Article 9 condition may be needed. Whether a particular inference constitutes special category data can be uncertain and is subject to assessment and evolving regulatory guidance.
Aggregating or transforming data automatically anonymises it, taking derived outputs outside the GDPR.
Transformation does not guarantee anonymity. Unless the output genuinely cannot be linked back to an individual, it typically remains personal data. The anonymisation threshold is high and should be evaluated against the risk of re-identification in context.

Best practices

Map where derived data is generated in your processing operations and record what source data feeds each derivation, so the data flows and their outputs are documented.
Identify and document a lawful basis under Article 6 for the derivation activity itself, rather than assuming the basis for the source data automatically covers it, and note that consent is only one of several available bases.
Assess whether any derived output could reveal special category data and, where it does, identify an appropriate additional Article 9 condition before relying on it.
Where you treat derived outputs as anonymous and out of scope, assess and document the risk of re-identification rather than assuming transformation is sufficient, recognising the anonymisation threshold is high.
Include derived data within data subject rights processes, since rights such as access, rectification, and erasure can generally extend to inferred or generated data that relates to the individual, subject to applicable limitations.
Consider a Data Protection Impact Assessment under Article 35 where derivation involves profiling or high-risk processing, and monitor evolving regulator guidance on inferred data, verifying your position against the current official text.