Skip to main content
Category: Security & Breach Notification

Categories and Approximate Number of Data Subjects

Also known as: Categories of Data Subjects, Data Subject Categories, Data Subject Types
Simply put

This refers to the way an organisation groups the individuals whose personal data it processes (for example, employees, customers, or children) and an estimate of how many people fall within each group. It is typically recorded so that an organisation can describe, at a glance, who is affected by its data processing. The figure is generally an approximation rather than an exact headcount.

Formal definition

A 'category of data subjects' is a grouping or classification of identified or identifiable natural persons whose personal data is processed, defined by reference to a shared characteristic or relationship to the organisation (such as employees, customers, or children, the latter generally being individuals under 18 in the UK context). Recording the categories and an approximate number of data subjects is a data-mapping and documentation practice associated with records of processing and processor arrangements; practitioners should verify the precise obligations and any exemptions against the current text of the applicable GDPR articles, as the exact scope of what must be documented can vary by role (controller or processor) and by national implementing law. This concept applies to personal data of living natural persons and does not extend to anonymous data, and the granularity of categorisation used is subject to assessment based on the nature and complexity of the processing.

Why it matters

Recording the categories and approximate number of data subjects gives an organisation a clear picture of who is affected by its processing activities. Different groups of individuals carry different risk profiles and, in some cases, different legal considerations. For example, processing data about children (generally individuals under 18 in the UK context) typically warrants heightened care, and employee records may include special category data that requires an additional condition under Article 9 alongside a lawful basis under Article 6. Knowing which categories are in play helps an organisation apply the right protections to the right people rather than treating all processing as uniform.

The approximate number of data subjects matters because it helps convey the scale of processing and can inform risk assessments and prioritisation. A processing activity affecting a large population may merit closer scrutiny than one touching a handful of individuals. Because the figure is generally an estimate rather than an exact headcount, it should be understood as an indicator of magnitude to support decision-making, not as a precise metric. Practitioners should note that the granularity of categorisation used is subject to assessment based on the nature and complexity of the processing.

This documentation is closely tied to records of processing and to processor arrangements. Maintaining accurate, current categories supports an organisation's ability to describe its processing to regulators, to individuals, and internally. Because the precise obligations and any exemptions can vary by role (controller or processor) and by national implementing law, organisations should verify what they are required to document against the current text of the applicable GDPR articles rather than relying on a fixed template.

Who it's relevant to

Data Protection Officers and compliance leads
DPOs and compliance teams rely on categories of data subjects to maintain records of processing and to describe, clearly and consistently, who is affected by the organisation's activities. Accurate categorisation supports risk assessment, helps flag where heightened protections may be needed (such as for children or where special category data is involved), and provides a foundation for responding to regulator enquiries. They should confirm the precise documentation obligations and any exemptions against the current applicable articles.
Privacy engineers and data-mapping teams
Those responsible for building and maintaining data inventories translate processing activities into structured records, including the categories of data subjects and their approximate numbers. The granularity they choose should reflect the nature and complexity of the processing, and they should ensure the mapping covers only personal data of identifiable living individuals, excluding anonymous data.
Controllers and processors
Both roles may need to document categories of data subjects in connection with records of processing and processor arrangements, but the exact scope of what must be recorded can differ by role and by national implementing law. Each should verify its specific obligations rather than assume the same requirements apply across the board.
Lawyers advising on data protection compliance
Legal advisers use categories of data subjects to help clients assess risk, structure documentation, and consider how legal bases and additional conditions may vary across different groups. They should treat the categorisation and figures as inputs subject to assessment and confirm requirements and any exemptions against the current text of the applicable law.

Inside Categories and Approximate Number of Data Subjects

Categories of Data Subjects
A classification of the types of individuals whose personal data is processed, such as employees, customers, prospects, website visitors, suppliers' contacts, or children. This element identifies who the data relates to rather than what data is held, and typically forms part of the record of processing activities maintained under Article 30 of the GDPR.
Approximate Number of Data Subjects
An estimate of how many individuals fall within each category. The GDPR generally accepts an approximation rather than an exact count for record-keeping purposes, though the reader should verify the specific expectations against the current official text and relevant supervisory authority guidance.
Link to Personal Data Scope
This concept applies only where the individuals concerned are identified or identifiable natural persons. It generally does not extend to anonymous data, and typically not to the data of legal entities or, subject to member state derogations, deceased persons.
Contextual Role in Documentation
Categories and numbers of data subjects commonly appear in records of processing activities and can inform, but are distinct from, a Data Protection Impact Assessment under Article 35 where processing is likely to result in a high risk to individuals.
Special Category Considerations
Some categories of data subjects may involve special category data under Article 9 or data relating to children, which typically warrants closer attention because additional conditions and safeguards may apply.

Common questions

Answers to the questions practitioners most commonly ask about Categories and Approximate Number of Data Subjects.

Does the GDPR require me to state the exact number of data subjects affected?
No. The relevant provisions generally call for describing the categories and an approximate number of data subjects, not a precise headcount. The term itself refers to an approximation. In most cases an estimate or a stated range is sufficient, provided it is reasonable and supportable. You should, however, verify the specific documentation or notification obligation you are addressing against the current official text, as the level of detail expected can vary by context (for example, records of processing versus breach notification).
Is 'categories of data subjects' the same thing as 'categories of personal data'?
No, these are distinct concepts that are frequently confused. Categories of data subjects describe the types of individuals whose data is processed (for example, employees, customers, or website visitors), whereas categories of personal data describe the types of information processed about them (for example, contact details or financial data). Both are typically documented, but they answer different questions and should not be conflated. Note that special category data under Article 9 attracts additional conditions and should be identified separately where relevant.
How should we estimate the approximate number of data subjects when we do not have exact figures?
Where exact figures are not readily available, a defensible estimate or range is generally acceptable. Organizations typically derive estimates from available sources such as system record counts, account totals, or transaction volumes, and document the basis for the estimate. The aim is a reasonable, good-faith approximation rather than false precision. You should record your methodology so the figure can be explained if queried, and revisit it as processing changes.
How often should we review and update the categories and approximate numbers we have documented?
There is no single fixed interval prescribed for this in the Regulation text, so review frequency is generally a matter of risk-based judgment. In most cases organizations review these entries periodically and also when there is a material change, such as a new processing activity, a significant shift in data subject volumes, or a change in the categories of individuals involved. Treat the documentation as a living record rather than a one-time exercise, and verify any specific timing expectations against applicable guidance.
How granular should the categories of data subjects be?
The appropriate level of granularity is typically driven by clarity and usefulness rather than a prescribed format. Categories should generally be specific enough to convey meaningfully who is affected and to support risk assessment, but need not be exhaustively subdivided. Where a category includes individuals who may warrant particular care, such as children or other potentially vulnerable groups, it is often useful to identify them distinctly. The right balance is context dependent and subject to assessment.
How do the categories and approximate numbers relate to breach notification content?
When notifying a personal data breach, the categories and approximate number of data subjects concerned are generally among the details expected to be communicated, alongside the categories and approximate number of affected records. If precise information is not available at the time of notification, it can typically be provided in approximate terms and supplemented as more becomes known. The exact information required, and any phased approach, should be confirmed against the applicable notification provisions and relevant regulator guidance, which may differ in emphasis.

Common misconceptions

An exact headcount of data subjects is required.
In most cases an approximate number is generally sufficient for record-keeping purposes; the emphasis is typically on a reasonable estimate rather than a precise figure. Confirm the specific requirement against the current official text and applicable guidance.
Documenting categories of data subjects means consent is the applicable legal basis for all of them.
Identifying who the data subjects are is separate from selecting a lawful basis. Any of the Article 6 bases (consent, contract, legal obligation, vital interests, public task, or legitimate interests) may apply depending on the processing, and special category data under Article 9 needs an additional condition.
Every category of data subject falls within the GDPR.
The concept generally applies only to identified or identifiable natural persons. Anonymous data, and typically the data of legal entities or deceased persons, fall outside scope, though member state derogations can vary the position for certain cases.

Best practices

Maintain a clear list of each category of data subject alongside a reasonable approximate number, and revisit both periodically as processing activities change.
Keep the identification of data subject categories distinct from the separate exercise of selecting the appropriate Article 6 lawful basis for each processing operation.
Flag categories that may involve special category data under Article 9 or data relating to children, and assess whether additional conditions or safeguards are needed.
Confirm that each listed category concerns identified or identifiable natural persons, excluding anonymous data and, subject to applicable national law, legal entities and deceased persons.
Document how figures were estimated and verify recording expectations against the current official GDPR text and relevant supervisory authority guidance, noting that positions can diverge between regulators.
Use the categorisation to inform, but not substitute for, a Data Protection Impact Assessment under Article 35 where processing is likely to result in a high risk to individuals.