Skip to content

Privacy, de-identification and sensitive data

Anonymised or pseudonymised? Classifying licensed training data under GDPR

Quick answer

Licensed training data is anonymous under GDPR only if no one, using all the means reasonably likely to be used, can identify the people in it; then GDPR stops applying [1]. If identifiers were replaced with tokens and a key or linkable auxiliary data exists, the dataset is pseudonymised and remains personal data for whoever can re-link it [2]. After EDPS v SRB, that question can also be asked from your position as recipient, so classification can depend on what you hold [4].

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

The legal test: Recital 26 and Article 4(5)

The dividing line is identifiability, not the technique a supplier used. Recital 26 GDPR says data protection principles do not apply to anonymous information, and it asks whether a person is identifiable taking into account all the means reasonably likely to be used, such as singling out, by the controller or by another person [1]. Article 4(5) defines pseudonymisation as processing so that data can no longer be attributed to a specific person without additional information that is kept separately and protected by technical and organisational measures [1].

Two consequences follow for a buyer. First, a supplier's label ("anonymized", "de-identified", "masked") is a description of a process, not a legal classification. Second, the US vocabulary your supplier uses does not map onto the GDPR categories; for the full cross-walk see de-identified vs anonymised legal definitions and the glossary entries on pseudonymization and data anonymization.

Why pseudonymised training data is still personal data

Pseudonymised data remains personal data wherever it can be attributed back to a person by the controller or by anyone else with access to the additional information [2][9]. The EDPB adopted Guidelines 01/2025 on pseudonymisation on 16 January 2025 and ran a public consultation that closed on 14 March 2025 [2][3]. As of October 2026, check the EDPB site for whether a final version has replaced the consultation text before quoting paragraph numbers [3].

The guidelines treat pseudonymisation as a safeguard, not an exit. It helps a controller meet data minimisation and security duties under Articles 5, 25 and 32, and it can weigh in favor of a legitimate-interest assessment or a purpose-compatibility analysis [2]. The UK ICO takes the same position for UK GDPR and flags attempts to reverse pseudonymisation as a specific risk [9].

Typical pseudonymised training records in business data look like this:

  • A CRM export where contact_name became cust_7f3a91 via a salted SHA-256 hash, and the supplier keeps the salt.
  • A support-ticket corpus where email addresses were swapped for consistent surrogates (user_0412@example.invalid) so threads still link across tickets.
  • An HR or payroll extract where employee_id was re-keyed, but job_title, site and hire_date remain intact.

Each keeps analytic value precisely because records still link to a stable individual. That linkability is what keeps them inside GDPR.

EDPS v SRB: identifiability judged from the recipient's side

The Court of Justice held in EDPS v SRB (C-413/23 P, 4 September 2025) that sufficiently strongly pseudonymised data may be personal data for the original controller but not for a recipient who cannot re-identify the individuals [4]. Bird & Bird describes it as the first CJEU judgment to confirm this recipient-relative reading explicitly [4]. Commentators have since examined how the reasoning applies to AI datasets, where the buyer is often a separate company with no access to the supplier's key [5].

For a licensed dataset, that turns classification into a factual inquiry about your own position:

  • Key custody. Does the supplier, a processor, or nobody hold the mapping table or salt? Is your access to it contractually and technically barred?
  • Auxiliary data. Could your own CRM, telemetry, scraped corpora or other licensed datasets be joined on quasi-identifiers such as dates, ZIP or postcode, employer and job title? See linkage and mosaic risk.
  • Free text. Do ticket bodies, emails or call transcripts still contain names, signatures or rare events that identify people directly? See indirect identifiers in business text.

The judgment does not make pseudonymised data non-personal by default; the original controller's obligations remain, and the recipient must actually lack reasonable means to re-identify [4]. The companion page on receiving pseudonymised data after EDPS v SRB covers transparency and contract consequences in depth.

What anonymity actually requires for training data

A dataset is anonymous only when re-identification is not reasonably likely for anyone, including a motivated outsider with public and commercially available data [1][8]. The ICO frames this as a risk assessment and recommends the motivated intruder test as a starting point [8]. The EDPB applies a similarly high bar to anonymity claims in the AI context, assessed case by case [7].

In practice, operational business records rarely clear that bar without heavy transformation. Common failure modes:

  • Consistent tokens across releases. A buyer who licenses a second extract with the same surrogate IDs can rebuild longitudinal profiles.
  • Residual quasi-identifiers. Exact timestamps, small branch locations, rare product SKUs or uncommon job titles single people out in small populations.
  • Unstructured fields. Named-entity redaction misses nicknames, signatures, quoted email chains and identifiers embedded in URLs or file paths; see PII redaction pipelines and how to measure misses.
  • Model-assisted inference. LLMs can combine weak signals in de-identified text; see LLM-assisted re-identification.

Techniques that move data toward anonymity include aggregation with minimum cell sizes, generalization of dates and locations, suppression of outliers, removal of persistent IDs, and per-release re-randomization of surrogates. Each degrades sequence, cohort and long-tail signal, so decide which property your pre-training, SFT or eval use actually needs before asking for it.

Consequences of each classification for an EU-based buyer

Classification decides whether you carry the full GDPR stack or almost none of it. The table below summarizes what typically follows.

QuestionPseudonymised (personal data for you)Anonymous (not personal data for you)
GDPR applies to your processingYes, including Art. 5 principlesNo, per Recital 26 [1]
Lawful basis neededYes; Opinion 28/2024 examines legitimate interest for development [6]No
Transparency and data subject rightsApply, subject to Art. 11 limits where you cannot identifyDo not apply
Records, DPIA, security measuresRequired where triggeredGood practice only
Model anonymity analysis laterNeeded; Opinion 28/2024 tests whether the trained model is itself anonymous [6][7]Lower risk, but document the input classification
Contract focusController roles, purpose limits, re-identification ban, onward transferWarranty on method, re-identification ban, change notice if new data arrives

EDPB Opinion 28/2024 matters even for anonymous inputs, because it treats the trained model as a separate classification question and examines the consequences of unlawful processing during development [6]. If your inputs were pseudonymised, document the legitimate-interest balancing and the safeguards that pseudonymisation provided [2][6].

US de-identification standards do not settle the GDPR question

A US supplier's compliance with a US de-identification standard is evidence for, not proof of, GDPR anonymity. HIPAA allows de-identification through Safe Harbor removal of 18 identifiers or through Expert Determination [11]. The CCPA defines "deidentified" information by reference to reasonable linkability plus three conditions on the business that holds it [12].

Neither standard asks Recital 26's question about any person using all reasonably likely means [1]. A Safe Harbor dataset that keeps year-level dates and three-digit ZIP prefixes can still be personal data for an EU buyer who holds linkable records. Treat the US certification or attestation as one input to your own assessment; the state privacy law comparison and Safe Harbor vs Expert Determination pages cover the US side.

Classification worksheet for a licensed dataset

Use a worksheet per dataset and per release, because identifiability changes when you add other data. The record below shows the fields counsel and the data team should complete together before ingestion.

Illustrative example: invented to show structure; it does not describe an available dataset.

dataset: support_tickets_2023_2025_release_2
source_systems: [Zendesk tickets, Salesforce Case, Gong call transcripts]
data_subjects: [customer contacts (incl. EU/UK residents), support agents]
direct_identifiers:
  removed: [name, email, phone, account_number, street_address]
  replaced_with: consistent surrogate per person (HMAC-SHA256)
  method_documented: yes  # supplier preparation note attached
key_custody:
  holder: supplier only
  buyer_access: contractually prohibited; not delivered
quasi_identifiers_remaining: [ticket_created_at (to minute), product_sku, agent_site, customer_company_size]
free_text_fields: [ticket_body, call_transcript]
free_text_residual_check: sample of 500 records reviewed; residual names found in signatures
buyer_auxiliary_data: [own CRM (no overlap expected), release_1 of same dataset]
cross_release_linkage: possible (same surrogates as release_1)
motivated_intruder_assessment: re-identification reasonably likely for agents via site + timestamps
classification_for_buyer: pseudonymised personal data
follow_ups: [request per-release re-keying, coarsen timestamps to day, drop agent_site, second redaction pass on signatures]

The worksheet's decisive rows are key custody, buyer auxiliary data and cross-release linkage. If any of them shows a reasonable path back to individuals, classify as pseudonymised and run the personal-data workflow.

Status notes as of October 2026

The legal baseline for this classification has not changed in 2026, though proposals are pending. The Commission's Digital Omnibus proposal would narrow the personal-data definition toward an entity-relative test, but as of September 2026 the GDPR amendments had not been adopted [10]. Until they are, rely on GDPR as written, the EDPS v SRB judgment and EDPB guidance [1][4][6]. For the UK, the ICO's anonymisation guidance is under review after the Data (Use and Access) Act 2025, so quote the live pages [8][9].

How sourcing affects the classification

Classification is easier when the supplier documents exactly what was removed and how. SourceX sources operational datasets from US companies on request, and each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked; no method is perfect, which is why the buyer's own Recital 26 assessment still matters.

Diligence materials on source, rights, preparation and allowed use are prepared per dataset, which gives counsel the inputs the worksheet above needs. Teams can describe the data they need at SourceX for buyers.

For the supplier-side view of GDPR, see GDPR and selling data to AI companies and can I license data if I am under GDPR. The privacy cluster hub and the AI data hub list related guides.

Licensing de-identified records for GDPR-sensitive training

SourceX looks for US businesses that hold the data you describe, and every release is approved by the supplying company. Nothing is contracted until a supplier agrees, and delivery runs through private, access-controlled workflows after an executed agreement. Describe the records, preparation and allowed uses you need at https://sourcex.si/buyers.

Frequently asked questions

Does hashing identifiers make training data anonymous?

No. A hash or HMAC applied consistently is a pseudonym: anyone holding the salt or key, or able to hash candidate values, can re-link records [2][9]. Unsalted hashes of low-entropy values such as phone numbers or emails can be reversed by enumeration even without a key.

If data is not personal data for us after EDPS v SRB, do we owe data subjects anything?

Your own GDPR duties depend on whether you can reasonably re-identify, but the supplier remains a controller of personal data and its transparency duties continue [4]. Document why re-identification is not reasonably likely for you, and reassess when you add new data sources.

Does an anonymous input dataset mean the trained model is anonymous?

Not automatically. Opinion 28/2024 assesses model anonymity separately, looking at whether personal data can be extracted or inferred from the model [6][7]. Anonymous inputs make that analysis easier but do not replace it.

Sources

  1. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  2. VPH Institute, "The European Data Protection Board (EDPB) adopts pseudonymisation guidelines" (2025). https://www.vph-institute.org/news/the-european-data-protection-board-edpb-adopts-pseudonymisation-guidelines.html
  3. European Data Protection Board, "Guidelines 01/2025 on Pseudonymisation (public consultation)" (2025). https://www.edpb.europa.eu/our-work-tools/documents/public-consultations/2025/guidelines-012025-pseudonymisation_lt
  4. Bird & Bird, "EU: The SRB decision - a new era for personal data and data processing agreements" (2025). https://www.twobirds.com/en/insights/2025/eu-the-srb-decision-a-new-era-for-personal-data-and-data-processing-agreements
  5. Lewis Silkin, "It's nothing personal - Reassessing pseudonymised data and AI after EDPS v SRB" (2025). https://www.lewissilkin.com/insights/2025/12/18/its-nothing-personal-reassessing-pseudonymised-data-and-ai-after-edps-v-srb-102ly5u
  6. CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
  7. Herbert Smith Freehills Kramer, "EDPB issues Opinion on personal data in AI models" (2025). https://www.hsfkramer.com/notes/data/2025-posts/EDPB-issues-Opinion-on-personal-data-in-AI-models
  8. Information Commissioner's Office, "How do we ensure anonymisation is effective?" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/
  9. Information Commissioner's Office, "Pseudonymisation" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/pseudonymisation/
  10. Acompli, "Digital Omnibus GDPR and Cookie Reforms Stall Without a Council Mandate" (2026). https://acompli.ie/news/digital-omnibus-gdpr-cookies-status-september-2026/
  11. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  12. California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data