Privacy, de-identification and sensitive data
When is a model trained on personal data anonymous? EDPB Opinion 28/2024 for data buyers
Quick answer
Under EDPB Opinion 28/2024, a model trained on personal data is anonymous only if, using all means reasonably likely to be used, the likelihood of extracting training subjects' personal data directly from the model, and of obtaining it through queries, is insignificant [4][3]. That is judged case by case by supervisory authorities, mostly from the controller's documentation [4]. For buyers, the evidence starts with the training data: source selection, minimisation and preparation records.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What the opinion actually decides about model anonymity
The opinion holds that whether a model trained on personal data is anonymous must be assessed case by case [4][2]. The European Data Protection Board adopted it on 17 December 2024 after the Irish supervisory authority requested it under Article 64(2) GDPR [4][2]. It answers four questions: when a model is anonymous, how a controller can show legitimate interest is an appropriate basis in development, the same question for deployment, and what follows when development used unlawfully processed data [2].
The anchor is GDPR Recital 26: information is outside the GDPR only if a person is no longer identifiable, taking into account all means reasonably likely to be used [1]. The EDPB applies that standard to model parameters rather than to a table of records. Its specific position is that personal data can be "absorbed" into weights, so the question is whether it can come back out [4][5].
As of October 2026, Opinion 28/2024 remains the EDPB's reference text on models, and the Digital Omnibus proposal to narrow the definition of personal data is not law. The opinion is guidance rather than binding law, but supervisory authorities use it in enforcement and cooperation [2].
The two-part test: direct extraction and query-based disclosure
A model passes only when both disclosure routes are insignificant, not just one [4][3]. Buyers and governance leads should treat these as separate engineering problems with separate evidence.
- Direct extraction from the model. Anyone with weights access could attempt membership inference, attribute inference, model inversion or reconstruction. This matters most for open-weight releases and for any model shared with deployers or licensees.
- Disclosure through queries. Users with API or chat access could prompt the model into regurgitating names, emails, account numbers or free-text details that appeared in training records.
Two consequences follow directly from the opinion. A model designed to output personal data about training subjects, such as a fine-tuned model that answers "what did customer X say," cannot be anonymous [4][3]. And a model that is anonymous for one release mode may not be for another, because "means reasonably likely to be used" changes when weights leave your control. The weights-release question is covered in releasing model weights trained on personal data.
Which training-data choices the EDPB expects you to show
The opinion lists design-stage elements an authority may weigh, and most of them sit upstream in the training data, not in the model [4]. That is why procurement records matter.
- Source selection. Whether the controller assessed the appropriateness of sources and limited or excluded sources with high exposure of personal data [4]. For licensed data, this is the supplier's description of origin and the categories withheld.
- Data preparation and minimisation. Whether personal data was filtered, removed, pseudonymised or replaced before training, and whether the volume collected was proportionate to the purpose [4].
- Methodological choices in training. Measures such as regularization or differential privacy that reduce memorization [4]. NIST SP 800-188 is useful context: traditional de-identification has inherent limits compared with formal privacy methods [8].
- Output-side measures. Filters and guardrails that reduce the chance of personal data appearing in responses [4].
- Testing. Resistance to attacks such as attribute and membership inference, exfiltration, regurgitation and model inversion, with testing scope proportionate to the threat [4].
- Documentation. Records an authority can review, which the opinion treats as central to its assessment [4].
Syntactic de-identification on its own is weak evidence. Sweeney's k-anonymity work shows that re-identification can still succeed on k-anonymous releases unless policies accompany the technique [9]. Free text in support tickets or legal workflows is especially hard: structured-field redaction leaves names in message bodies, signatures and quoted email threads. See PII redaction for LLM training data for measuring what redaction pipelines miss.
The documentation file an authority will ask for
If a controller claims anonymity, the opinion expects it to show its work through documentation, and authorities assess that file [4]. Elements the EDPB mentions include data protection impact assessments, the data protection officer's advice, the technical and organizational measures taken in design, evidence of theoretical resistance to re-identification techniques, and documentation given to deployers [4].
For a lab that licenses rather than collects, much of this file depends on what the supplier hands over. A model card that says "PII was removed" without a method, a recall estimate or a sample-check result will not carry an anonymity claim. Ask for the de-identification method, the fields affected, and the residual-risk statement, and keep them with the model's version history. Our de-identification evidence package checklist lists the documents to request.
The documentation also overlaps with other regimes. General-purpose model providers under AI Act Article 53(1)(d) publish a training-content summary on the Commission's 24 July 2025 template [7]; reconcile it with your anonymity file so the two descriptions of sources do not conflict.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Opinion element | Evidence to request from the data supplier | Evidence you generate | Common failure mode |
|---|---|---|---|
| Source selection | Origin description, system of record (e.g., Zendesk, Salesforce, Jira export), excluded categories | Source register linking dataset ID to model version | Source described only as "business records" |
| Preparation and minimisation | De-identification method, fields removed or replaced, tokenization scheme, sample-check results | Pre-training filter logs, dedup stats | Redaction of structured fields only; free text untouched |
| Training method | Not applicable | DP-SGD settings or regularization choices, epochs per record | Many epochs over small fine-tuning sets |
| Output measures | Not applicable | Output PII filter configuration and test set | Filter tuned for US formats only |
| Testing | Canary or known-record list, if available | Membership inference and extraction test reports | Testing on base model, not the fine-tuned release |
| Documentation | Rights and preparation record per dataset | DPIA, DPO advice, deployer documentation | Records not versioned with the model |
A preparation record to keep with each model version
The simplest durable artifact is one preparation record per licensed dataset, versioned alongside each model checkpoint that consumed it. That lets you answer an authority, a deployer or a takedown request without reconstructing history. The dataset-to-model traceability guide covers lineage mechanics.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"dataset_id": "ds-support-tickets-2026-03",
"source_description": "B2B SaaS support tickets, ticket body and agent replies",
"personal_data_categories_in_source": ["names", "emails", "phone numbers", "account IDs"],
"preparation": {
"method": "NER + regex replacement with consistent surrogates",
"fields_processed": ["subject", "body", "signature", "attachments_text"],
"sample_check": {"records_reviewed": 500, "residual_identifiers_found": 3, "notes": "two names in quoted threads"},
"residual_risk_statement": "No method is perfect; free-text residuals possible"
},
"license_ref": "license-2026-041",
"models_trained": [
{"model": "support-assist-ft", "version": "1.4.0", "epochs": 2, "dp": false}
],
"extraction_tests": ["membership-inference-2026-05.pdf", "regurgitation-probe-2026-05.csv"]
}
How anonymity interacts with legitimate interest and unlawful processing
If a model is not anonymous, the GDPR applies to it, and you need a legal basis and the rest of the accountability stack [4][6]. The opinion accepts that legitimate interest can be a basis for development and deployment, subject to the three-step test of a legitimate interest, necessity and balancing, which the companion page on documenting a legitimate interest assessment for licensed data covers [6].
Anonymity also changes the unlawful-processing analysis. Where personal data was processed unlawfully in development but the model is then properly anonymized, the opinion indicates the GDPR may not apply to the model's subsequent operation [4]. Where it is not anonymized, the unlawfulness can affect later processing, and a deployer is expected to assess whether the model was lawfully developed [4][5]. Those scenarios are detailed in unlawfully processed training data consequences.
Note the separate question of whether the licensed dataset itself is personal data for you. After the CJEU's 2025 EDPS v SRB judgment, pseudonymized data may not be personal data for a recipient who cannot re-identify it; see pseudonymised data from the recipient's perspective. That does not settle the model question, because the EDPB test looks at what can be extracted from the model.
Buyer checklist before relying on a supplier's de-identification
A buyer can strengthen an anonymity position before training starts by setting requirements in the request and the license. Use this list in data requests and diligence reviews.
- Describe the minimum personal data the use case needs; ask the supplier to exclude the rest at source.
- Require the de-identification method, affected fields and sample-check results per dataset, in writing.
- Ask how free text, attachments, call transcripts and screenshots were handled, not just structured columns.
- Ask whether any EU or UK data subjects appear in US-origin records; see EU personal data in US datasets.
- Plan extraction and memorization tests on the released model, not only the base model; see training-data extraction and memorization risk and what model memorization is.
- Agree how record-level withdrawals will be handled; see record-level takedown obligations.
The privacy and de-identification hub collects related guidance, and GDPR and selling data to AI companies covers the supplier side of the same rules.
Where SourceX fits for EU-facing model teams
SourceX sources operational datasets from US companies on request and manages licensing and ongoing purchases; it serves AI teams wherever they are based and does not train models. Every dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which is the kind of record your anonymity file needs. You can describe the data your model needs without naming suppliers.
Getting documented training data for EU-facing models
SourceX finds US businesses that hold the data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is transacted. Each release is approved by the supplying company and delivered through private, access-controlled workflows after an executed agreement; a request does not guarantee a match. Start a data request on the SourceX buyers page.
Frequently asked questions
Is an AI model personal data under the GDPR?
It can be. The EDPB says models trained on personal data cannot always be treated as anonymous, and personal data may be absorbed into parameters [4][2]. Whether a specific model is personal data depends on the extraction and query-disclosure likelihood assessed case by case.
Does removing PII from training data make the model anonymous?
Not automatically. Preparation and minimisation are evidence the EDPB weighs, but it also looks at training methods, output measures and attack testing [4]. Residual identifiers in free text can still be memorized and regurgitated.
Is Opinion 28/2024 binding?
No. It is an Article 64(2) opinion, but supervisory authorities use it to assess cases, so it shapes enforcement practice [4][2].
Sources
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
- CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
- Herbert Smith Freehills Kramer, "EDPB issues Opinion on personal data in AI models" (2025). https://www.hsfkramer.com/notes/data/2025-posts/EDPB-issues-Opinion-on-personal-data-in-AI-models
- Securiti, "Summary of EDPB Opinion 28/2024 concerning AI models processing of personal data" (2025). https://securiti.ai/summary-of-edpb-opinion-282024-concerning-ai-models-processing-of-personal-data
- Osborne Clarke, "EDPB delivers view on using personal data to train and deploy AI models" (2024). https://www.osborneclarke.com/insights/edpb-delivers-view-using-personal-data-train-and-deploy-ai-models
- Kennedys, "Training AI models: European Data Protection Board's opinion and recent developments" (2025). https://kennedyslaw.com/en/thought-leadership/article/2025/training-ai-models-european-data-protection-board-s-opinion-and-recent-developments
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- Latanya Sweeney, Data Privacy Lab, "k-Anonymity: A Model for Protecting Privacy" (2002). https://dataprivacylab.org/people/sweeney/kanonymity.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.