Skip to content

Privacy, de-identification and sensitive data

Unlawfully processed personal data in training sets: downstream consequences for model buyers and deployers

Quick answer

If personal data was processed unlawfully to train a model, that defect can follow the model downstream. Under EDPB Opinion 28/2024, a controller that deploys a third-party model is expected to have assessed, as part of its own GDPR accountability, whether the model was developed lawfully; supervisory authorities may weigh that assessment when judging the deployment [1][2]. The main exception is a model that is genuinely anonymous. Practically, buyers need documented provenance diligence, contract remedies and a remediation plan before deployment.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What Opinion 28/2024 actually decided about unlawful training

The EDPB treats unlawful development as a fact that supervisory authorities assess case by case, not as an automatic ban on every later use of the model [1]. The opinion was requested by the Irish supervisory authority and adopted in December 2024; it answers four questions, the last of which concerns how unlawful processing during development affects later processing or operation of the model [1][3]. Law-firm commentary on the opinion discusses what this question means for models trained on unlawfully processed personal data, which is where licensees and deployers are exposed [4][5].

The analysis rests on ordinary GDPR mechanics rather than a new AI rule. Lawfulness under Article 5(1)(a) and a legal basis under Article 6 attach to each processing operation, and Article 24 makes each controller responsible for demonstrating its own compliance [6]. The opinion's core move is to ask whether personal data is still "in" the model: if the model memorizes or can regurgitate training records, deployment is itself processing of that personal data, and the upstream defect is in play [1].

Two framing points matter for buyers. First, the AI Act does not displace this: Article 2(7) keeps Union data protection law fully applicable to personal data processed in connection with AI systems [7]. Second, as of October 2026 the Digital Omnibus proposals that would narrow the definition of personal data were still proposals, not law, so the Opinion 28/2024 analysis remains the operative reference [11].

The three scenarios and where a buyer sits

Your exposure depends on which scenario describes your relationship to the model and whether personal data survives in the weights. The opinion distinguishes three situations [1][2].

ScenarioWho processed unlawfullyPersonal data retained in model?EDPB position (paraphrased)Typical buyer position
1The same controller later deploysYesWhether development and deployment are separate purposes, and whether the missing legal basis taints deployment, is assessed case by caseIn-house team that trained on licensed or collected data it should not have used
2Another controller (the developer)YesThe deploying controller should have carried out an appropriate assessment that the model was not developed by unlawfully processing personal data; the depth of that assessment can vary with the risksEnterprise licensing a model, fine-tuned checkpoint or embedding model
3Any controller, but the model was then anonymisedNoIf later operation of the model does not involve personal data, the GDPR does not apply to that operation and the unlawfulness should not affect it; new personal data processed in deployment is assessed on its ownBuyer of a model whose anonymity has been demonstrated, not asserted

Scenario 3 is the escape hatch, and it is narrow. The anonymity bar in the opinion requires that the likelihood of extracting personal data from the model, directly or through queries, be insignificant, which is a high threshold under Recital 26 [1][6]. Our companion guide on when a model trained on personal data counts as anonymous covers the tests; this page assumes you cannot yet rely on them.

What "appropriate assessment" means for a deploying controller

The deployer's duty is evidentiary: you must be able to show you asked the right questions about the model's training data and acted on the answers. The opinion points supervisory authorities toward factors such as the source of the data and whether the developer's processing has already been found to infringe the GDPR by an authority or a court [1]. A model built from a personal-data breach, or from a scrape an authority has already condemned, is a red flag a deployer cannot claim not to have seen.

The opinion also notes that the depth of the assessment can scale with risk [1]. A customer-facing assistant handling health or financial queries, or a model that will be fine-tuned further on employee records, warrants more scrutiny than a narrow classifier with no generative output. Where the provider is subject to AI Act documentation duties, those artifacts are useful inputs but not a substitute for your own GDPR accountability [1][7].

Useful upstream evidence usually exists if you ask for it. General-purpose model providers must publish a training-content summary under Article 53(1)(d) using the Commission's July 2025 template [9]. For high-risk systems, Article 10 requires data governance covering data collection processes and the origin of data, which gives buyers a concrete document set to request [8]. As of October 2026, Article 10's high-risk dates have reportedly moved to December 2027 (Annex III) and August 2028 (Annex I) under Regulation (EU) 2026/1744, but the documentation it describes is already a sound diligence baseline [8].

Why provenance claims fail under scrutiny

Most training-data defects surface as broken provenance chains rather than obvious misconduct. The Data Provenance Initiative's audit of more than 1,800 text datasets found license omission above 70% and license error rates above 50% on popular hosting sites [10]. If licenses are this unreliable, the lawful-basis and consent metadata behind personal data is unlikely to be better.

Recurring failure modes buyers should test for:

  • Purpose drift. CRM, ticketing or call-recording data collected for service delivery was repurposed for model training without a compatibility assessment or a new legal basis.
  • Broken processor chains. A vendor acting as a processor under Article 28 used client data to train its own model, which can make it a controller for that processing without a legal basis [6].
  • Laundering through intermediaries. A reseller or aggregator cannot hand over rights it never received; see our guide to buying training data through a broker or reseller.
  • Ignored objections and erasure. Article 21 objections or Article 17 erasure requests honored in the source system but not propagated to training snapshots.
  • Special-category leakage. Health, biometric or union-membership data inside free-text fields, where Article 9 conditions were never met [6].
  • "Public" equals "lawful." Scraped forum posts or profile pages treated as free to use; the opinion's legitimate-interest analysis does not support that shortcut [1].

These are also the reasons many buyers prefer licensed operational data with documented rights over scraped corpora; our comparison of licensed, synthetic and scraped training data sets out the trade-offs.

Deployer diligence request for a third-party model or dataset

A written request, answered in writing and kept on file, is the core of a defensible assessment. The template below adapts the opinion's factors into questions a governance lead can send to a model developer or dataset supplier [1][8][9]. Pair it with our broader data provider due diligence questionnaire.

Illustrative example: invented to show structure; it does not describe an available dataset.

#RequestEvidence that answers itRed flag
1List personal-data sources used in pre-training, fine-tuning and RLHF, by source systemSource register with collection dates, jurisdictions, record counts by category"Publicly available web data" with no further breakdown
2Legal basis per source under GDPR Article 6 (and Article 9 condition where relevant)Legitimate-interest assessments, consent records, contractsOne blanket basis for all sources
3Any supervisory authority or court finding about the training processingSigned statement plus copies of decisions or undertakingsRefusal to answer or "not applicable" without explanation
4How objections and erasure requests are handled for data already trained onWritten procedure, suppression lists, retraining or unlearning cadenceNo mechanism after the snapshot date
5Pre-training minimization and de-identification appliedMethod description, tool versions, sample QA resultsDetection-only tooling with no residual-rate measurement
6Memorization and extraction testing on the released modelRed-team reports, canary or membership-inference resultsNo testing, or tests only on non-personal fields
7AI Act documentation (GPAI training-content summary, Article 10 governance where applicable)Published summary URL, data governance recordsSummary contradicts answers to items 1-2
8Contract remedies if a defect is foundWarranty, indemnity, notification and cooperation clausesRemedies capped at fees with no cooperation duty

For memorization testing in item 6, see training-data extraction and memorization risk; for item 5 evidence, see the de-identification evidence package checklist.

Corrective measures and what remediation looks like

Supervisory authorities have the full Article 58(2) toolkit, and the opinion lists outcomes up to erasure of the model itself [1][6]. Measures named include fines, temporary limitations on processing, erasure of the unlawfully processed part of the dataset and, where that is not possible, erasure of the whole dataset or the model, depending on the facts [1]. For a deployer, the practical consequence is operational: a model you depend on may have to be withdrawn or retrained on a timetable you do not control.

Remediation options, roughly in order of cost:

  1. Suspend the affected feature while the developer's position is clarified, and document the decision.
  2. Block regurgitation with output filtering and retrieval controls; this reduces risk but does not cure unlawful development.
  3. Demand a cleaned retrain excluding the tainted sources, with a written description of what was removed and how.
  4. Demonstrate anonymity of the retrained or existing model under the opinion's tests, which moves you toward Scenario 3 [1].
  5. Replace the model with one whose provenance you can evidence end to end.

If the defect is in a dataset you licensed directly, follow your response playbook for personal data found in a licensed dataset and check the record-level takedown obligations in your license. Weights already trained on the data are the hard part: deleting the source records does not, by itself, remove what the model learned.

Allocating the risk in contracts

Contracts cannot make unlawful processing lawful, but they decide who pays and who must cooperate. Because each controller carries its own accountability, an indemnity from the developer does not discharge your duty to assess; it only funds the response [1][6]. Draft for the response you will need, not just for damages. Teams that request licensed datasets through SourceX's buyer intake still need these clauses in every license they sign.

Clauses buyers commonly negotiate with model and data suppliers:

  • Provenance warranty that personal data in training sets was processed with a valid legal basis, tied to the source register in the diligence request.
  • Prompt notification of any supervisory inquiry, complaint or finding touching the training processing.
  • Cooperation and retraining duty, including delivery of a cleaned checkpoint or dataset version on request.
  • Suspension right for the buyer without breach if an authority's action makes continued use doubtful.
  • Audit or attestation rights over de-identification and memorization testing.

Where your own lawful basis for further fine-tuning rests on legitimate interest, document it separately; see the legitimate-interest assessment buyers must document. The privacy and de-identification hub collects the related guides.

Sourcing training data with documented rights

The cheapest way to avoid an unlawful-training problem is to start from data whose ownership, consents and preparation are documented before you train. SourceX sources operational datasets from US companies on request, rights-reviews each one for ownership and consents, removes or replaces personal details before delivery with the method recorded, and prepares diligence materials per dataset, all under a license that defines records, uses, term and delivery. Describe the data you need at sourcex.si/buyers.

Frequently asked questions

Does an unlawfully trained model become unusable for every deployer?

Not automatically. The EDPB leaves the consequences to case-by-case assessment by supervisory authorities, and a deployer that carried out an appropriate assessment is in a materially different position from one that did not [1][2].

If the developer anonymised the model afterward, are we in the clear?

For processing that involves no personal data, the opinion says the GDPR does not apply to that operation and the earlier unlawfulness should not affect it [1]. The burden is to evidence anonymity under Recital 26 standards, not to accept a vendor label [6].

Does the Digital Omnibus change this analysis?

As of October 2026, the GDPR amendments in the Digital Omnibus remained proposals that had not been adopted [11]. Plan on the current GDPR text and Opinion 28/2024.

Sources

  1. European Data Protection Board, "Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models" (2024). https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
  2. European Data Protection Board, "EDPB opinion on AI models: GDPR principles support responsible AI" (2024). https://www.edpb.europa.eu/news/edpb-opinion-on-ai-models-gdpr-principles-support-responsible-ai_en
  3. CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
  4. Osborne Clarke, "EDPB delivers view on using personal data to train and deploy AI models" (2024). https://www.osborneclarke.com/insights/edpb-delivers-view-using-personal-data-train-and-deploy-ai-models
  5. Kennedys, "Training AI models: European Data Protection Board's opinion and recent developments" (2025). https://kennedyslaw.com/en/thought-leadership/article/2025/training-ai-models-european-data-protection-board-s-opinion-and-recent-developments
  6. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  7. EUR-Lex, "Regulation (EU) 2024/1689 (AI Act), Article 2: Scope" (2024). https://eur-lex.europa.eu/eli/reg/2024/1689/art_2/oj
  8. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  9. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  10. Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  11. Acompli, "Digital Omnibus GDPR and Cookie Reforms Stall Without a Council Mandate" (2026). https://acompli.ie/news/digital-omnibus-gdpr-cookies-status-september-2026/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data