Skip to content

Privacy, de-identification and sensitive data

Data subject requests and licensed training data: erasure, objection and opt-out after delivery

Quick answer

If a licensed training dataset still contains personal data for you, erasure, objection and deletion rights follow the records after delivery. Under the GDPR, a supplier that erases must notify recipients like you [1]; under the CCPA, a business must notify third parties it sold data to [6]. Your job is a routing path from supplier to buyer, a defined action on the dataset and its derivatives, and a documented position on models already trained. That position depends on whether the model is anonymous [2].

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

When do data subject rights still apply to licensed training data?

Rights apply only while the data is personal data in your hands, so the first decision is the legal status of what you received. GDPR Recital 26 puts truly anonymous information outside the regulation, judged by the means reasonably likely to be used to identify someone [1]. Pseudonymised data, where tokens replace names but a key exists somewhere, stays personal data under UK guidance [5]. The CJEU in EDPS v SRB (C-413/23 P) held that pseudonymised data can be personal for the disclosing controller but not for a recipient that cannot reasonably re-identify anyone [4].

For a buyer, that ruling cuts both ways. If the supplier holds the key, never shares it and contractually bars re-identification, you may argue the data is not personal data for you, which narrows your rights exposure. If free-text fields in support tickets, call transcripts or engineering notes leak names, addresses or account numbers, that argument fails record by record. Our guide on receiving pseudonymised data after EDPS v SRB walks through the recipient test.

In California, information that meets the 1798.140 "deidentified" definition, including a public commitment not to re-identify and contractual bars on recipients, is outside the CCPA's personal information rules [6]. As of October 2026, the proposed Digital Omnibus change to the GDPR definition of personal data is not law [9]. Plan on current rules.

Which requests reach a buyer, and through which channel?

Most requests reach a buyer indirectly, as a notice from the supplier, because the individual has no relationship with you. Map each right to its legal trigger before you write the contract clause.

  • GDPR erasure (Art. 17). Grounds include data no longer necessary, withdrawn consent, a successful objection and unlawful processing. Article 19 requires the controller to communicate erasure to each recipient unless that is impossible or involves disproportionate effort [1].
  • GDPR objection (Art. 21). Where the supplier or you rely on legitimate interests, an objection stops processing unless compelling legitimate grounds override the individual's interests [1]. Training is a common legitimate-interest use; see documenting legitimate interest for licensed training data.
  • CCPA deletion (1798.105). A business that receives a verifiable request must delete and notify its service providers and contractors, and notify third parties to whom it sold or shared the data, unless impossible or disproportionate [6]. A data buyer is usually a third party, not a service provider.
  • CCPA opt-out (1798.120). An opt-out of sale or sharing stops future sales by the supplier [6]. It does not, by itself, reach data already delivered, which is why the opt-out mechanics belong in your license.
  • Delete Act and DROP. California's Request and Opt-out Platform lets one consumer request reach every registered data broker [7]. Check whether your supplier, or your own resale activity, puts either party in that regime.

Timing matters downstream. GDPR controllers answer within one month, extendable by two further months for complexity [1], and CCPA businesses within 45 days, with one 45-day extension [6]. Your contractual window to act on a forwarded notice should sit inside the supplier's remaining time.

What should happen to the delivered dataset and its derivatives?

The dataset action is the easy part: delete or suppress the affected records everywhere they were copied, and record that you did. The hard part is knowing where "everywhere" is. A support-ticket corpus licensed in Parquet typically spawns a cleaned copy, a tokenized shard set, a deduplicated training mix, an eval split, an embedding index for retrieval and cached feature stores.

Each copy needs a join key back to the supplier's record identifier. If your pipeline drops the supplier's stable record_id during cleaning, you cannot honor a forwarded request without re-matching on content, which is slow and error-prone. Build lineage into the access controls for licensed training data so every derived table carries the source ID and license ID.

Retrieval indexes deserve special handling because they return source text verbatim. Removing a vector without removing the chunk store, or the reverse, leaves the record retrievable; the RAG privacy controls guide covers deletion at retrieval time.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "dsr_ticket_id": "DSR-2026-0412",
  "received_from": "supplier",
  "supplier_notice_date": "2026-09-02",
  "regime": "GDPR",
  "right": "erasure_art17",
  "ground": "objection_art21_upheld",
  "license_id": "LIC-0193",
  "supplier_record_ids": ["tkt_88213", "tkt_88214"],
  "match_method": "supplier_record_id",
  "copies_actioned": [
    {"store": "raw_parquet", "action": "deleted"},
    {"store": "train_mix_v7_shards", "action": "suppressed_for_next_build"},
    {"store": "rag_index_support_v3", "action": "vector_and_chunk_deleted"},
    {"store": "eval_holdout_q3", "action": "deleted"}
  ],
  "models_trained_on_records": ["assist-ft-2026-08"],
  "model_action": "filtered_output_and_retrain_next_cycle",
  "closed_date": "2026-09-19",
  "confirmation_sent_to_supplier": true
}

What happens to models already trained on the records?

Whether a trained model is itself within reach of a request depends on whether it is anonymous. EDPB Opinion 28/2024 says model anonymity is assessed case by case, considering the likelihood of extracting personal data from the model and of obtaining it through queries [2]. Commentary on the Opinion reads it to mean that a model capable of outputting training subjects' personal data is not anonymous, which keeps rights in play for the model [3]. Our explainer on when a model trained on personal data is anonymous sets out the evidence the EDPB expects.

If your memorization and extraction tests show low regurgitation risk, document that as your basis for treating the model as anonymous, and act on the data only. If they do not, choose from the options below and record why. Extraction testing methods are covered in training-data extraction and memorization risk.

Illustrative example: invented to show structure; it does not describe an available dataset.

OptionWhat it doesWhen it fitsMain weakness
Output filteringBlocks named identifiers or matched strings at inferenceLow extraction risk, request concerns a few recordsModel still encodes the data; filters miss paraphrase
Suppress and retrain at next cycleRemoves records from the next training mixRegular retraining cadence, no urgent harmOld checkpoints remain in use until replaced
Sharded retraining (SISA)Retrains only the shard that held the recordsPipeline built for it in advanceRequires sharded training from the start [8]
Approximate unlearningGradient-based removal of specific examplesResearch-stage use with evaluationHard to verify complete removal
Retire the checkpointWithdraws the model or adapterUnlawful processing found, high-risk outputHighest cost; affects downstream deployers

Machine unlearning research shows why retraining cost is a design choice: SISA training splits data into isolated shards so removing a point means retraining one constituent model, not the whole system [8]. If you fine-tune adapters on licensed data, keeping one adapter per license or per supplier gives you a similar unit to retrain. For weights already shared externally, see releasing model weights trained on licensed personal data.

How should the license route requests from supplier to buyer?

The license should turn the statutory notices into an operational duty with named contacts, a format and a deadline. Without it, the supplier's Article 19 or 1798.105 notice may arrive as an informal email that no one on your ML data team sees. The retention and deletion guide covers end-of-term deletion; the clauses below cover requests during the term.

Illustrative example: invented to show structure; it does not describe an available dataset.

Buyer checklist: data subject request clauses for a training data license

  1. Notice channel: a monitored privacy inbox or API endpoint on both sides, never a single named employee.
  2. Notice content: supplier record IDs, regime, right exercised, ground, and whether the individual also objected to model use.
  3. Buyer action window: a fixed number of business days, set inside the supplier's statutory deadline.
  4. Scope of action: raw delivery, derived copies, eval splits, embedding indexes and backups, with backup deletion at next rotation stated plainly.
  5. Model treatment: which option from the table applies by default, and who decides escalation.
  6. Confirmation: a closure record returned to the supplier with copies actioned and dates.
  7. Identifier retention: keep supplier record IDs through every pipeline stage so requests can be matched.
  8. No re-identification: buyer agrees not to re-identify pseudonymised records, which supports the EDPS v SRB recipient position [4].
  9. Volume trigger: a renegotiation or retraining threshold if requests exceed an agreed share of records.
  10. Unlawful processing: what happens if the supplier discovers the original collection lacked a basis; see downstream consequences of unlawfully processed training data.

Expect pushback on item 3 from suppliers that batch requests monthly. A reasonable compromise is suppression from new builds immediately and physical deletion on the next pipeline run.

How do de-identification and exceptions reduce the request burden?

Strong de-identification at the source is the most effective way to keep requests from reaching you at all. GDPR Article 11 says that if a controller can show it cannot identify the individual, the access-to-portability rights in Articles 15 to 20 do not apply unless the individual supplies extra identifying information [1]. That helps only if the buyer genuinely lacks the means to link records to people.

The research exception is narrower than many teams assume. Article 17(3)(d) exempts processing for scientific research under Article 89(1) only where erasure is likely to render impossible or seriously impair the research objectives [1]. A commercial fine-tuning project should not plan around it without counsel's analysis.

Before signing, ask suppliers how identifiers were removed from free text, what sample was checked and what residual risk remains. The de-identification evidence package checklist lists the documents to request, and the privacy cluster hub links the related guides. If you discover personal data that should not be there, follow the found personal data playbook.

How SourceX approaches personal data in licensed datasets

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. For an operational view of whether data is deleted once training ends, read do AI labs delete data after training, and describe your requirements to SourceX through the buyer intake page.

Scope licensed data with rights handling in mind

Tell SourceX what data you need, the regimes your program operates under and how you plan to handle requests after delivery. SourceX looks for US businesses that hold the described data, and every release is approved by the supplying company. Nothing is contracted until a supplier agrees. Start a data request at sourcex.si/buyers.

Sources

  1. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  2. European Data Protection Board, "Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models" (2024). https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
  3. Herbert Smith Freehills Kramer, "EDPB issues Opinion on personal data in AI models" (2025). https://www.hsfkramer.com/notes/data/2025-posts/EDPB-issues-Opinion-on-personal-data-in-AI-models
  4. Court of Justice of the European Union (EUR-Lex), "Judgment in Case C-413/23 P, EDPS v SRB" (2025). https://eur-lex.europa.eu/eli/C/2025/5551/oj/eng
  5. Information Commissioner's Office, "Pseudonymisation" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/pseudonymisation/
  6. California Privacy Protection Agency, "California Consumer Privacy Act of 2018 (statute text)". https://cppa.ca.gov/regulations/pdf/ccpa_statute.pdf
  7. California Privacy Protection Agency (CalPrivacy), "About DROP and the Delete Act". https://privacy.ca.gov/drop/about-drop-and-the-delete-act/
  8. Bourtoule et al., IEEE Symposium on Security and Privacy (arXiv), "Machine Unlearning" (2019). https://arxiv.org/abs/1912.03817v2
  9. Acompli, "Digital Omnibus GDPR and Cookie Reforms Stall Without a Council Mandate" (2026). https://acompli.ie/news/digital-omnibus-gdpr-cookies-status-september-2026/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data