Getting started
Can data be removed from an AI model after training?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
Data generally cannot be cleanly removed from an AI model after training. Machine unlearning research offers partial methods, but the dependable options today are retraining without the data or controlling what goes in. For a company licensing records, the real safeguards are de-identification before delivery, tight scope and contract terms on deletion and future training.
Key takeaways
- Machine unlearning tries to make a model act as if certain records were never used, but current methods are partial and hard to verify.
- Deleting dataset copies is simple and certifiable; removing what a trained model learned is not.
- Most licenses separate the dataset from the model: delete the data, bar future training, keep existing models under use limits.
- The only data a model can never reveal is data it never received, so preparation carries the most weight.
- Automated personal-data and secrets scanners need human review, as their own documentation acknowledges.
What is machine unlearning?#
Machine unlearning is the set of techniques that try to make a trained AI model behave as if specific training examples had never been used, without retraining it from scratch. It is an active research field, and today the methods are partial, hard to verify and rarely offered as a contractual service.
Researchers distinguish exact unlearning, which produces the same result as retraining without the data, from approximate unlearning, which adjusts the model so it is less likely to reflect the removed examples. Exact methods usually depend on how the model was built, for example training separate parts on separate slices of data so one part can be retrained. Approximate methods are cheaper but leave open how much was actually forgotten.
For a general counsel, the question usually comes up in two places: a customer or employee asks for their data to be deleted after a license is signed, or a license ends and the company wants its records gone. Both are easier to answer when the contract and the preparation were designed with the limits of unlearning in mind.
Why is removing data from a trained model so hard?#
Removing data from a trained model is hard because what the model learned from any one record is spread across a very large number of parameters, mixed with everything else it learned. There is no index that maps a support ticket to the part of the model that holds it.
The practical picture is messier still. Models are saved as checkpoints, fine-tuned into variants, distilled into smaller models and copied into backups. Even if one version is retrained, earlier versions and derived models may remain, which is why contracts focus on what may be trained in the future rather than on what can be erased from the past.
Removal options compared#
Each removal option works on a different layer: the dataset, the model or the outputs. The table compares what each can realistically achieve.
| Option | What it does | Strength | Limits |
|---|---|---|---|
| Delete dataset copies | Removes raw and derived copies of your records | Simple to perform and certify | Does not touch models already trained |
| Retrain without the data | Builds a new model version that never saw your records | Closest to true removal | Costly for large models; older versions may persist |
| Approximate unlearning | Adjusts a model to reduce the influence of specific records | Faster than retraining | Hard to verify how much was removed |
| Output filters | Blocks outputs that reproduce specific content | Reduces visible leakage | The information may still sit in the model |
| Model retirement | Withdraws model versions trained on the data | Clear endpoint | Rarely accepted for widely used models |
| No future training | Bars use of the records in later training runs | Practical and checkable by certificate | Leaves existing models in place |
What deletion clauses can and cannot promise#
A deletion clause can reliably require the buyer to delete every copy of your dataset, including derived datasets and backups, and to certify that in writing. It cannot realistically promise that a model already trained on the data will forget it.
Most licenses therefore separate the dataset from the model. Typical terms let models trained during the license term continue, bar any further training on the records after termination, restrict outputs that reproduce records and require any successor to the models to accept the same limits. Counsel who understands this split negotiates for terms that work instead of a removal promise no one can verify.
Ask the buyer to describe where copies will live before signing: training storage, evaluation sets, checkpoints, backups and any subcontractor environments. A certificate that names each location is worth more than a general statement that all data was deleted.
Why pre-delivery controls are the real safeguard#
Pre-delivery controls are the real safeguard because the only data a model can never reveal is data it never received. For a supplier, decisions made during preparation carry more weight than any clause about removal afterward.
Tools help but do not finish the job. Gitleaks and TruffleHog are open-source scanners for secrets such as passwords, API keys and tokens (Gitleaks' maintainer announced in May 2026 that it is feature complete and will receive security patches only), and Presidio is an open-source SDK for finding and anonymizing personal data in text and images. Presidio's own documentation says there is no guarantee it will find all sensitive information and that additional systems and protections should be used, which is why human review stays in the process.
- Remove names, contact details, account numbers and other personal data from tickets, emails, chats and notes.
- Remove confidential business details such as pricing, margins and client-identifying project names where the scope calls for it.
- Exclude records whose customer contracts bar reuse, and files or code that belong to customers.
- Scan engineering records for credentials, API keys and tokens before export.
- Review samples by hand after automated detection, and keep a record of what was checked.
- Limit scope to the record types and date ranges the buyer actually needs.
Illustrative: a scheduling software company plans for deletion#
Illustrative: a fictional field service scheduling software company is preparing to license years of Zendesk tickets linked to Jira issues and GitHub code reviews. Its general counsel asks the buyer to commit to removing the company's data from any model on request. The buyer agrees to delete datasets and certify it, but says it cannot remove learned information from trained models.
Rather than stall, the company moves its protection upstream. It runs secrets scanning on code reviews and comments, strips customer names and emails from tickets, and excludes every ticket from one enterprise customer whose contract bars reuse. The license bars training on the records after the term ends and requires a deletion certificate for raw and derived datasets. The general counsel signs off because what reaches the buyer no longer contains anything the company would need to recall.
How SourceX handles scope and deletion#
SourceX puts most of the protection into the Preparation and Approval steps of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. Customer names, contact details and confidential business details are stripped out before any export, and the supplier signs off on the final release.
The SourceX Evidence Packet records the privacy record and permitted use for each package, including deletion terms, so the supplier has a written baseline to check a buyer's deletion certificate against when the license ends.
Frequently asked questions
Does a right to erasure require AI companies to unlearn data?
How erasure rights under laws such as GDPR apply to trained models is still being worked out by regulators and courts. Buyers generally respond by deleting datasets, filtering outputs and limiting future training. For a supplier, the safer course is to remove personal data before delivery so that erasure requests never reach the model.
Can a model reproduce my records word for word?
It can happen, especially with text that appears many times or is unusual, such as a distinctive template or a repeated signature block. Removing personal details, deduplicating records and excluding boilerplate reduce the chance. Contract terms that restrict outputs reproducing licensed records add a further layer.
Should I ask for an audit right to verify deletion?
A signed deletion certificate is the common baseline. An audit right gives stronger assurance but adds cost and friction, so suppliers tend to reserve it for sensitive packages or larger licenses. Whatever you choose, define derived datasets and backups so the certificate covers them.
Is synthetic data a way around the removal problem?
Not automatically. Synthetic records generated from your data can still carry its patterns and occasionally its details. If a buyer will create synthetic data from your records, treat the output as derived data in the contract, with the same use limits and deletion terms.
What should a deletion certificate include?
A useful certificate names the dataset and license, lists the locations deleted, including derived datasets, evaluation sets, backups and subcontractor copies, states the date and method of deletion, and is signed by an authorized officer. It should also confirm that no further training on the records will take place.
Sources
- Presidio is an open-source, MIT-licensed SDK for PII identification and anonymization in text and images; its documentation states there is no guarantee it will find all sensitive information and that additional systems and protections should be employed. Source
- Gitleaks is an MIT-licensed tool for detecting secrets such as passwords, API keys and tokens in git repositories, files and stdin; on May 21, 2026 its README was updated to state that it is feature complete and future releases will be security patches only. Source
- TruffleHog is an AGPL-3.0 open-source secret scanner from Truffle Security that scans sources including Git, chats, wikis, logs, object stores and filesystems. Source
Related resources
- QuestionShould companies sell or license their data?
- InsightSharing data licensing revenue with customers: a model for vertical SaaS
- QuestionIs selling company data legal?
- InsightIs it safe to license company data for AI training?
- InsightHow do I de-identify contracts and legal documents for AI training?
- IndustryBPO & contact centers data
See if your company qualifies
A short company assessment. No data uploads are needed.