Software companies
Does deleting customer data remove it from your AI training sets?
By SourceX Editorial · Updated
Short answer
Deleting customer data removes it from AI training sets only if those sets are rebuilt from source or purged by customer ID; it rarely removes the data's influence from a model already trained. Selective unlearning is not yet dependable. Rule: trace each customer to each dataset and model version, purge training data, and retrain or retire on a schedule.
Key takeaways
- Training sets can follow deletion when they are rebuilt from source and each example carries its customer of origin.
- Trained models usually cannot forget one customer on request; retraining without the data or retiring the version is the practical route.
- Derived copies, such as embeddings, feature stores, evaluation sets and backups, each need their own deletion rule.
- Promise only what your pipeline can do, in the same words in the DPA, the trust page and the offboarding checklist.
Short answer: datasets yes, models rarely#
Deleting customer data removes it from AI training sets only when the sets are rebuilt from live source records or purged by a lineage-driven job. A static snapshot copied into a training bucket keeps the data until someone deletes that snapshot as well.
A model trained on the data is a different matter. Its weights do not hold records in a form you can find and remove, and machine unlearning research has not yet produced methods most SaaS teams can rely on to erase one customer's influence. The practical options are to retrain or fine-tune again without the data, or to retire the model version.
The distinction matters because customers rarely ask about models in those words. A request to delete our data, including from AI features, covers both, so the answer has to separate what is purged now from what changes at the next model version.
Where customer data hides after deletion#
Customer data survives deletion in every copy the pipeline made along the way. Production deletion is usually well understood; the derived copies are where promises fail.
| Copy | Example | Deletion approach |
|---|---|---|
| Production database | Records in the application | Normal offboarding deletion |
| Training snapshots | Parquet or JSONL exports in a data lake | Purge by customer ID, or rebuild from source |
| Embeddings and vector indexes | Chunks of tickets or documents for retrieval | Delete vectors along with their source records |
| Feature stores | Precomputed per-customer features | Recompute, or drop by customer key |
| Evaluation and test sets | Hand-picked examples used for quality checks | Often forgotten; tag them and purge |
| Fine-tuned models | Weights trained on the data | Retrain without the data, or retire the version |
| Backups and logs | Database backups, prompt and output logs | Expire on the backup and log schedule |
What published vendor practices show#
Published vendor practices show that derived copies can be tied to deletion when the pipeline is designed for it. Notion, for example, commits to removing the vector embeddings derived from a page or workspace within 60 days of that page or workspace being deleted, which treats a derived artifact as part of the customer's data rather than the vendor's.
Backups usually lag behind. Greenhouse's security overview says database backups are kept for 30 days and a deleted record is purged on the 31st day, with remnants of a departing organization's data persisting until backups expire. A deletion promise that ignores backup windows overstates what actually happens.
How to build a training pipeline that can honor deletion#
A training pipeline can honor deletion when every example carries its origin and every dataset can be rebuilt. The work is mostly bookkeeping, and it is far cheaper to add before the first model ships than after.
- Tag every training example with tenant ID, source record ID and extraction date.
- Keep a manifest for each dataset version listing the tenants that contributed to it.
- Rebuild datasets from source on a schedule instead of appending to old snapshots.
- Register each model version with the dataset versions it was trained on.
- Run deletion jobs that purge by tenant across the lake, feature stores and vector indexes.
- Set a retraining or retirement rule for models trained on a departed customer's data.
- Log each deletion run so the customer's request can be evidenced later.
Options when a departing customer's data was used for training#
A departing customer whose data was used for training leaves the vendor with a handful of options, each with a trade-off. Choose the default in advance and write it into the DPA, so the answer is ready before anyone asks.
| Option | What it achieves | Trade-off |
|---|---|---|
| Purge training sets only | Future models exclude the data | The current model still reflects it |
| Retrain on a fixed cycle | Model influence removed at the next cycle | Compute cost and validation effort |
| Retire the model version | Immediate removal of that model | Feature quality may drop until a replacement ships |
| Train only on opted-in customers | Deletion requests become rare and planned | A smaller training set |
| Approximate unlearning methods | Reduced influence without full retraining | Hard to verify; still an active research area |
What to promise in contracts and offboarding#
Promises about AI training data in contracts and offboarding should match the pipeline exactly. A DPA can commit to deleting customer data from training datasets within the deletion period, while stating that models already trained are retrained or retired on a defined cycle rather than edited.
Some vendors sidestep the problem by never training on customer content, or by training only on content from customers who opted in. Others train only on aggregated or de-identified data, which reduces the deletion question without removing it. Individuals may also hold deletion rights under the GDPR or US state privacy statutes that reach training data; how those rights apply to trained models is unsettled and worth reviewing with counsel.
Illustrative: a help desk vendor builds a deletion-ready pipeline#
Illustrative: a fictional help desk software vendor trains a reply-suggestion model on tickets from customers who opted in. A large customer leaves and asks for full deletion, including from AI features.
The vendor's manifests show which dataset versions included the customer's tickets. A deletion job purges those tickets from the lake, the vector index and the evaluation set, and backups expire on their normal schedule. The current model was trained on the data, so the vendor confirms in writing that it will be replaced at the next scheduled retraining, exactly as its DPA describes.
When the customer's security team asks for evidence, the vendor sends the deletion log, the list of dataset versions purged and the retraining date. The request closes without escalation, because the DPA, the trust page and the written confirmation all say the same thing.
Why lineage matters when you license records too#
Lineage matters just as much when a software company licenses its own records to an AI developer. In the SourceX five-step transaction, Preparation removes personal and confidential details before delivery, and the SourceX Evidence Packet records provenance and the privacy record, so the company can show which source records went into each licensed package and answer a later deletion question with facts.
License terms can also set what the licensee must do if identifiable records were delivered and a deletion request arrives later, which is one more reason to de-identify before delivery.
Frequently asked questions
Is machine unlearning ready for production use?
For most SaaS teams, not as a dependable deletion mechanism. Research methods can reduce a data subset's influence, but proving that a specific customer's data no longer affects outputs is hard. Treat unlearning as a supplement to retraining rather than a replacement, and do not promise it to customers.
Does de-identifying data before training solve deletion requests?
It reduces them, because records that cannot be linked to a person or customer are hard to find and may fall outside some deletion rights. Customer contracts can still require deletion of a customer's content, de-identified or not, so check exactly what each contract promised.
What about the third-party model providers we call?
Check their retention and training terms for your plan. If a provider does not train on your API traffic and keeps it only briefly, deletion work focuses on your own copies. Record the provider's terms alongside your sub-processor list so the answer is documented.
Should source retention rules apply to training copies?
They should. If tickets are deleted from the help desk after a set retention period, training snapshots built from those tickets should follow the same rule, or the training copy becomes the longest-lived record you hold. Map each source retention rule to every derived copy.
Do we need to retrain every time a customer leaves?
Not necessarily. Many teams retrain on a regular cycle and exclude departed customers at the next run, and they say so in the DPA. Faster removal is needed only where a specific contract or law requires it.
How do we prove deletion to a customer?
Send a written confirmation listing the systems purged, the dataset versions affected and how the model is being handled, backed by deletion logs. A confirmation that names systems and versions is more credible than a general statement that all data was deleted.
Sources
- Notion's AI security page states that by default Notion and its AI Subprocessors do not use Customer Data to train any models, and that embeddings stored in vector databases are deleted within 60 days after a page or workspace is deleted. Source
- Greenhouse's security overview says database backups are kept for 30 days, and a deleted candidate record is completely purged on the 31st day after deletion. On request, Greenhouse will delete a departing organization's entire data tree from the production database, with remnants persisting until backups expire. Source
Related resources
- QuestionShould companies sell or license their data?
- InsightSharing data licensing revenue with customers: a model for vertical SaaS
- DataSales call transcripts
- GlossaryRevenue share
- QuestionData licensing vs data selling: what's the difference?
- InsightCan fitness and wellness businesses sell their data to AI companies?
See if your company qualifies
A short company assessment. No data uploads are needed.