Skip to content

Data licensing for AI training

Deletion and return clauses for licensed training data

Quick answer

A workable deletion clause in an AI training data license names every copy that must go (delivered files, processed shards, tokenized caches, embeddings, annotation exports and derived datasets), says who holds them (you, affiliates, contractors, cloud processors), sets a realistic window, treats backups as aging out on a fixed schedule, and ends with a signed certificate. It should state plainly whether trained weights are inside or outside scope, because silence there is where most end-of-term disputes start. For the wider set of terms, start at the AI training data licensing guide.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What the deletion obligation should cover

The obligation should cover every materialization of the licensed records, not just "the Data" as delivered. By the end of a project, one supplier export typically exists in five or six forms, and a clause that only names the delivered files leaves most of them in place.

Map the copies before you negotiate. A typical pipeline produces:

  • Delivered originals: the Parquet, JSONL or CSV files in a landing bucket, plus any staging copies.
  • Processed training shards: WebDataset-style tar shards (for example train-{000000..004095}.tar) that group files by basename into samples, which are new files with new names [5].
  • Tokenized caches: pre-tokenized arrays, packed sequences and dataloader caches on training nodes or shared file systems.
  • Embeddings and indexes: vectors in a vector store, which must be deleted by key from the index; deleting the source file does not remove them [4].
  • Annotation and eval artifacts: labeling-tool exports, preference pairs, gold sets and red-team prompts built from licensed records.
  • Derived datasets: filtered subsets, synthetic data seeded from licensed records, and summaries that reproduce record content.

Rights to keep or delete derived datasets connect to the rights grant itself; see synthetic data from licensed data and writing the AI training rights grant so the definitions line up.

Where trained weights sit

Trained weights should be addressed by name, either carved out of deletion or deliberately brought in, because ambiguity here is the most expensive gap in the clause. Most buyers need the carve-out: a model trained during the term continues to exist after the data is gone, and a clause that defines "Derived Data" as anything "generated from or using" the records can be read to reach checkpoints.

Some consumer platform terms show the carve-out in practice: analysts note, for example, that deleting a Luma AI account does not revoke the training license for content already used [9]. Suppliers of operational data may push back, so be ready to trade something (a field-level restriction, a no-memorization test, or limits on successor models) for a clean weights carve-out. The separate question of what happens to models when the license ends is covered in model retention after license termination and derivative and successor model rights.

A contractual carve-out does not bind a regulator. The FTC's Everalbum order required deletion of models and algorithms developed using improperly obtained photos [2], and FTC staff have warned that AI companies must keep their privacy and confidentiality commitments about data use [3]. In the EU, the EDPB's Opinion 28/2024 addresses when a model can be treated as anonymous and what follows from unlawful processing during development [6]. The carve-out protects you against the licensor; only clean provenance protects you against a disgorgement remedy. See algorithmic disgorgement.

Backups should be handled by an age-out rule rather than a promise to scrub them, because restoring and rewriting immutable snapshots is rarely feasible. The standard pattern is: the licensee need not delete licensed data from backups that cannot be selectively edited, provided those backups expire under a documented schedule, are not restored for any use other than disaster recovery, and remain under the license's confidentiality terms until they expire.

Write three things into the clause:

  1. The maximum retention period for backups, tied to your actual snapshot policy (object-lock periods, database point-in-time recovery windows, tape rotation). Do not accept a number your infrastructure team has not confirmed.
  2. A restoration rule: if a backup is restored, licensed records in it are deleted again within the main deletion window.
  3. A legal-hold carve-out: copies preserved under a litigation hold or regulatory obligation are retained only for that purpose and deleted when the hold lifts.

Residual copies on decommissioned hardware fall under media sanitization practice. NIST SP 800-88 Rev. 2 groups sanitization methods into clear, purge and destroy [1]; naming it as the reference standard gives both parties an objective test instead of a debate over what "securely deleted" means.

Contractors, affiliates and cloud processors

The clause should make the licensee responsible for deletion by everyone it gave access to, and require it to obtain written confirmation from them. Labeling vendors, eval contractors, offshore annotation teams and managed training providers often hold full copies, and a licensor will hold you to their deletion.

Keep an access register from day one: entity, purpose, data subset, storage location, date access granted. Without it you cannot certify deletion at term end. Who may receive data in the first place is a separate negotiation, covered in affiliates, contractors and cloud processors in access clauses.

Return or destroy, window and certificate

Choose "destroy and certify" as the default and reserve "return" for cases where the licensor actually needs the files back, because returning terabytes of data the licensor already holds adds cost without adding protection. Return makes sense for unique physical media or when the licensor lacks a master copy.

Set the window by copy type rather than one number for everything. Hot storage and vector indexes can usually be purged quickly; shards spread across training clusters and contractor environments take longer; backups follow the age-out rule. Give the licensee a short extension with notice if a contractor confirmation is late; how such windows are tracked alongside other supplier commitments is covered in data supplier SLAs.

The certificate should be signed by an officer or the accountable data owner, not by whoever ran the delete command. Certificate mechanics and evidence formats belong to the delivery workflow; the contract only needs to name the required fields. The supplier-side view of retention and deletion terms is covered in the SourceX guide to retention and deletion in data licensing, and the term itself is defined in the retention policy glossary entry.

Aligning deletion with privacy obligations

Where licensed data includes personal information, statutory deletion duties sit alongside the contract and can be stricter. A license deletion clause is a floor, and the clause should say that applicable law prevails where it requires earlier or broader deletion.

Three regimes come up repeatedly with operational data:

  • HIPAA limited data sets remain protected health information and may be used only under a data use agreement [8]; the DUA's restrictions govern what you may keep alongside the commercial license.
  • Biometric data under Illinois BIPA: Section 15 requires a published retention schedule and destruction guidelines for entities in possession of biometric identifiers [7], so voiceprints or face-geometry scans derived from recordings carry their own destruction clock.
  • GDPR and the EDPB opinion: if a model may not be anonymous, erasure questions can reach beyond the dataset [6].

De-identified data reduces but does not remove this overlap; see de-identified vs anonymized legal definitions and the de-identification evidence package.

Clause scoping table and sample language

The table and sample clause below show one way to scope each copy type; adapt the windows to your own infrastructure.

Illustrative example: invented to show structure; it does not describe an available dataset.

Copy typeTypical locationIn deletion scope?Suggested handling
Delivered filesLanding bucket, stagingYesDelete within main window
Training shards, tokenized cachesObject storage, node-local NVMe, shared FSYesDelete; sanitize decommissioned drives per SP 800-88 [1]
Embeddings and vector indexesVector storeYesDelete vectors by key; rebuild or compact the index where the store requires it [4]
Annotation exports, eval setsLabeling tool, eval repoYes, unless licensed separatelyDelete or carry over by written election
Synthetic or filtered derived datasetsData lakePer rights grantDefined list; default delete
BackupsSnapshots, PITR, tapeAge-outExpire on schedule; no restore except DR
Legal-hold copiesHold repositoryCarved outRetain for hold only
Trained weights and checkpointsModel registryExpressly excluded (or included)State it either way

Illustrative example: invented to show structure; it does not describe an available dataset.

Deletion on expiry or termination. Within 30 days after expiry or termination, Licensee shall delete all Licensed Data and Derived Data Copies in its possession or control, including copies held by Affiliates and Service Providers, and sanitize any media on which they were stored and which is leaving Licensee's control using a method consistent with NIST SP 800-88. "Derived Data Copies" means any copy, shard, cache, embedding, index, annotation, subset or synthetic record that contains or reproduces Licensed Data, but excludes Model Weights. Copies in backup systems that cannot reasonably be selectively deleted need not be deleted if they expire within 90 days, are not restored except for disaster recovery, and remain subject to Section [Confidentiality]. Within 15 days after the deletion period, an officer of Licensee shall deliver a certificate identifying the systems and Service Providers covered, the method used, and any copies retained under a legal hold.

Certificate fields to require: license ID; data release identifiers; systems and buckets purged; contractor confirmations attached; sanitization method; backup expiry date; retained copies and legal basis; signer name and title; date.

How SourceX approaches deletion in licensed deals

SourceX sources operational datasets from US companies on request and manages the commercial process, including the license. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, allowed uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Buyers can describe the data they need.

Request licensed training data with clear end-of-term scope

SourceX looks for US businesses that hold the operational data you describe, and nothing is contracted until a supplier agrees and approves the release. Licenses define records, uses, term and delivery for each deal. Tell SourceX what data you need.

Frequently asked questions

Should "Derived Data" ever include trained weights?

Only if you intend to retrain or retire the model when the license ends. If you plan to keep the model, exclude "Model Weights" by definition and negotiate any compensating restriction separately; see the negotiation checklist.

Do we have to delete eval sets built from licensed data?

Yes, unless the license grants a surviving right to keep them. Eval sets are a common oversight because they live in code repositories rather than data platforms; list them in the access register.

Is a deletion certificate enough evidence?

For most commercial licensors, yes, if it names systems, contractors and methods. Where the licensor has audit rights, keep the underlying logs (bucket lifecycle events, vector delete calls, sanitization records) for the audit period.

Sources

  1. National Institute of Standards and Technology (NIST), "SP 800-88 Rev. 2, Guidelines for Media Sanitization" (2025). https://csrc.nist.gov/pubs/sp/800/88/r2/final
  2. Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
  3. Federal Trade Commission (Office of Technology), "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  4. Amazon Web Services (Amazon S3 User Guide), "Deleting vectors from a vector index". https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-vectors-delete.html
  5. WebDataset project, "webdataset (GitHub repository)". https://github.com/webdataset/webdataset
  6. CMS (law firm summary), "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
  7. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  8. eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  9. ConductAtlas, "Luma AI Terms of Service: account deletion does not revoke AI training license". https://conductatlas.com/platform/luma-ai/luma-ai-terms-of-service/account-deletion-does-not-revoke-ai-training-license/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data