Skip to content

Data licensing for AI training

What happens to trained models when a data license ends

Quick answer

Data license termination does not settle, by itself, whether a trained model can stay in use: the contract does. If the license is silent, or requires deletion of "all copies and derivatives," the licensor can argue the weights must go. Before signing, separate data copies, which you delete and certify at the end, from models trained during the term, which should survive under a no-further-training covenant. Then agree how breach, a supplier's rights failure or a regulator's order changes that outcome.

By SourceX Editorial · Updated

Data copies and trained weights are different assets at termination

Termination normally ends your right to hold and process the licensed data; whether it also ends your right to use a model depends on how the license classifies weights. Map both sides of that line before you read the termination clause, because each side needs different contract language.

Data copies are what a deletion duty can reach and what you can certify as gone (drafting that duty is covered in deletion and return clauses for licensed training data):

  • delivered files (for example Parquet or JSONL in your bucket) and staging copies
  • normalized, filtered and de-identified versions
  • tokenized shards, deduplication indexes and data-loader caches
  • held-out evaluation splits and annotation exports
  • embeddings and vector indexes built from the records, covered in removing licensed content from vector indexes
  • backups and snapshots

Model artifacts are results the license should treat separately: base checkpoints and final weights, LoRA (low-rank adaptation) and other adapter weights, reward models, distilled, quantized and merged variants, and the outputs and synthetic data those models produce. Which of these count as the licensed "Model" belongs in derivative and successor model rights.

Deals can settle the line in writing: a published practitioner case study of an imagery-for-training license describes a survivability clause covering what happens to the trained model when the license ends [1]. Without one, the line is contested, because "return or destroy" language written for databases and documents does not say whether weights are a copy. Licensors also have a technical argument: researchers extracted hundreds of verbatim training sequences from GPT-2 [2] and thousands of training examples from aligned production models [3]. Expect licensors to reject any assurance that weights "contain no licensed data"; concede measurable output controls and keep the weights.

Five model outcomes a license can specify

A license can let models trained during the term survive indefinitely, survive while the data stops being used, survive for a run-off period, be replaced by retrained versions, or be "unlearned." The first three can be verified from records, retraining is costly but provable, and unlearning has no agreed acceptance test.

OutcomeWhat the clause saysCost to the buyerEvidence either side can check
Perpetual survivalModels trained before the cut-off date may be used, modified and deployed indefinitely within the permitted fieldNone at termination; the licensor may price it inTraining-run logs dating each model's training before the cut-off
Survival plus no further trainingAs above, and no training, fine-tuning or evaluation on the data after the cut-offLow: the next model generation uses other dataRun logs, manifests, deletion certificate
Run-off (sunset) periodExisting models stay in production for a fixed period, then are retired or replacedA replacement on the licensor's timetableModel registry and deployment dates
Retrain obligationModels are replaced by versions trained without the data by a set dateA full run for pre-training; far less for an adapter or small fine-tuneLineage records for the replacement model
Unlearning obligationThe data's influence is removed from existing weightsOpen-ended: no agreed method or testNone agreed

The "cut-off date" is the effective date of expiry or termination. For most buyers the strongest structure pairs perpetual rights in models trained before the cut-off with a finite term for holding the dataset itself: the licensor decides how long you keep its data, and you decide how long your model lives. A fixed-term license with no survival language does the opposite and can strand a model mid-deployment. Choose the outcome deliberately, then set the license term with term length and perpetual rights in mind.

How the cause of termination should change the outcome

Natural expiry and termination for convenience should leave trained models in place; only your uncured breach of the use restrictions justifies a run-off or retirement. A supplier's rights failure or a regulator's order can override any survival clause, which is why survival needs warranties and an indemnity behind it.

CauseReasonable model outcomeBuyer position
Expiry at end of termSurvival plus no further trainingInsist: this is the case survival clauses exist for
Buyer terminates for convenienceSurvival plus no further trainingAccept a termination fee rather than a model loss
Licensor terminates for convenienceSurvival, plus a refund or credit for the unused termResist the right itself in a training license
Buyer's uncured material breach of use restrictionsRun-off period, then retire the affected modelsLimit to use-restriction breaches, add a cure period, exclude payment disputes
Supplier lacked rights or consentsSet by the third party's rights, not by your contractBack survival with data warranties and an IP indemnity
Record withdrawn during the termRemove from data copies; earlier models unchanged unless the license says otherwiseHandle under record-level takedown terms
Regulator or court orderWhatever the order requiresNo clause helps; diligence on collection does

The FTC has required model deletion where it alleged that data was collected or used unlawfully. Its 2021 final order against photo-app developer Everalbum required it to delete the models and algorithms developed using users' photos and videos [4]. The proposed Rite Aid order described in December 2023 would require deletion of photos and videos from its facial recognition system and any data, models or algorithms derived from them [5]. A January 2024 FTC staff post says the agency has required such deletions for models built in whole or in part from unlawfully obtained data [6].

These statements reflect the Commission's leadership at the time, and current priorities may differ. The contract point holds either way: a survival clause binds only your licensor.

Personal data adds a separate test in the EU. The EDPB's Opinion 28/2024, adopted in December 2024, says whether a model trained on personal data is anonymous must be assessed case by case, and it is anonymous only if the likelihood of extracting personal data from it, directly or through queries, is insignificant [7]. See EDPB Opinion 28/2024 for data buyers before relying on survival for a model trained on personal data.

Why unlearning is a weak contractual remedy

Retraining without the data is the only removal method a buyer can readily demonstrate to a licensor. A promise to "unlearn" a dataset from existing weights creates an obligation with no agreed method, no acceptance test and no way for either side to prove completion.

Exact removal means retraining without the records, either from scratch or from a checkpoint saved before they were first used. Approximate unlearning adjusts existing weights to suppress what the model learned from target records. Research on it is active, but there is no widely accepted test showing that a model no longer reflects a given dataset, and the extraction results above show that training content can persist [2][3].

Access controls do not change this. Delivering data through a live share, such as the Delta Sharing protocol, where the data provider runs the sharing server and manages recipient access [8], lets a licensor end access at the cut-off; it does nothing to weights trained earlier, and rows already exported into training shards are copies again.

Offer these instead:

  • No further training after the cut-off date, evidenced by run logs and the deletion certificate. A periodic audit limited to logged training inputs, as in the published case study [1], lets the licensor check compliance without inspecting your model.
  • Output tests: a filter against verbatim reproduction of licensed records, checked with canaries, membership inference or extraction tests from model leakage audits.
  • Removal at the next scheduled model version rather than by a fixed deadline, when the licensor needs eventual removal.
  • Priced removal on request. The same case study describes a tiered removal-on-request mechanism with costs tied to the removal method [1], which prices the licensor's request instead of promising an outcome nobody can verify.
  • Architectural isolation for short terms. If a fine-tuning set trains only an adapter that stays unmerged from the base model, deleting the adapter removes its direct contribution, provided no outputs, synthetic data or distilled students from it were reused. In pre-training, saving a checkpoint before licensed data enters the mixture shortens any later retrain.

The trade-off depends on the application. A retrain obligation can be tolerable for supervised fine-tuning (SFT) on a narrow dataset. It is rarely acceptable for a pre-training corpus inside a base model that many products depend on.

What a survival clause must name, with sample language

A survival clause should define protected models by the cut-off date, state the rights that survive, exclude models and compliance records from deletion, add a no-further-training covenant and spell out the breach exception. Leave any of these out and a broad deletion duty can swallow the model.

Survival also has to reach copies you cannot recall: models customers host or fine-tune, weights you sublicense to platform customers or release as open weights, and outputs already delivered, whose ownership is covered in who owns model outputs.

Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.

14.1 "Cut-off Date" means the effective date of expiry or termination.
     "Trained Model" means any model, checkpoint, adapter or reward model
     whose parameters were updated using Licensed Data before the Cut-off
     Date, and any distilled, quantized, merged or further fine-tuned
     variant of such a model, whenever made.

14.2 Survival. Except under 14.5, Licensee's rights to use, host, modify,
     deploy and commercialize Trained Models within the Permitted Field
     survive expiry or termination for any reason. These rights are
     perpetual, irrevocable and fully paid-up, and bind Licensor's
     successors and assigns.

14.3 No further training. After the Cut-off Date, Licensee will not use
     Licensed Data to train, fine-tune or evaluate any model.

14.4 Deletion. Within [X] days after the Cut-off Date, Licensee will delete
     all copies of Licensed Data, including tokenized, deduplicated, cached
     and embedded forms, and certify deletion in writing. Backup copies are
     deleted on normal rotation and never restored. Trained Models and
     Retained Records are excluded. "Retained Records" means manifests,
     file hashes, record counts, license and run identifiers, evaluation
     results and documentation required by law.

14.5 Breach. If Licensor terminates for Licensee's uncured material breach
     of Section [Use Restrictions], Licensee may use Trained Models deployed
     before the Cut-off Date for a run-off period of [X] months, will not
     deploy them in new products, and will then retire or replace them.

14.6 Outputs. Outputs generated in accordance with this Agreement remain
     usable after the Cut-off Date, subject to Section [Output Restrictions].

Records to keep after the licensed data is deleted

Deleting the data must not delete your ability to show what trained the model, because transparency duties continue for as long as the model is offered. As of October 2026, three regimes illustrate why:

  • EU AI Act Article 53(1)(d) requires providers of general-purpose AI models to publish a sufficiently detailed summary of training content using the AI Office template [9].
  • California AB 2013 requires developers of generative AI systems made available to Californians to post training-data documentation, including the sources or owners of the datasets, whether they include copyrighted material, whether they were purchased or licensed, and whether they include personal information. It was due by 1 January 2026 and is due again before each later release or substantial modification [10].
  • Colorado SB26-189, signed 14 May 2026, requires developers of automated decision-making technology that materially influences consequential decisions to give deployers documentation including training data categories, starting 1 January 2027 [11].

The confidentiality clause must permit those disclosures after termination (see confidentiality clauses versus transparency duties), and the deletion clause must carve out the records behind them. A per-dataset ledger entry linked from your model registry is enough; tracing which models trained on which dataset covers the tooling, and exiting a data contract covers certificates and handover.

Illustrative example: invented to show structure; it does not describe an available dataset.

dataset_ref: DS-0142
license_ref: DLA-0142 (amendment 2)
cut_off_date: 2027-06-30
delivery_manifest: manifest-DS-0142.json    # file list, byte sizes, SHA-256 per file
records_delivered: <count from acceptance report>
first_training_use: 2026-11-04
last_training_use: 2027-05-18
training_runs: [sft-run-0412, sft-run-0430]
trained_models: [support-agent-v3-adapter, support-agent-v3.1-adapter]   # survive under 14.2
derived_artifacts: [synthetic-qa-batch-07]
copies_deleted: [raw bucket prefix, parquet staging, tokenized shards, eval split, vector index]
deletion_certificate: cert-DS-0142-2027-07-14.pdf
disclosures_citing_dataset: [AB 2013 documentation page, EU training content summary]

Opening asks, responses and fallbacks for counsel

Answer each licensor opening position on trained models with a response that protects the model, and keep a fallback that still gives the licensor something it can verify.

Licensor's opening askBuyer responseFallbackWalk-away signal
"Delete all copies and derivatives"Exclude Trained Models and Retained Records by definitionAdd no further training plus a deletion certificateWeights stay inside the deletion duty
"Models end with the license"Perpetual survival for models trained before the cut-offRun-off period matched to your release cycleRetirement deadline shorter than one training run
"Unlearn on request"No further training plus output testsRemoval at the next scheduled retrainUnlearning with acceptance left to the licensor
"Any breach ends model rights"Only uncured material breach of use restrictionsRun-off instead of immediate shutdownPayment disputes can retire models
"Survival for internal research only"Survival matches the production field of useSurvival for models deployed before the cut-offProducts already shipped fall outside

When SourceX sources a dataset, the license defines which records are included, what they can be used for, how long the license runs and how delivery happens, and nothing is delivered until an agreement is executed and the supplier approves the terms. State the model outcome you need when you describe your data and licensing requirements. For the rest of the agreement, work from the AI data license negotiation checklist and the AI training data licensing hub. SourceX's owner pages explain AI data license terms, how long buyers keep licensed data and revocation from the data owner's side.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Know what your models need after the license ends?

Describe the data you need and the uses that must outlast the license term, including models trained during it. SourceX looks for US companies that hold that data, checks the data and the supplier's licensing permissions, and manages the license, in which pricing and allowed uses are agreed; nothing is contracted until a supplier agrees. Submit your licensing requirements.

Sources

  1. terms.law, "AI and data licensing (case study: archive imagery licensed for AI training)". https://terms.law/case-studies/ai-data-licensing-archive-imagery.html
  2. Carlini et al., USENIX Security, "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
  3. Nasr et al., ICLR, "Scalable Extraction of Training Data from Aligned, Production Language Models" (2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/cce0e917b050208170151f77b497fc71-Abstract-Conference.html
  4. Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
  5. Federal Trade Commission, "Coming face to face with Rite Aid's allegedly unfair use of facial recognition technology" (2023). https://www.ftc.gov/business-guidance/blog/2023/12/coming-face-face-rite-aids-allegedly-unfair-use-facial-recognition-technology
  6. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  7. CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024 opinion). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
  8. Databricks, "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing" (2021). https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
  9. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  10. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  11. Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data