Data licensing for AI training
Licensing data for foundation-model pre-training: the rights grant you need
Quick answer
A pre-training data license must grant more than a right to "train." It should name every processing step your pipeline runs on the corpus, define "Model" so the grant covers checkpoints, successor generations and distilled or quantized variants, and let trained weights survive after the data copy is deleted. It also has to settle the output controls the licensor wants and leave room for the training-data disclosures that EU and California rules require.
By SourceX Editorial · Updated
How a pre-training grant differs from a fine-tuning-only grant
Pre-training turns a corpus into a base model that later becomes the parent of many models, so the grant has to reach across more processing, more models and more time than a grant for adapting one model. After pretraining, the base model is fine-tuned, distilled and redeployed for years, and a grant written for one fine-tuning project tends to break along that path.
| Term | What a pre-training grant needs | What a fine-tuning-only grant may limit itself to |
|---|---|---|
| Models covered | The base model family, its successors and everything derived from them | One named base model and its adapters |
| Training runs | Unlimited runs, including ablations, scaling experiments and restarts | The runs needed for the stated model |
| Processing | Filtering, deduplication, tokenizer training, mixing and upsampling with other sources | Formatting into prompt-response pairs |
| Data copy | Kept for the term, with an archival manifest after deletion | Deleted when the project ends |
| Weights after the term | Perpetual and irrevocable | Tied to the term or a named deployment |
| Deployment | Any product, API or customer channel, and open weights if negotiated | One defined application |
| Price unit | Corpus fee, or tokens counted with a named tokenizer at an agreed point | Records or examples |
For general use, exclusivity and deletion terms, see SourceX's guide to AI data license terms.
Processing rights: name what your pipeline does to the corpus
A pre-training grant should list the transformations your pipeline performs, because a right to "use the data for training" does not clearly cover copying, altering, discarding or combining it before training starts. List them in a schedule your data engineers can check:
- Copy and convert. Mirroring to object storage, extracting text from HTML or PDF, writing JSONL or Parquet, and re-extracting when the parser improves.
- Filter and classify. Language identification, quality classifiers, toxicity filters and PII redaction. These choices change the model: a controlled study that pretrained 28 models of 1.5B parameters found quality filtering raised downstream performance while removing 10% or more of the data, and toxicity filtering traded generalization for fewer toxic generations [1]. Expect to rerun quality filters for each model generation.
- Deduplicate. Within the corpus and against everything else you train on. One sentence appeared more than 60,000 times in C4, and deduplicated training made models emit memorized text about ten times less often [2].
- Tokenize and shard. Converting text to token IDs and packing it into shards, often in streaming formats such as MDS.
- Train tokenizers. A tokenizer's vocabulary is learned from corpus text and outlives any single model. One adaptation method replaces a model's least frequent tokens with tokens built from target-language text before training continues [3], so the grant should cover tokenizer training and keep vocabularies outside any deletion duty.
- Mix, upsample and repeat. Weighting the corpus against other sources and training on it for more than one epoch.
- Create derived data. Classifier scores, annotations, embeddings, held-out evaluation splits and rephrased or synthetic versions of licensed data.
Check the licensor's sample for near-duplicates against your existing corpus before agreeing price: tokens you already hold add no coverage. For a per-token fee, name the tokenizer and whether tokens are counted before or after deduplication (per-token pricing).
Defining "Model" so the grant survives the next generation
The Model definition decides whether the grant covers one checkpoint or a model family. For pre-training it should reach every model that inherits weights or learned behavior from one trained on the corpus:
- intermediate checkpoints and later training stages such as mid-training or annealing runs;
- successor models trained on a mixture that includes the corpus, whatever their name or version;
- models created by continued training from any of those checkpoints;
- fine-tuned, preference-tuned, distilled, quantized, pruned and merged models, and adapters;
- hosting, API and on-device deployment, with open-weight release addressed expressly either way.
The usual drafting patterns (named model, model family with successors, any model the licensee trains) are compared in derivative and successor model rights. Pre-training buyers generally need the broadest one, because the corpus cannot be separated from the base model once trained.
Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.
"Licensed Corpus" means the documents listed in Manifest A (document IDs and
SHA-256 hashes), as delivered and as corrected under Section 4.
"Training Use" means copying, storing, converting, filtering, classifying,
deduplicating, tokenizing and mixing the Licensed Corpus with other data, and
using it (a) to train tokenizers and Models, (b) to run experiments, ablations
and evaluations, and (c) to create Derived Data.
"Model" means any machine learning model, including all checkpoints, weights and
parameters, trained in whole or in part on the Licensed Corpus or Derived Data
by Licensee or its Affiliates, and any model created from such a model by
fine-tuning, continued training, distillation, quantization, pruning, merging
or adapter training, regardless of name or version.
"Derived Data" means tokenized shards, tokenizer vocabularies, filter scores,
annotations, embeddings and other data generated from the Licensed Corpus,
excluding Models.
Survival. Licensee's rights in Models trained before expiry or termination are
perpetual and irrevocable and survive deletion of the Licensed Corpus.
Fuller grant language is in writing the AI training rights grant. If contractors or a cloud provider run your pipeline, cover them in the affiliates, contractors and cloud processors clause.
Retention: deleting the corpus without losing the model
Separate copies of the data, which may have to be deleted or returned, from trained weights, which cannot be returned and should survive the license. In a pre-training pipeline, data copies live in raw downloads, extracted text, filtered subsets, MinHash deduplication indexes, tokenized shards, caches and backups. A deletion clause that covers "all copies and derivatives" without excluding Models can be read to reach the weights (deletion and return clauses).
Retraining without the data is the only certain way to remove its influence, and at pre-training scale that means a new training run. Offer "no further training on the corpus after termination" instead of unlearning promises (what happens to trained models when a license ends). After deletion, keep a manifest of document IDs, hashes, license IDs and the runs that used them, so you can answer audits and disclosure questions without the text.
Survival clauses bind the licensor, not regulators or third-party rights holders. The FTC's 2021 final order against photo-app developer Everalbum required it to delete the models and algorithms developed using users' photos and videos [4]. In Bartz v. Anthropic, the class settlement received final approval in July 2026 [5]. Back weight survival with warranties of title and lawful collection and an IP indemnity.
Memorization: the output controls licensors ask for
Because large models can reproduce training text verbatim, expect licensors of valuable text to ask for output-side controls; agree to controls you can measure, not a promise that no output will match the corpus. Carlini et al. extracted hundreds of verbatim training sequences from GPT-2, including personal contact details, and found larger models more vulnerable [6]. Nasr et al. recovered thousands of training examples from aligned production chat models [7]. Deduplication reduces memorized output but does not eliminate it [2].
Some deals already include such terms: the 2024 HarperCollins book program was described as an opt-in license with a commitment to limit verbatim reproduction [8]. The copyright chapter of the EU's voluntary General-Purpose AI Code of Practice covers mitigating the risk of infringing outputs and a complaint mechanism for rightsholders [9]. Workable buyer positions:
- an output filter that checks generations against an index of the licensed corpus for runs of matching tokens above a length both parties agree and you can test;
- exclusions for public-domain text, facts, short quotations and text the user supplies in the prompt;
- a notice-and-filter process for passages the licensor flags, rather than a retraining obligation;
- separate terms for retrieval or display, which is a different grant (grounding license vs training license).
Chain of title, EU opt-outs and disclosure duties
A license is evidence of permission only if the licensor holds the rights it grants, and its confidentiality terms must still let you describe the corpus in the training-data summaries EU and California law require. As of October 2026, the U.S. Copyright Office's Part 3 report remains a May 2025 pre-publication version; it concludes that copying works into training datasets may be prima facie infringing unless a defense such as fair use applies [10]. Provenance labels are weak evidence: the Data Provenance Initiative found license omission above 70% and license errors above 50% on popular dataset hosting sites [11]. Ask for document-level source and license metadata and a warranty of title; the AI data provenance guide covers the records to keep.
EU opt-outs. The text and data mining exception in Article 4 of Directive (EU) 2019/790 does not apply where rightsholders have expressly reserved their rights, for example by machine-readable means, and AI Act Article 53(1)(c) requires general-purpose AI model providers to keep a copyright policy that identifies and complies with those reservations [12]. A license from the rightsholder is what lets you train on a reserved source. Tag licensed documents so your opt-out filter does not drop them, and ask whether the grant also covers copies of the same works already in your web crawl (DSM Article 4 opt-outs).
Disclosure. Article 53(1)(d) requires a public summary of training content [12], following the template the Commission published on 24 July 2025 [13]. California's AB 2013 required developers of generative AI systems offered to Californians to post, by 1 January 2026 and before each later release or substantial modification, a high-level dataset summary covering sources or owners, whether the data includes copyrighted or licensed material, and whether it includes personal information [14]. Write confidentiality so you may name the licensor or describe the corpus at that level (confidentiality clauses vs AI transparency duties; what buyers need from suppliers for EU summaries). Other duties are in the AI training data compliance guide.
Continued pre-training on a licensed domain corpus
Continued pre-training starts from an existing base model, so the data license must permit training on top of weights you may not own, and the base model's own license may also govern the result. One study of continual pre-training for finance built its domain corpus from SEC filings and financial news [15]; sourcing such corpora is covered in domain corpora for continued pre-training. Check three points:
- the grant covers continued training of third-party base models, not only models you train from scratch;
- tokenizer extension or replacement is a permitted use [3];
- if a hosted provider runs the training, the data reaches a processor, which the hosted training and API terms guide covers.
Steps and owners: from use statement to signed grant
Work from your pipeline outward: document what you will do with the corpus, make each act and artifact a defined term, then discuss price.
- Pre-training data lead: write the use statement. Model families, training stages, modalities, deployment channels and whether open weights are possible.
- Data engineering: list pipeline acts and artifacts. This becomes the Training Use and Derived Data schedules.
- Data engineering: measure the sample. Token count with your tokenizer, near-duplicate rate, overlap with your existing corpus, language mix and date range.
- Counsel: draft definitions, survival and access. Model, Training Use, Derived Data, affiliates and processors.
- Privacy and security reviewers: check personal data and storage. De-identification method, shard and backup locations, deletion evidence.
- Counsel with the licensor: agree output controls and disclosure wording.
- Procurement: record terms in the buyer's term sheet and attach the manifest to each run's provenance record.
Records held inside operating businesses must be found before they can be licensed. SourceX sources operational datasets from US companies, such as support and sales histories, engineering records and documents, and manages the commercial process, including licensing agreements and ongoing purchases. Each dataset goes through rights review and is delivered under a license that defines the records included, permitted uses, term and delivery; datasets are sourced on request, so a request does not guarantee a match (how SourceX works with data buyers).
Drafting gaps that strand a pre-training corpus
License language written for a narrower use strands a pre-training corpus; check the draft for these gaps.
| Drafting gap | What breaks | Fix |
|---|---|---|
| Grant names one model version | The next generation needs a new license | Define Model as family, successors and derivatives |
| "Training" left undefined | Tokenizer training, ablations and synthetic rephrasing are disputed | Schedule of Training Use acts |
| Deletion covers "all derivatives" | Weights arguably fall within the deletion duty | Exclude Models; state perpetual survival |
| Access limited to licensee employees | Contractors and cloud processors running the pipeline fall outside | Affiliates and processors clause |
| Field of use "internal research" | Commercial deployment falls outside the grant | Name deployment channels |
| Absolute non-reproduction warranty | Untestable, so every output is a potential breach | Measurable filter and notice process |
The AI training data licensing hub maps other clauses; for corpus sources, see sourcing for pre-training teams and licensed text corpora for LLM pre-training.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Describe the pre-training corpus you need licensed
If the text or records you want to pre-train on sit inside US companies, describe them on the buyers page: record types, volume, history and the uses you need licensed. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, and manages the license, delivery and future purchases. Start a data request with SourceX.
Sources
- Longpre et al., NAACL 2024, "A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity" (2023). https://arxiv.org/pdf/2305.13169
- Lee, Ippolito, Nystrom, Zhang, Eck, Callison-Burch, Carlini, "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- arXiv:2311.05741, "Efficiently Adapting Pretrained Language Models To New Languages" (2023). https://arxiv.org/pdf/2311.05741
- Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
- Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
- Carlini et al., USENIX Security 2021, "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
- Nasr et al., ICLR 2025, "Scalable Extraction of Training Data from Aligned, Production Language Models" (2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/cce0e917b050208170151f77b497fc71-Abstract-Conference.html
- Authors Guild, "HarperCollins AI licensing deal" (2024). https://authorsguild.org/news/harpercollins-ai-licensing-deal/
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- arXiv:2311.08545, "Efficient Continual Pre-training for Building Domain Specific Large Language Models" (2023). https://arxiv.org/pdf/2311.08545
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.