Data licensing for AI training
Fine-tuning-only data licenses: scope, limits and when they are enough
Quick answer
A fine-tuning data license lets you adapt an existing model with the licensed records, through supervised fine-tuning, preference tuning or adapter training, without the right to pre-train a base model on them. A fine-tuning-only grant is enough when five terms are written down: which base models it covers, whether adapters, merged weights and distilled models count, where the tuned model may be deployed, whether tuned weights survive the term, and what broader rights would cost later.
By SourceX Editorial · Updated
What "fine-tuning" has to mean in the grant
Define fine-tuning in the license as a list of training acts, because one post-training project uses the same records in ways a one-word grant does not clearly cover. In the InstructGPT pipeline, fine-tuning meant supervised training on human-written demonstrations, then a separate reward model trained on human rankings of outputs, then reinforcement learning against it [1]. Direct Preference Optimization (DPO) fits the model to preference pairs with no separate reward model [2].
| Training act | Why a bare "fine-tuning" grant is ambiguous | State in the license |
|---|---|---|
| Converting records to chat JSONL, splitting train and validation sets | Copies and alters records before training | Preparation acts, including tokenization and caching |
| Full-parameter supervised fine-tuning | Changes every base-model weight | Covered for the base models in a schedule |
| Adapter training (LoRA and similar) | Base weights stay frozen; output is a separate weight file | Covered; adapters are Tuned Models |
| Preference optimization on preference data | Pairs may be priced apart from SFT examples | Covered, with pair counts in the manifest |
| Reward-model training for RLHF | Produces a second model that can score other models | Covered or excluded expressly |
| Hyperparameter sweeps | One dataset yields dozens of checkpoints | Count deployed models, not checkpoints |
| Scoring models on a held-out split | Benchmarking third-party models is a different use | Name the models (evaluation-only terms) |
| Continued pre-training on raw text | Pre-training under another name | Excluded unless named |
General terms such as exclusivity and deletion are explained in SourceX's guide to AI data license terms.
Base-model scope: one named model, a family, or any model you choose
The base-model clause decides whether you can move to a better model without a new license, so settle it before the first training run.
| Pattern | How it is written | Fits | Buyer risk |
|---|---|---|---|
| Named base model | Name plus a pinned revision, commit hash or checkpoint checksum | One vertical model on one open-weight checkpoint | A newer base, or retirement of a hosted model version, needs a new grant |
| Model family | A publisher's named family and later versions | Teams tracking one publisher's releases | Disputes when successors are renamed or change modality |
| Any base model, capped | Any model the licensee picks, up to a set number of deployed Tuned Models | Teams comparing bases or serving several products | The licensor prices in the cap, so the counting rule matters |
A substitution right is the middle path: you may re-tune the same records on a different base model after notice, provided the Tuned Model on the replaced base is retired within an agreed period. It suits hosted base models, whose versions retire on the provider's schedule.
The base model's own license and acceptable-use policy still apply to the tuned model; a data license cannot widen them. Hosted training sends the records to a third party (hosted fine-tuning and API terms), and general model-definition patterns are compared in derivative and successor model rights.
Adapters, merged weights and distilled students: name every artifact
List each artifact a fine-tune can produce and say whether it is a Tuned Model, a new model or prohibited, because disputes start when one is used in a way neither side priced.
- Adapters. A LoRA-style adapter is a small set of trained weights loaded on top of an unchanged base model. Teams serving several tasks often train one adapter each, so a per-model price needs a counting rule: per deployed adapter, per use case, or unlimited adapters within one licensed purpose.
- Merged, quantized and pruned copies. Folding an adapter into the base or compressing a model produces a new checkpoint. Treat each as the same Tuned Model so serving optimizations are not license events.
- Distilled students. Distillation trains a smaller model on a tuned model's outputs, so what the teacher learned from licensed records passes to a model that never saw them. One published method fine-tunes a teacher with differential privacy, generates synthetic text from it and distills a student on that text [3]. Decide whether students are allowed, counted or excluded (distillation datasets; synthetic data from licensed data).
- Per-customer fine-tunes. If each enterprise customer receives a model tuned on the licensed records plus its own data, licensed value reaches third parties. That is sublicense territory (sublicensing on an AI platform).
A register kept per license makes the counting rule auditable.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"license_id": "LIC-0001",
"dataset_manifest": "support-histories-v3, manifest sha256 recorded",
"base_model": {"name": "example-8b-instruct", "revision": "a1b2c3d", "model_license": "publisher license v2"},
"techniques": ["LoRA SFT", "DPO"],
"artifacts": [
{"id": "triage-adapter-v4", "type": "adapter", "status": "deployed", "deployment_tier": "internal_tool"},
{"id": "triage-v4-merged-int8", "type": "merged_quantized_copy", "counts_as": "triage-adapter-v4"},
{"id": "sweep-17-ckpt-0412", "type": "training_checkpoint", "status": "deleted", "counted": false}
],
"deployed_tuned_models": 1,
"cap_under_license": 3,
"post_term_rule": "deployed models survive; no further training"
}
Deployment scope: internal tools, customer features, APIs and shipped weights
Deployment terms should follow exposure: the further a tuned model reaches beyond your staff, the more chances others have to extract licensed records, and the more a licensor will want to price it. Treat each tier as a named field of use (drafting field-of-use restrictions).
| Tier | Example | What outsiders reach | Terms to settle |
|---|---|---|---|
| Internal tool | Drafting assistant in your support desk | Nothing | Internal-use definition, contractors (internal-use-only licenses) |
| Customer feature | Ticket-routing suggestions in your SaaS product | Outputs, in a constrained interface | Commercial deployment named; output controls |
| Open API | Third parties prompt the model freely | Outputs at volume | API deployment named; monitoring, rate limits |
| Shipped weights | On-premises or on-device installs | The weights | Downstream restrictions, sublicense terms |
| Public release | Tuned weights published | The weights, with no user contract | Express permission (open-weight release) |
Extraction risk is why licensors care: Carlini et al. recovered individual training examples, including personal contact details, by querying a language model [4], and the Janus Interface study examines how fine-tuning can amplify recovery of personal information [5]. Scan inputs as well as targets for identifiers before training, and agree measurable output controls rather than a promise that no output will match a record.
Public deployment can also trigger disclosure duties the confidentiality clause must allow. As of October 2026, California's AB 2013 requires developers of generative AI systems offered to Californians to post a high-level summary of training datasets, including sources or owners, whether they contain copyrighted or licensed material and whether they contain personal information; the first deadline was 1 January 2026, and the duty recurs before each new release or substantial modification [6]. Whether fine-tuning makes you a developer under AB 2013 or a provider under the EU AI Act is covered in provider and developer duties when fine-tuning.
Tuned weights after the term: what survives deletion
Separate the training data, which you can delete or return, from tuned weights and adapters, which cannot be returned and must survive the term if you plan to keep serving them. Data copies sit in JSONL training files, tokenized caches, validation splits, experiment trackers and backups. A deletion duty covering "all derivatives" that does not exclude Tuned Models can be read to reach the adapter itself (deletion and return clauses).
Survival positions, strongest first for the buyer:
- Tuned Models trained during the term survive indefinitely, including redeployment.
- Tuned Models deployed before expiry survive; no further training after expiry.
- Tuned Models survive for a wind-down period, then are retired.
- Tuned Models are deleted with the data.
Position 2 answers a licensor's fear of quiet retraining; the sample clause below uses it. Survival binds only the licensor: a January 2024 FTC Office of Technology post notes that the agency has previously required companies that unlawfully obtained consumer data to delete products, including models and algorithms developed in whole or in part with that data [7]. Pair survival with warranties on lawful collection; more options are in what happens to trained models when a license ends.
When a fine-tuning-only grant is enough, and when to pay for more
A fine-tuning-only grant is enough when the licensed records teach a task, format or judgment on top of knowledge the base model already has; it falls short when the model must absorb a large domain corpus or the records should feed your next base model. LIMA fine-tuned a 65B-parameter LLaMA model on 1,000 curated prompt-response pairs and concluded that almost all knowledge is learned in pre-training, with limited instruction data teaching output format [8]. If your gap is knowledge rather than behavior, you may need a corpus for continued pre-training, which this grant excludes (fine-tuning vs RAG vs continued pre-training).
| Your plan | Fine-tuning-only enough? | What to add |
|---|---|---|
| LoRA or full SFT on one base model for one workflow | Yes | A substitution right |
| DPO or reward-model training on licensed pairs | Yes, if both acts are named | Reward model as a covered artifact |
| One adapter per enterprise customer | Only with customer deployment and a cap | Sublicense terms for delivered adapters |
| Distilling into a smaller student | Only if distillation is named | Students defined as Tuned Models |
| Continued pre-training on raw domain documents | No | Continued pre-training rights |
| Mixing the records into your next base model | No | Pre-training rights |
| A retrieval index over the same documents | No | A retrieval grant (RAG content license terms) |
Can licensed data be used for both fine-tuning and pre-training? Only if the grant names both; a fine-tuning-only license excludes pre-training by design.
Price the upgrade before you train. An option clause adds pre-training or wider deployment rights later on terms agreed now: an exercise window, a fixed price or formula, the records covered, and the effect on Tuned Models already built. At signing you can still walk away to other data; once a product depends on a model tuned on these records, the licensor knows your switching cost. Price structures are compared in AI data license pricing structures.
Open SFT and preference sets: the license chain behind the download
For open or marketplace SFT data, a commercial fine-tuning right depends on every upstream license and on the terms of any model that generated responses, not only on the download page's label. The Data Provenance Initiative audited more than 1,800 text datasets from widely used fine-tuning collections and reported license omission above 70% and license error rates above 50% on popular hosting sites [9]. On the Hugging Face Hub, a dataset's displayed license comes from a YAML metadata block in its README dataset card [10], which the uploader writes.
Responses generated by commercial models also carry the provider's terms; as of October 2026, Anthropic's help center, for example, says its terms do not allow using outputs to train models that compete with its own [11]. See open datasets that allow commercial fine-tuning and due diligence for synthetic fine-tuning data.
Scoping steps, owners and sample definitions
Write the use statement before discussing price, so the license names what your training actually does.
- ML lead: write the use statement. Base models with revisions, techniques, expected number of deployed Tuned Models, and deployment tiers.
- ML engineers: list artifacts and copies. This list becomes the definitions schedule.
- Counsel: draft definitions, survival and the option, using the sample below as a checklist.
- Privacy and security reviewers: check identifiers and exposure, including PII scans of inputs and targets and where weights are stored.
- Procurement: record the terms in the buyer's term sheet and keep a register per deployed model.
Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.
"Fine-Tuning Use" means copying, formatting, tokenizing and splitting the
Licensed Data and using it to (a) train Tuned Models by supervised
fine-tuning, adapter training or preference optimization, (b) train reward
models used only to train Tuned Models, and (c) evaluate Tuned Models.
Fine-Tuning Use excludes pre-training or continued pre-training of any base
model and any use of the Licensed Data in a retrieval index.
"Base Model" means a model listed in Schedule B, identified by name and
revision. Licensee may add or substitute a Base Model on written notice,
provided any Tuned Model built on a replaced Base Model is retired within
[number] days.
"Tuned Model" means a model or adapter created by Fine-Tuning Use from a
Base Model, together with merged, quantized or pruned copies of it.
Checkpoints that are never deployed do not count toward the cap in
Schedule C. Models trained on outputs of a Tuned Model are [included /
excluded].
Survival. Tuned Models deployed before expiry may continue to be used in
the channels listed in Schedule D after expiry. No Fine-Tuning Use is
permitted after expiry.
Operational records suited to domain-specific fine-tuning, such as support histories, engineering records and finance workflows, often sit inside companies rather than on dataset hubs. SourceX sources operational datasets from US companies and manages the commercial process: it checks the data and each supplier's licensing permissions, and allowed uses are agreed in a license that defines which records are included, what they can be used for, how long it runs and how delivery happens. Datasets are sourced on request, so a request does not guarantee a match (see how SourceX works with data buyers).
Fine-tuning license checklist before signature
Before signing, confirm the draft settles these nine points; each one left open invites renegotiation.
- Training acts listed, including preference optimization and reward models
- Base models named with revisions, plus a substitution right
- Counting rule based on deployed Tuned Models, with compressed copies counted once
- Distillation and training on Tuned Model outputs addressed
- Deployment tiers named, including per-customer adapters
- Deletion duty excludes Tuned Models; survival position stated
- Option terms for pre-training or wider deployment
- Confidentiality allows AB 2013 and EU training-data summaries
- Warranties on lawful collection and consents
For every other clause, see the AI data license negotiation checklist, the AI training data licensing guide and the fine-tuning datasets guide.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Know which fine-tuning uses you need licensed?
Describe the records you want to fine-tune on and the uses you need covered: base models, techniques, deployment tiers and how long tuned models must stay in service. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, and manages the license, delivery and future purchases. Submit your licensing requirements.
Sources
- Ouyang et al., OpenAI, "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
- Rafailov, Sharma, Mitchell, Ermon, Manning, Finn, "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
- arXiv:2403.00932, "Differentially Private Knowledge Distillation via Synthetic Text Generation" (2024). https://arxiv.org/html/2403.00932v2
- Carlini et al., USENIX Security 2021, "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
- arXiv:2310.15469, "The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks" (2023). https://arxiv.org/pdf/2310.15469
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
- Zhou et al., "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
- Anthropic (Claude Help Center), "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.