Skip to content

Data licensing for AI training

Fine-tuning-only data licenses: scope, limits and when they are enough

Quick answer

A fine-tuning data license lets you adapt an existing model with the licensed records, through supervised fine-tuning, preference tuning or adapter training, without the right to pre-train a base model on them. A fine-tuning-only grant is enough when five terms are written down: which base models it covers, whether adapters, merged weights and distilled models count, where the tuned model may be deployed, whether tuned weights survive the term, and what broader rights would cost later.

By SourceX Editorial · Updated

What "fine-tuning" has to mean in the grant

Define fine-tuning in the license as a list of training acts, because one post-training project uses the same records in ways a one-word grant does not clearly cover. In the InstructGPT pipeline, fine-tuning meant supervised training on human-written demonstrations, then a separate reward model trained on human rankings of outputs, then reinforcement learning against it [1]. Direct Preference Optimization (DPO) fits the model to preference pairs with no separate reward model [2].

Training actWhy a bare "fine-tuning" grant is ambiguousState in the license
Converting records to chat JSONL, splitting train and validation setsCopies and alters records before trainingPreparation acts, including tokenization and caching
Full-parameter supervised fine-tuningChanges every base-model weightCovered for the base models in a schedule
Adapter training (LoRA and similar)Base weights stay frozen; output is a separate weight fileCovered; adapters are Tuned Models
Preference optimization on preference dataPairs may be priced apart from SFT examplesCovered, with pair counts in the manifest
Reward-model training for RLHFProduces a second model that can score other modelsCovered or excluded expressly
Hyperparameter sweepsOne dataset yields dozens of checkpointsCount deployed models, not checkpoints
Scoring models on a held-out splitBenchmarking third-party models is a different useName the models (evaluation-only terms)
Continued pre-training on raw textPre-training under another nameExcluded unless named

General terms such as exclusivity and deletion are explained in SourceX's guide to AI data license terms.

Base-model scope: one named model, a family, or any model you choose

The base-model clause decides whether you can move to a better model without a new license, so settle it before the first training run.

PatternHow it is writtenFitsBuyer risk
Named base modelName plus a pinned revision, commit hash or checkpoint checksumOne vertical model on one open-weight checkpointA newer base, or retirement of a hosted model version, needs a new grant
Model familyA publisher's named family and later versionsTeams tracking one publisher's releasesDisputes when successors are renamed or change modality
Any base model, cappedAny model the licensee picks, up to a set number of deployed Tuned ModelsTeams comparing bases or serving several productsThe licensor prices in the cap, so the counting rule matters

A substitution right is the middle path: you may re-tune the same records on a different base model after notice, provided the Tuned Model on the replaced base is retired within an agreed period. It suits hosted base models, whose versions retire on the provider's schedule.

The base model's own license and acceptable-use policy still apply to the tuned model; a data license cannot widen them. Hosted training sends the records to a third party (hosted fine-tuning and API terms), and general model-definition patterns are compared in derivative and successor model rights.

Adapters, merged weights and distilled students: name every artifact

List each artifact a fine-tune can produce and say whether it is a Tuned Model, a new model or prohibited, because disputes start when one is used in a way neither side priced.

  • Adapters. A LoRA-style adapter is a small set of trained weights loaded on top of an unchanged base model. Teams serving several tasks often train one adapter each, so a per-model price needs a counting rule: per deployed adapter, per use case, or unlimited adapters within one licensed purpose.
  • Merged, quantized and pruned copies. Folding an adapter into the base or compressing a model produces a new checkpoint. Treat each as the same Tuned Model so serving optimizations are not license events.
  • Distilled students. Distillation trains a smaller model on a tuned model's outputs, so what the teacher learned from licensed records passes to a model that never saw them. One published method fine-tunes a teacher with differential privacy, generates synthetic text from it and distills a student on that text [3]. Decide whether students are allowed, counted or excluded (distillation datasets; synthetic data from licensed data).
  • Per-customer fine-tunes. If each enterprise customer receives a model tuned on the licensed records plus its own data, licensed value reaches third parties. That is sublicense territory (sublicensing on an AI platform).

A register kept per license makes the counting rule auditable.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "license_id": "LIC-0001",
  "dataset_manifest": "support-histories-v3, manifest sha256 recorded",
  "base_model": {"name": "example-8b-instruct", "revision": "a1b2c3d", "model_license": "publisher license v2"},
  "techniques": ["LoRA SFT", "DPO"],
  "artifacts": [
    {"id": "triage-adapter-v4", "type": "adapter", "status": "deployed", "deployment_tier": "internal_tool"},
    {"id": "triage-v4-merged-int8", "type": "merged_quantized_copy", "counts_as": "triage-adapter-v4"},
    {"id": "sweep-17-ckpt-0412", "type": "training_checkpoint", "status": "deleted", "counted": false}
  ],
  "deployed_tuned_models": 1,
  "cap_under_license": 3,
  "post_term_rule": "deployed models survive; no further training"
}

Deployment scope: internal tools, customer features, APIs and shipped weights

Deployment terms should follow exposure: the further a tuned model reaches beyond your staff, the more chances others have to extract licensed records, and the more a licensor will want to price it. Treat each tier as a named field of use (drafting field-of-use restrictions).

TierExampleWhat outsiders reachTerms to settle
Internal toolDrafting assistant in your support deskNothingInternal-use definition, contractors (internal-use-only licenses)
Customer featureTicket-routing suggestions in your SaaS productOutputs, in a constrained interfaceCommercial deployment named; output controls
Open APIThird parties prompt the model freelyOutputs at volumeAPI deployment named; monitoring, rate limits
Shipped weightsOn-premises or on-device installsThe weightsDownstream restrictions, sublicense terms
Public releaseTuned weights publishedThe weights, with no user contractExpress permission (open-weight release)

Extraction risk is why licensors care: Carlini et al. recovered individual training examples, including personal contact details, by querying a language model [4], and the Janus Interface study examines how fine-tuning can amplify recovery of personal information [5]. Scan inputs as well as targets for identifiers before training, and agree measurable output controls rather than a promise that no output will match a record.

Public deployment can also trigger disclosure duties the confidentiality clause must allow. As of October 2026, California's AB 2013 requires developers of generative AI systems offered to Californians to post a high-level summary of training datasets, including sources or owners, whether they contain copyrighted or licensed material and whether they contain personal information; the first deadline was 1 January 2026, and the duty recurs before each new release or substantial modification [6]. Whether fine-tuning makes you a developer under AB 2013 or a provider under the EU AI Act is covered in provider and developer duties when fine-tuning.

Tuned weights after the term: what survives deletion

Separate the training data, which you can delete or return, from tuned weights and adapters, which cannot be returned and must survive the term if you plan to keep serving them. Data copies sit in JSONL training files, tokenized caches, validation splits, experiment trackers and backups. A deletion duty covering "all derivatives" that does not exclude Tuned Models can be read to reach the adapter itself (deletion and return clauses).

Survival positions, strongest first for the buyer:

  1. Tuned Models trained during the term survive indefinitely, including redeployment.
  2. Tuned Models deployed before expiry survive; no further training after expiry.
  3. Tuned Models survive for a wind-down period, then are retired.
  4. Tuned Models are deleted with the data.

Position 2 answers a licensor's fear of quiet retraining; the sample clause below uses it. Survival binds only the licensor: a January 2024 FTC Office of Technology post notes that the agency has previously required companies that unlawfully obtained consumer data to delete products, including models and algorithms developed in whole or in part with that data [7]. Pair survival with warranties on lawful collection; more options are in what happens to trained models when a license ends.

When a fine-tuning-only grant is enough, and when to pay for more

A fine-tuning-only grant is enough when the licensed records teach a task, format or judgment on top of knowledge the base model already has; it falls short when the model must absorb a large domain corpus or the records should feed your next base model. LIMA fine-tuned a 65B-parameter LLaMA model on 1,000 curated prompt-response pairs and concluded that almost all knowledge is learned in pre-training, with limited instruction data teaching output format [8]. If your gap is knowledge rather than behavior, you may need a corpus for continued pre-training, which this grant excludes (fine-tuning vs RAG vs continued pre-training).

Your planFine-tuning-only enough?What to add
LoRA or full SFT on one base model for one workflowYesA substitution right
DPO or reward-model training on licensed pairsYes, if both acts are namedReward model as a covered artifact
One adapter per enterprise customerOnly with customer deployment and a capSublicense terms for delivered adapters
Distilling into a smaller studentOnly if distillation is namedStudents defined as Tuned Models
Continued pre-training on raw domain documentsNoContinued pre-training rights
Mixing the records into your next base modelNoPre-training rights
A retrieval index over the same documentsNoA retrieval grant (RAG content license terms)

Can licensed data be used for both fine-tuning and pre-training? Only if the grant names both; a fine-tuning-only license excludes pre-training by design.

Price the upgrade before you train. An option clause adds pre-training or wider deployment rights later on terms agreed now: an exercise window, a fixed price or formula, the records covered, and the effect on Tuned Models already built. At signing you can still walk away to other data; once a product depends on a model tuned on these records, the licensor knows your switching cost. Price structures are compared in AI data license pricing structures.

Open SFT and preference sets: the license chain behind the download

For open or marketplace SFT data, a commercial fine-tuning right depends on every upstream license and on the terms of any model that generated responses, not only on the download page's label. The Data Provenance Initiative audited more than 1,800 text datasets from widely used fine-tuning collections and reported license omission above 70% and license error rates above 50% on popular hosting sites [9]. On the Hugging Face Hub, a dataset's displayed license comes from a YAML metadata block in its README dataset card [10], which the uploader writes.

Responses generated by commercial models also carry the provider's terms; as of October 2026, Anthropic's help center, for example, says its terms do not allow using outputs to train models that compete with its own [11]. See open datasets that allow commercial fine-tuning and due diligence for synthetic fine-tuning data.

Scoping steps, owners and sample definitions

Write the use statement before discussing price, so the license names what your training actually does.

  1. ML lead: write the use statement. Base models with revisions, techniques, expected number of deployed Tuned Models, and deployment tiers.
  2. ML engineers: list artifacts and copies. This list becomes the definitions schedule.
  3. Counsel: draft definitions, survival and the option, using the sample below as a checklist.
  4. Privacy and security reviewers: check identifiers and exposure, including PII scans of inputs and targets and where weights are stored.
  5. Procurement: record the terms in the buyer's term sheet and keep a register per deployed model.

Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.

"Fine-Tuning Use" means copying, formatting, tokenizing and splitting the
Licensed Data and using it to (a) train Tuned Models by supervised
fine-tuning, adapter training or preference optimization, (b) train reward
models used only to train Tuned Models, and (c) evaluate Tuned Models.
Fine-Tuning Use excludes pre-training or continued pre-training of any base
model and any use of the Licensed Data in a retrieval index.

"Base Model" means a model listed in Schedule B, identified by name and
revision. Licensee may add or substitute a Base Model on written notice,
provided any Tuned Model built on a replaced Base Model is retired within
[number] days.

"Tuned Model" means a model or adapter created by Fine-Tuning Use from a
Base Model, together with merged, quantized or pruned copies of it.
Checkpoints that are never deployed do not count toward the cap in
Schedule C. Models trained on outputs of a Tuned Model are [included /
excluded].

Survival. Tuned Models deployed before expiry may continue to be used in
the channels listed in Schedule D after expiry. No Fine-Tuning Use is
permitted after expiry.

Operational records suited to domain-specific fine-tuning, such as support histories, engineering records and finance workflows, often sit inside companies rather than on dataset hubs. SourceX sources operational datasets from US companies and manages the commercial process: it checks the data and each supplier's licensing permissions, and allowed uses are agreed in a license that defines which records are included, what they can be used for, how long it runs and how delivery happens. Datasets are sourced on request, so a request does not guarantee a match (see how SourceX works with data buyers).

Fine-tuning license checklist before signature

Before signing, confirm the draft settles these nine points; each one left open invites renegotiation.

  • Training acts listed, including preference optimization and reward models
  • Base models named with revisions, plus a substitution right
  • Counting rule based on deployed Tuned Models, with compressed copies counted once
  • Distillation and training on Tuned Model outputs addressed
  • Deployment tiers named, including per-customer adapters
  • Deletion duty excludes Tuned Models; survival position stated
  • Option terms for pre-training or wider deployment
  • Confidentiality allows AB 2013 and EU training-data summaries
  • Warranties on lawful collection and consents

For every other clause, see the AI data license negotiation checklist, the AI training data licensing guide and the fine-tuning datasets guide.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Know which fine-tuning uses you need licensed?

Describe the records you want to fine-tune on and the uses you need covered: base models, techniques, deployment tiers and how long tuned models must stay in service. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, and manages the license, delivery and future purchases. Submit your licensing requirements.

Sources

  1. Ouyang et al., OpenAI, "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  2. Rafailov, Sharma, Mitchell, Ermon, Manning, Finn, "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  3. arXiv:2403.00932, "Differentially Private Knowledge Distillation via Synthetic Text Generation" (2024). https://arxiv.org/html/2403.00932v2
  4. Carlini et al., USENIX Security 2021, "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
  5. arXiv:2310.15469, "The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks" (2023). https://arxiv.org/pdf/2310.15469
  6. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  7. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  8. Zhou et al., "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  9. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
  10. Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
  11. Anthropic (Claude Help Center), "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data