Rights and contracts
Does a training license cover fine-tuning, RAG and evaluation?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
A training license does not automatically cover fine-tuning, retrieval-augmented generation (RAG) or evaluation; it covers whatever the grant clause defines as training. Because RAG keeps your records readable at answer time and evaluation turns them into test sets, a supplier-friendly license names each use separately and treats anything unnamed as not granted.
Key takeaways
- A grant clause covers only the uses it defines, and an undefined word like training invites a dispute over scope.
- Pre-training and fine-tuning change model weights; RAG stores your records and retrieves them while answering users.
- RAG is closer to an ongoing content license than to training, so it deserves its own clause, term and deletion rule.
- Evaluation use can still expose records if test sets are shared, published or reused across many model versions.
- A strong grant lists permitted uses, excluded uses, covered models and derived data such as embeddings and synthetic records.
What does 'training' mean in a data license?#
Training in a data license means what the definitions clause says it means, and nothing more certain than that. Where the word is undefined, each side reads it in its own favor: the buyer as any use that improves a model, the supplier as one pass of model training on the delivered files.
In technical usage, training means adjusting a model's weights by exposing it to examples. Fine-tuning is training in that sense, but it happens later, on a smaller and targeted dataset, often for a specific product. RAG and evaluation do not change weights at all, which is why a bare reference to training may or may not reach them.
Supplier-side counsel should assume that any ambiguity will be resolved later, under pressure, by people who were not in the negotiation. Defining each use at signing costs far less than arguing about it afterward.
How do pre-training, fine-tuning, RAG and evaluation differ?#
Pre-training, fine-tuning, RAG and evaluation differ in what happens to your records, how long they stay in active use and how close they come to the people using the model. The table sets out the differences a license should account for, including two adjacent uses that buyers sometimes fold into a training grant.
| Use | What happens to your records | Stay in active use? | Exposure to end users | Name it separately? |
|---|---|---|---|---|
| Pre-training | Mixed into a very large general corpus to build a base model | Absorbed into weights; source files may be retained | Indirect; memorized passages are possible | Yes |
| Fine-tuning | Used to adapt a base model to a task, domain or style | Absorbed into the weights of a specific model | Indirect but more targeted | Yes, with covered models listed |
| RAG or grounding | Indexed, often as embeddings, and retrieved to answer queries | Yes, read whenever a query matches | Direct; outputs can quote records | Yes, with its own term and deletion rule |
| Evaluation | Held out as test sets to measure model performance | Reused across model versions | Low unless test sets are published | Yes, with publication limits |
| Synthetic data generation | Used as seeds or examples to generate new training records | Lives on in derived datasets | Indirect | Yes, or exclude |
| Prompting at inference | Pasted into prompts as examples or context | Yes, per request | Direct | Usually exclude unless needed |
Why RAG rights deserve their own clause#
RAG rights deserve their own clause because retrieval keeps your records working as a live reference library inside someone else's product. A model grounded in your support tickets or project files can quote them, summarize them for third parties and keep doing so for as long as the index exists.
That makes RAG closer to a content license than to a training license. The questions change: who can query the index, whether outputs may reproduce records verbatim, how quickly the index is rebuilt when you withdraw records, and whether embeddings count as copies. Vector indexes usually store the source passages next to their embeddings, and research has shown that embeddings alone can sometimes be used to approximate the original text, so treat both as derived data that is deleted with the records.
If a buyer signs for training rights and later asks to add retrieval, that is a new negotiation, not a clarification of the old grant.
Is evaluation a lighter use?#
Evaluation is a lighter use in one sense, because test sets do not change model weights, but evaluation data can end up more exposed than training data. Developers reuse test sets across many model versions, share them with contractors who score outputs, and sometimes publish benchmarks so others can compare results.
Evaluation also carries an interest that suppliers and buyers share: a test set loses its usefulness once it leaks into training corpora. Buyers therefore often want evaluation records kept out of any training, which suits a supplier that wants to limit spread. Write both points into the license: evaluation only, no training on the test set, no publication, and a named list of who may see it.
What a supplier-friendly grant clause includes#
A supplier-friendly grant clause lists what is permitted and reserves everything else. The elements below are the ones most often missing from a buyer's first draft.
Wording choices inside each element matter as much as the list itself. The comparison shows narrow phrasings to prefer and broad ones to question. Provenance metadata standards point the same way: the Data & Trust Alliance's Data Provenance Standards include metadata elements for license to use and intended data use, so the permitted uses can travel with the dataset rather than sit only in the contract.
- Defined terms for pre-training, fine-tuning, retrieval or grounding, evaluation and synthetic data generation.
- A permitted-use list naming which of those uses the buyer receives, with everything unnamed reserved to the supplier.
- Covered models, stated plainly: a named model, a model family or the buyer's general models.
- Derived data rules covering embeddings, vector indexes, labels, synthetic records and evaluation results.
- Who may use the records: the buyer only, or also affiliates, contractors and the buyer's own customers.
- Term and post-termination rules for each use, since a retrieval index and a trained model end differently.
| Grant term | Narrow wording to prefer | Broad wording to question |
|---|---|---|
| Purpose | Fine-tuning and evaluation of named models | Any machine learning or AI purpose |
| Models | The buyer's models listed in a schedule | Any current or future model |
| Users | The buyer and named contractors under confidentiality | The buyer, its affiliates and its customers |
| Derived data | Deleted or restricted along with the records | Owned by the buyer without limit |
| Retrieval | Excluded unless separately granted | Included within training |
Illustrative: a SaaS company splits its grant#
Illustrative: a fictional vertical software company serving property managers holds years of Zendesk tickets linked to Jira issues and release notes. A model developer asks for training rights to build a support agent and to measure that agent against real resolutions.
The company's counsel splits the request into three uses. Fine-tuning on de-identified tickets and issues is granted for the developer's support-agent models. A held-out set of resolved escalations is licensed for evaluation only, with no training and no publication. Retrieval is excluded, because the developer's product would let other software companies query the index.
Later, the developer asks to add retrieval. The company treats it as a new grant with its own term, an obligation to rebuild the index when records are withdrawn and deletion of the vector index at termination.
How SourceX scopes permitted use#
SourceX records permitted use as one of the five elements of the SourceX Evidence Packet, alongside provenance, licensing rights, the privacy record and release authorization. Each use a buyer requests, whether fine-tuning, retrieval or evaluation, is listed separately, so the supplier approves it by name during the Approval step of the SourceX five-step transaction.
SourceX's role is to make the scope visible to the supplier, the supplier's counsel and the buyer before Delivery.
Frequently asked questions
Are embeddings of our records a copy of the data?
Treat them as derived data that carries your content. Embeddings are numeric representations, but retrieval systems usually store the original passages alongside them, and embeddings alone can sometimes be used to approximate the text. A license should say whether embeddings may be created, who may hold them and that they are deleted along with the records.
Can a buyer use our records to generate synthetic training data?
Only if the license permits it. Synthetic records generated from your tickets or project files can carry your workflows, terminology and occasionally fragments of the originals into datasets that outlive the license. Either exclude synthetic generation or permit it with limits on use, transfer and retention after termination.
Can the buyer fine-tune models for its own customers with our records?
That is a separate question from fine-tuning the buyer's own models, and the license should answer it directly. Fine-tuning on behalf of enterprise customers can place your records inside models you never see, held by companies you never approved. Suppliers can limit fine-tuning to the buyer's own models unless they approve each end customer.
Can we grant evaluation rights first and add training later?
Yes, and starting narrow is a reasonable way to build trust on both sides. An evaluation-only license lets a buyer test whether the records are useful while limiting the supplier's exposure. Write any later expansion as a new grant with its own terms rather than an automatic upgrade.
What if an older contract says 'machine learning purposes'?
Read it broadly when assessing your exposure and narrowly when negotiating. A phrase like machine learning purposes could be argued to cover training, fine-tuning, retrieval and evaluation alike. If the license is still in force, consider an amendment that defines the uses; if it is up for renewal, replace the phrase with named uses.
Sources
- The Use group of the Data & Trust Alliance Data Provenance Standards includes elements for license to use and intended data use, among others. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.