AI data market
Training license vs retrieval license: what's the difference?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
A training license lets an AI developer use your records to change a model's weights, so their influence stays in the model after the copies are deleted. A retrieval license lets the developer keep records in an index and look them up at answer time, often showing passages to users. Training is harder to undo; retrieval exposes more text.
Key takeaways
- A training license grants lasting influence: deleting the dataset does not remove what a model learned.
- A retrieval license keeps records in a searchable index that can be updated or switched off.
- Retrieval usually shows passages to end users, so confidential business records rarely suit it.
- Retention, refresh, display, attribution and termination clauses differ sharply between the two.
- Name the permitted use precisely, because a vague grant for any AI purpose covers both.
What is the difference between a training license and a retrieval license?#
A training license permits a buyer to use records to train or fine-tune a model, while a retrieval license permits the buyer to keep records in an index that a model consults when it answers. The first changes the model itself. The second leaves the model unchanged and treats your records as a reference library it can search.
Retrieval is often called retrieval-augmented generation, RAG or grounding. Publisher deals show why the distinction gets spelled out: the December 2023 Axel Springer agreement with OpenAI stated that content from Axel Springer brands would be used to advance the training of OpenAI's large language models, a point worth stating because publisher content can also be displayed in AI answers. Business records raise the same question with different stakes, because most of them were never meant to be read by outsiders.
How the two licenses compare, term by term#
The two licenses differ on buyer rights, retention, refresh, display and what ending the license achieves. Neither is safer in general: training is harder to reverse, while retrieval is easier to reverse but shows more of the original text.
| Term | Training license | Retrieval license |
|---|---|---|
| What the buyer gets | Right to use records to change model weights | Right to store, search and draw on records at answer time |
| Retention | Training copies deleted or returned on schedule; the trained model remains | Records kept in the index for the license term |
| Refresh | Usually a fixed delivery, sometimes periodic updates | Often an ongoing feed so answers stay current |
| Display to users | Usually restricted; outputs should not reproduce records | Passages may be shown, quoted or cited |
| Ending the license | Stops new training; past models usually keep what they learned | Records removed from the index and access ends |
| Main risk to the licensor | Permanent influence and possible memorization | Verbatim text exposed to outside users |
| Records that fit | Internal workflows: tickets, code reviews, job and quality records | Content written to be read: public help articles, documentation, catalogs |
Why most business records suit training, not retrieval#
Most internal business records suit training licenses because their value lies in patterns of work, not in being quoted. A model learns how support agents calm an upset customer, how engineers review a risky change or how dispatchers reorder a day, and the individual tickets never need to appear in an answer.
Retrieval puts the records themselves in front of users. Even after personal and confidential details are removed, a quoted support thread or engineering discussion can reveal pricing, product plans or a customer's situation. Content written for outside readers, such as public knowledge base articles, product catalogs and technical documentation, fits retrieval far better.
There is a third case: evaluation. Buyers sometimes license records only as a test set to measure model performance, with no training and no display, which calls for its own terms again.
Clauses that change with the license type#
The clauses that change most between the two license types are permitted use, retention, refresh, display and what survives termination. A license that simply grants use for artificial intelligence purposes covers both, so the grant should name each permitted use and exclude the rest.
- Permitted use: training, fine-tuning, evaluation, retrieval or a defined combination, with anything unnamed excluded.
- Model scope: which models or products may use the records, and whether successor models are covered.
- Retention and deletion: when raw copies, processed copies and indexes are deleted, with written certification.
- Refresh: delivery schedule, update format and who bears the cost of an ongoing feed.
- Display and attribution: whether passages may be shown, how much, and whether a source credit is required.
- Output controls: obligations to prevent a model from reproducing records verbatim.
- Exclusivity and field of use: whether others may license the same records, and in which markets.
- Termination effects: what happens to trained models, indexes and cached copies when the license ends.
- Audit, security and indemnity: how compliance is checked and who carries which risks.
Where permitted use gets recorded#
Permitted use should be recorded with the dataset itself, not only in the contract. Metadata standards now reflect the distinction: the Data & Trust Alliance's Data Provenance Standards include intended data use and license to use among the elements of their Use group, so a buyer's data team can see what a package allows without reading the agreement.
That matters inside large buyers, where the people who negotiate a license are rarely the people who load the data. A package labeled for training only is less likely to end up in a retrieval index by mistake.
How do the economics differ?#
The economics of the two licenses differ mainly in timing. Training licenses are commonly priced for a defined delivery, paid at signing or in stages, because the buyer gets the benefit once the model is trained. Retrieval licenses more often resemble subscriptions, because their value depends on continued access and fresh content.
For a CFO, that changes the revenue profile and the accounting questions. Under ASC 606, a license to functional intellectual property is generally a right to use the IP as it exists when granted, with revenue recognized at a point in time, while sales- or usage-based royalties are recognized only when the later usage occurs. A one-time training delivery and a usage-priced retrieval feed may therefore be recognized quite differently, though the treatment of a specific dataset license depends on its terms.
Refresh obligations also carry real costs for the supplier, from staff time to export tooling, which belong in the pricing discussion. Review the structure with your accountants before agreeing to it.
Illustrative: a SaaS company splits its records by license type#
Illustrative: a fictional field service software company is approached about its records. It holds a public help center of how-to articles, years of Zendesk tickets linked to Jira issues, and GitHub pull requests with review comments.
Its general counsel splits the scope. The help center articles, already public, are offered under a retrieval license with attribution and a regular content feed. The tickets, issues and code reviews, after personal and confidential details are removed, are offered only under a training license with no display rights, deletion of training copies after use and output controls against verbatim reproduction.
Each package carries its own permitted-use definition, so neither license can be read to cover the other.
How SourceX documents permitted use#
SourceX documents permitted use for every package in the SourceX Evidence Packet, alongside provenance, licensing rights, the privacy record and release authorization. Whether a buyer may train on records, evaluate with them or retrieve from them is written down before delivery, not inferred afterward.
Within the SourceX five-step transaction, the Rights step identifies which record families can support which uses, and the Approval step lets the supplier accept or decline each one. Data is licensed, not sold outright, so the company keeps ownership whichever license type it grants.
Frequently asked questions
Can a buyer remove what a model learned from our records?
Not reliably today. Deleting the training copies is straightforward and can be certified, but removing the influence of specific records from a trained model is an open technical problem. That is why training licenses focus on what is delivered, how copies are handled and how outputs are controlled, rather than promising removal from the model.
Can one contract cover both training and retrieval?
Yes, if each use is named and given its own terms. A combined license should say which record families may be used for which purpose, with separate retention, refresh and display rules. A single vague grant is the common mistake, because it leaves the buyer free to use internal records in ways the supplier never intended.
Is fine-tuning a training use?
Fine-tuning is a training use because it changes model weights, even when the model is small or specialized. Some licenses distinguish pretraining, fine-tuning and evaluation and price or restrict them differently. If the distinction matters to you, define each term in the contract rather than relying on industry usage.
What should a retrieval license say about attribution?
A retrieval license should state whether answers that draw on your content must credit or link to you, how long quoted passages may be, and whether the buyer may summarize without credit. For public content such as help articles, attribution can be part of the value; for anything confidential, retrieval is usually the wrong license.
Does a retrieval license let the buyer train on the same records?
Only if the contract says so. Records sitting in a retrieval index are technically easy to reuse for training, which is why the license should prohibit any use it does not name and require the buyer to keep retrieval copies separate. Audit and certification rights help confirm that the boundary holds.
Sources
- The Use group of the Data & Trust Alliance Data Provenance Standards includes elements for confidentiality classification, consent documentation location, privacy-enhancing technologies applied, allowed and excluded processing and storage geographies, license to use, intended data use, and copyright, patent and trademark status. Source
- The December 2023 Axel Springer-OpenAI agreement states that quality content from Axel Springer media brands will be used to advance the training of OpenAI's large language models. Source
- Under ASC 606, a license to functional intellectual property is generally a right to use the IP as it exists when the license is granted, with revenue recognized at a point in time, subject to exceptions. Source
- ASC 606-10-55-65 requires revenue for a sales-based or usage-based royalty on a license of IP to be recognized only when the later of the subsequent sale or usage occurs or the related performance obligation is satisfied. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.