AI data license terms explained: permitted use, exclusivity, snapshots and deletion
The terms that matter most in an AI data license are permitted use, field of use, exclusivity, duration and territory, ownership of derivatives and trained models, retention and deletion, and the audit, warranty, indemnity and payment terms that allocate risk. Permitted use should name each activity you need, from training and fine-tuning to evaluation and commercial deployment. Settle in writing whether models trained during the term survive its end, because that is hard to fix after training.
In this guide
Key takeaways
- Name every activity you need in the permitted use, because a license usually grants only the uses it lists.
- Exclusivity is a scope, not a yes-or-no, so define the snapshot, use and period it covers.
- Decide in writing whether models trained during the term survive its end, before any training starts.
- Deletion duties should say whether they reach derived datasets and model weights, and how you prove deletion.
- Rights and de-identification warranties, backed by indemnities, decide who pays if the data turns out not to be cleared for your use.
This guide is general information, not legal advice. Have counsel review any license against your intended use and the laws that apply to you.
The terms at a glance
| Term | In plain language | Ask before signing |
|---|---|---|
| Permitted use | What you may do with the data | Are training, fine-tuning, evaluation, synthetic data and commercial deployment each named? |
| Field of use | Where those uses and the resulting models may be applied | Does it cover products you may build later? |
| Exclusivity | Whether others can license the same data | Exclusive for which snapshot, which use and how long? |
| Duration | How long you may use the data | What survives the end of the term? |
| Territory | Where data may be stored, accessed and used | Do residency, transfer or export rules apply? |
| Derivatives and models | Who owns what you build from the data | Do you own your models, labels and synthetic data? |
| Retention and deletion | What you must destroy, when, and how you prove it | Does deletion reach derived datasets or model weights? |
| Audit | How the licensor checks your compliance | What scope, notice and frequency, and who pays? |
| Attribution and confidentiality | What you may say about the data and the deal | Can you name the source or publish eval results? |
| Warranties and indemnities | Promises about the data, and who pays if they prove wrong | Who gives the rights warranty, and is it capped? |
| Refreshes | Whether new data arrives during the term | Do schema and pseudonyms stay stable across deliveries? |
| Payment | How and when the licensor is paid | Is payment tied to delivery and acceptance? |
Permitted use
Permitted use is the list of activities the license allows; unlisted activities are usually not licensed. Name each one you need:
- Training: pretraining or continued pretraining.
- Fine-tuning: supervised fine-tuning, preference training and reinforcement learning.
- Evaluation: benchmarking and regression testing, internal or published.
- Retrieval: indexing records for use at inference time, which exposes them to users more directly than training does.
- Synthetic data generation: using records as seeds or references for generated data.
- Commercial deployment: offering models trained on the data in products, not only internal research.
Also settle who may touch the data: affiliates, contractors such as annotators, and cloud or model providers, including third-party fine-tuning APIs.
Field of use
Field of use limits where the permitted uses apply: a domain, product line or market, such as support automation for financial services. It can also exclude uses, such as products that compete with the licensor. A narrow field may cost less but can follow your model for life, so check it against products you may build later.
Exclusivity and dataset snapshots
Exclusivity limits whom else the licensor may license the same data to. A non-exclusive license leaves the licensor free to license the same records to other buyers; an exclusive license rules that out within the scope it defines. The licensor keeps using its records to run its business, so exclusivity concerns other licensees and, where the license says so, the licensor's own AI development.
Exclusivity is a scope, not a yes-or-no. It can be limited to:
- A dataset snapshot: a defined extract, such as records from named systems over a set date range, frozen at delivery.
- A permitted use or field: for example, training for one domain.
- A period: after which others may license the snapshot.
Unless the license says otherwise, exclusivity over a snapshot does not cover records created after it or uses outside the exclusive scope. Buyers seek it to differentiate their models and to keep evaluation snapshots out of competitors' training data. Some SourceX programs include exclusivity for an agreed dataset snapshot or permitted use over a defined period; it is negotiated in the license and usually affects price.
Duration and territory
Duration is how long you may use the data, whether a fixed, renewable or perpetual term. Separate two questions: how long you may use the raw data, and what survives the end of the term, such as models trained during it and the right to keep serving them. Check whether the license transfers if your company is acquired.
Territory is where the data may be stored, accessed and used. Personal data leaving the European Economic Area needs a GDPR transfer mechanism, such as an adequacy decision or standard contractual clauses. Export controls and the licensor's customer contracts can also restrict access and location.
Derivatives and model ownership
Derivatives are anything you create from the data: cleaned and relabeled versions, annotations, embeddings, evaluation sets, synthetic data, model weights and outputs. A common split is that the licensor owns the data and any derivative that contains or can reconstruct records, while you own your models, outputs and annotations. Do not assume it; get explicit answers on three points:
- Whether models trained during the term may be used and commercialized after it ends.
- Whether synthetic data and evaluation sets built from the records count as the data, with its restrictions and deletion duties, or as your derivatives.
- Whether you must take reasonable steps to stop models reproducing licensed records verbatim.
Retention and deletion
Retention and deletion clauses say what you must destroy, when and how you prove it. Expect deletion of raw data and copies when the license ends, with a written deletion certificate and a carve-out for backups that expire on normal rotation. Negotiate whether deletion reaches derived datasets that contain records, and whether it reaches models, since deleting weights after training is costly.
Personal data raises the stakes. The European Data Protection Board has said that AI models trained on personal data cannot in all cases be considered anonymous and must be assessed case by case (EDPB Opinion 28/2024), one reason to de-identify thoroughly before delivery. Also agree how mid-term removal requests work, for example when a licensor's customer withdraws authorization.
Audit
An audit clause lets the licensor check your compliance. Agree the scope (usage records, access logs and deletion records, not model weights or unrelated systems), notice, frequency, whether an independent auditor does the work, and who pays; often the licensor pays unless the audit finds a material breach.
Attribution and confidentiality
Attribution terms say whether you may, or must, name the data source; many licensors prefer not to be named. Confidentiality clauses treat the data, and often the deal, as confidential, so check that they still allow what you need to disclose: evaluation results or examples, training data descriptions in model documentation and, if you provide a general-purpose AI model in the EU, the public summary of training content that the EU AI Act requires.
Warranties and indemnities
Warranties are promises about the data; indemnities say who pays if a third party makes a claim. Ask the licensor to warrant that it may license the data for the permitted use, that collection followed applicable law and the notices given at the time, that de-identification followed the agreed spec, that no earlier grant conflicts with yours, and that the delivery matches the manifest. You will usually warrant that you stay within scope, do not re-identify anyone and keep the data secure.
Indemnities typically follow the warranties: the licensor covers claims that it lacked rights, and you cover claims arising from misuse. Check liability caps and whether rights and privacy indemnities sit outside them. With an intermediary, confirm who gives each warranty: the data owner, the intermediary or both. Licensors usually disclaim fitness for purpose, so quality protection comes from acceptance criteria and remedies such as re-delivery.
Refreshes and updates
A license can cover one snapshot or a series of deliveries. For refreshes, agree the cadence, a stable schema and pseudonym mapping so IDs join across deliveries, the same de-identification, pricing, whether exclusivity extends to new data, and how corrections and takedowns of earlier records work.
Payment structures
Payment terms decide when money moves and who carries the risk that the data disappoints. Common structures, often combined:
- One-time fee for a snapshot: simple, but quality risk sits with you after delivery.
- Milestones tied to delivery and acceptance of each tranche: shifts some quality risk back to the licensor.
- Subscription for ongoing refreshes: spreads cost, with renewal risk on both sides.
- Volume-based pricing per record, hour or file: needs a precise counting rule.
- Exclusivity premium: an added fee for exclusive scope.
- Royalties or revenue share: hard to administer, because model revenue is hard to attribute to one dataset.
Where SourceX fits
SourceX agrees scope, permitted use, pricing and obligations in writing for each dataset it sources. Delivery follows the partner's approval, an executed agreement and explicit authorization of buyer access. Supply is not guaranteed. Before signing, work through the due diligence checklist; the full process is in how to license proprietary data for AI training. To start, send a data request with your licensing requirements.
Related dataset types
- Enterprise workflow and task execution histories
Linked task trajectories from request to outcome, across every tool the work touched
- Human feedback and QA-scored work
Work items with scores, verdicts and corrections from the people who reviewed them
- First-person video of skilled manual work
First-person video of skilled workers doing real tasks at partner businesses
Questions
Who owns a model trained on licensed data?
The license decides, and a common position is that the licensee owns its model weights and outputs while the licensor keeps ownership of the data. The license can still restrict the model: a field-of-use limit can follow it, a deletion clause can reach it, and some licenses require steps to stop it reproducing licensed records. Read the derivatives clause together with the termination and survival clauses, since together they decide what you may do with the model after the term.
What happens to a trained model when the data license ends?
Whatever the license says, which is why it should say so explicitly. Typical positions are that models trained during the term survive while the raw data is deleted; that models survive but no further training on the data is allowed; or, in stricter licenses, that deletion extends to the models. Negotiate this before training starts, because afterwards the only remedy may be retraining without the data.
What does exclusivity for a dataset snapshot mean?
It means the licensor will not license a defined extract of its data, such as records from named systems over a set date range, to other buyers for the agreed use during the agreed period. Unless the license says otherwise, it does not cover records created after the snapshot or uses outside the exclusive scope. Some SourceX programs include this kind of exclusivity; it is negotiated in the license and usually affects price.
What is the difference between permitted use and field of use?
Permitted use lists the activities you may perform with the data, such as training, fine-tuning, evaluation or synthetic data generation. Field of use limits where those activities and the resulting models may be applied, such as a domain, product line or market. A license can allow training broadly but restrict deployment to one field, so check both clauses against the products you plan to build.
How does an evaluation-only license differ from a training license?
An evaluation-only license lets you test models on the data but not train on them, so the licensor gives up less and the terms are narrower. Keep the data out of training pipelines, restrict access, and agree whether you may publish scores or examples. If you later want to train on the same records, you need an amended license, and the evaluation set loses its value as held-out data.
Ready to source data?
Send the domain, modality, volume, format, timeline and permitted use you need. SourceX will match it against partner businesses.
Updated 3 October 2026.