Skip to content

Rights and contracts

Who owns an AI model trained on your licensed data?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

The AI developer that trains a model on your licensed data typically owns the model and its weights, while your company keeps ownership of the records themselves. Your control comes from the license: which models may use the data, whether fine-tunes and distilled models count, and what happens to data and derivatives when the term ends.

Key takeaways

  • A data license grants a right to use records; it does not make the supplier a co-owner of the model.
  • Your company keeps ownership of its records and can license them again unless it grants exclusivity.
  • Fine-tuned versions, distilled models, embeddings and synthetic data need to be addressed by name in the license.
  • Once records shape a model's weights they cannot be pulled back out, so scope and preparation must be settled before delivery.
  • Sharing in a model's commercial success is a negotiated payment term, not an ownership right.

Why does the buyer usually own the model?#

The buyer usually owns the model because the weights are the product of its own compute, engineering and many other data sources, of which your records are one input. A data license gives the developer permission to use your records for named purposes; it does not transfer a share of whatever the developer builds with them. Developers typically protect their weights as trade secrets and through contract, and nothing in a standard data license gives the supplier a claim on them.

Founders sometimes expect licensing to work like a joint venture, with an ownership stake in the resulting model. Developers generally resist this, because a model trained on many suppliers' records cannot practically have many co-owners. What suppliers negotiate instead are limits on use, transfer and survival, plus payment terms.

Ownership of a model also covers things suppliers rarely think about: the training code, the evaluation results, the safety tuning and the product built around the model. None of these passes to a data supplier under a standard license, and asking for them tends to slow negotiations without adding protection.

What does your company keep?#

Your company keeps ownership of the licensed records, every right the license does not grant, and the contract rights that limit how the buyer uses what it received. Those reserved rights are where most of a supplier's protection sits, so they deserve as much attention as the grant itself.

A useful test is to ask what you could do the day after the license ends. You could still use the records internally, license them to another developer if the first license was non-exclusive, and enforce any surviving terms. You could not reach into the buyer's model and take back what it learned, which is why the terms agreed before delivery matter most.

  • Ownership of the original records, such as support tickets, CRM histories, code reviews or project files.
  • The right to license the same records to other developers, unless the license is exclusive for a field or period.
  • Confidentiality protection for anything the license treats as confidential information.
  • Any use not granted, including retrieval, synthetic data generation or resale, where the license reserves it.
  • Remedies if the buyer uses the records outside scope, from deletion demands to damages.

Which derivatives should the license address?#

The license should address every derivative that can carry your records' influence beyond the delivered files. A derivative left unmentioned tends to be treated by the buyer as its own property without limits.

Distilled models are the gap suppliers most often miss. A developer can train a smaller model on the outputs of a larger one, and if the larger model was fine-tuned on your records, the smaller one inherits some of that learning without ever touching your files.

Which derivatives should the license address?
DerivativeWhat it isTypical buyer positionSupplier ask to consider
Base model weightsParameters of a general model trained on many sourcesOwned outright and retained after terminationNo transfer of your records with the model; confidentiality survives
Fine-tuned modelsA model adapted using your recordsOwned outrightLimited to named purposes; not marketed as built on your expertise
Distilled or student modelsSmaller models trained on a larger model's outputsOutside the license entirelyCovered if trained on a model fine-tuned on your records
Embeddings and vector indexesNumeric representations used for retrievalWorking files owned by the buyerDeleted with the records at termination
Synthetic dataNew records generated using yours as examplesOwned and reusableExcluded, or limited in use and retention
Labels and annotationsTags or ratings the buyer adds to your recordsOwned by the buyerLabeled copies follow the same use and deletion rules as the originals

Can a trained model reveal your records?#

A trained model can sometimes reproduce passages from its training data, especially text that appeared often or was unusual. That is why, for a supplier, what went into the model matters more than who owns it.

Remove personal details, customer names, credentials and confidential terms before delivery, because no contract term can pull them back out of the weights afterward. Then add output terms: the buyer takes reasonable measures against reproducing records verbatim and acts on takedown requests if a passage does surface.

Preparation also decides how much of your business a model can learn. Customer names, pricing and named employees rarely add value for a model developer, and removing them limits what a model could ever reveal. Keep the workflows and decisions that make the records useful, and leave out the identifiers that make them sensitive.

Illustrative: a founder asks for a share of the model#

Illustrative: the founder of a fictional field service software company is negotiating a license for engineering history from Jira, GitHub pull requests and code reviews. The founder's opening request is co-ownership of any model fine-tuned on those records.

The developer declines, because its coding models draw on many suppliers. Counsel reframes the ask into terms the developer accepts: the records may be used only to fine-tune and evaluate the developer's coding models, may not be transferred or sublicensed, and may not be used to build a product marketed as trained on the company's code. Source files and embeddings are deleted at termination, with a signed certificate.

The founder ends up with no ownership of the model but clear control over the records and their derivatives, which was the underlying concern all along.

What terms protect you without owning the model?#

The terms that protect a supplier without model ownership are permitted use, transfer limits, naming limits and survival. Each answers a specific worry that founders and CEOs raise in the first conversation.

What terms protect you without owning the model?
Founder worryClause that addresses it
Our records end up with a competitorNo transfer, resale or sublicensing; restrictions on competitor-facing products
Our name is used to market the modelNo use of the supplier's name or marks without written consent
The buyer keeps using our data indefinitelyA defined term, with deletion of records and derived data at the end
Promises lapse when the deal endsSurvival of confidentiality, use limits and output terms
We cannot check what happenedDeletion certificates and a limited audit or attestation right
The buyer is acquiredAssignment only with consent or to a successor bound by the same terms

How SourceX handles model rights#

SourceX records the agreed model rights in the SourceX Evidence Packet under permitted use, next to provenance, licensing rights, the privacy record and release authorization. The supplier approves the permitted models, derivative rules and post-termination terms during Approval, before anything moves in Delivery.

The rights SourceX holds in a deidentified dataset are set by the signed supplier agreement. Ownership, exclusivity and payment terms are set between the supplier and the buyer.

Frequently asked questions

Can we be paid more if the model succeeds commercially?

Some licenses include payments that continue over time, such as renewal fees or usage-linked amounts, but they are negotiated case by case and depend on whether the model's use can be measured. They are contract terms, not ownership. There is no standard rate, and value becomes clear only once a buyer engages.

Does licensing to one developer stop us licensing to others?

Not unless you grant exclusivity. A non-exclusive license leaves you free to license the same records elsewhere. If a buyer asks for exclusivity, narrow it by field of use, record type or period, so the rest of your archive stays available for other buyers.

Who owns what the model generates?

Model outputs generally belong to the developer or its users under their own terms, not to the data supplier. A supplier's interest is narrower: outputs should not reproduce its records or confidential details. Address that with output terms and careful preparation rather than an ownership claim.

Could our company be named as a training data source?

Possibly. Some laws and voluntary practices lead developers to publish summaries of their training data sources. If confidentiality about your participation matters, say so in the license, and ask counsel how any applicable disclosure rules interact with that promise.

Does it matter whether the buyer is a model developer or an application company?

It changes which derivatives exist. A model developer produces base and fine-tuned weights. An application company may only fine-tune or build retrieval on top of someone else's model, so its main derivatives are fine-tunes, embeddings and prompt libraries. Draft the derivative clause around the buyer's actual workflow.

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify