Skip to content

Rights and contracts

Synthetic data made from your records: who owns it and can it be sold?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

Who owns synthetic data generated from licensed data is decided mainly by the license, not by whoever ran the generator. The safest default for a supplier is to treat synthetic outputs as derived data, governed by the same permitted use, confidentiality and deletion terms as the source records, unless the license expressly says otherwise. Resale needs explicit permission.

Key takeaways

  • Synthetic records made from your data can still carry your trade secrets, business patterns and sometimes personal details.
  • Define synthetic data by name in the license; generic derived-data wording may not clearly reach it.
  • A buyer's right to generate synthetic data for its own training is different from a right to sell or share it.
  • Deletion clauses should state whether synthetic outputs must be destroyed when the license ends.

Who owns synthetic data made from your records?#

Ownership of synthetic data made from your records is decided mainly by the license between you and the buyer, because the law on who owns machine-generated output is unsettled and depends on the facts. Where the license is silent, each side will argue for the reading that suits it, which is a poor position to be in after delivery.

Copyright offers little help to either side. In the US, copyright generally protects human authorship, so purely machine-generated records may carry little copyright of their own. Control therefore rests mostly on contract terms and on trade secret and confidentiality protection for the source records, not on an ownership label.

For a supplier, the safe default is to treat synthetic outputs as derived data. They then sit under the same permitted use, confidentiality, security and deletion terms as the records they came from, unless the license grants something broader in plain words.

Why synthetic records are not automatically free of your rights#

Synthetic records are not automatically free of your rights because they are built to reproduce the patterns in your originals. A generator trained on a software company's support history learns its product names, failure modes, escalation paths and tone. Those patterns are often the know-how the license was written to protect.

Privacy risk does not disappear either. Generative models can memorize and reproduce rare records, so a synthetic set made from call transcripts may echo a real customer's situation closely enough to identify them. Whether synthetic records count as personal information depends on how they were made and tested, and on the law that applies.

The easy assumption fails on both counts: that synthetic means new, and new means owned by whoever generated it. Treat that assumption as a claim to be settled in the contract.

Three common contract positions on synthetic outputs#

Data licenses tend to settle synthetic outputs in one of three ways, and each suits a different deal. The right choice depends on how sensitive the source records are and what the buyer plans to build.

Use language matters more than ownership language. A buyer that owns synthetic outputs but cannot distribute them, must test them for similarity and must destroy them at termination holds a narrow right in practice.

Three common contract positions on synthetic outputs
PositionWhat the license saysWhen it fits
ProhibitedNo synthetic generation from the licensed recordsHighly sensitive records, or a first license with a new buyer
Permitted as derived dataGeneration allowed for the licensed purpose; outputs follow source terms, including deletionThe common middle ground for training and evaluation licenses
Permitted and owned by the buyerBuyer owns outputs, subject to listed limits on distribution and similarityLower-sensitivity records and commercial terms that reflect broader rights

Can synthetic data made from your records be sold?#

Synthetic data made from your records can be sold or relicensed by a buyer only if the license expressly allows it, and most supplier-friendly licenses do not. A right to train internal models is not a right to build a dataset product, and the license should say so in a sentence that leaves no room for argument.

The question also runs the other way. A supplier that owns its source records may be able to commission synthetic data itself and license that instead of, or alongside, the originals. Customer contracts, privacy notices and confidentiality duties still apply to the source, so the same rights review is needed before a synthetic version can be offered.

Which outputs count as derived data?#

Derived data covers more than synthetic records, and a license works best when it sorts each kind of output into a category. Buyers produce several artifacts from licensed records during training and evaluation, and they carry different levels of risk for the supplier.

The sorting below is a starting point for negotiation. The categories that most resemble the source records deserve the tightest terms, while outputs that cannot reproduce any record can usually be left with the buyer.

Which outputs count as derived data?
OutputUsual treatmentWhy
Synthetic records that mimic the sourceDerived data under source termsCan echo real records and business know-how
Labels, annotations and summaries of source recordsDerived data under source termsCarry the content of the records in condensed form
Evaluation sets built from source recordsDerived data, often with stricter sharing limitsFrequently circulated between teams and partners
Embeddings or indexes of source recordsDerived data; delete with the sourceCan support retrieval of original text
Trained model weightsOwned by the buyer, subject to use limitsRarely reproduce records in ordinary use, but memorization is possible, so use limits still apply
Aggregate statistics with no record-level detailUsually the buyer's to keepHard to trace back to individual records when groups are large enough

Clause checklist for synthetic and derived data#

A clause checklist keeps synthetic data from falling between definitions. Counsel can compare a draft license against each point below.

Read the checklist against the deletion and survival clauses together. A license that deletes source records at termination but lets synthetic copies survive has moved the value of the records into a form the supplier no longer controls.

  • Definitions: separate definitions for licensed data, derived data, synthetic data and model outputs, with synthetic data named expressly.
  • Permitted generation: whether synthetic generation is allowed at all, and only for the licensed purpose.
  • Ownership: who owns synthetic outputs, stated directly rather than left to general IP clauses.
  • Distribution: no sale, sharing or publication of synthetic data derived from the licensed records without written consent.
  • Similarity testing: a duty to test outputs for near-copies of source records and to remove them.
  • Deletion: whether synthetic data must be destroyed, with a certificate, at termination alongside the source records.
  • Audit and notice: a right to ask which synthetic sets exist and where they are held.

Illustrative: a software company fields a synthetic data request#

Illustrative: a fictional vertical software company licenses de-identified Zendesk tickets linked to Jira issues to a developer building a support agent. Midway through negotiation, the developer asks for the right to generate synthetic tickets from the records and to own the outputs outright.

The company's counsel agrees to generation for training and evaluation only. Outputs are treated as derived data, cannot be shared outside the developer's affiliates, must be tested for near-copies of real tickets and are deleted with the source records when the license ends. The developer gets the training flexibility it wanted, and the company keeps control of anything that could be sold.

How SourceX records rights in derived data#

SourceX settles synthetic and derived data in the Rights step of the SourceX five-step transaction. The permitted use recorded in the SourceX Evidence Packet states whether synthetic generation is allowed and under which limits, so the supplier, the buyer and any later reviewer read the same terms.

Because the supplier approves every step, a request for synthetic generation is treated as a scope change that goes back to the supplier for a decision, not a detail settled quietly between lawyers.

Frequently asked questions

Does de-identifying records before licensing remove the synthetic data question?

No. De-identification lowers privacy risk in the source records, but synthetic outputs can still reproduce your business patterns, product details and know-how. The license should govern synthetic data whether or not the source was de-identified.

Who owns a model trained on synthetic data made from our records?

Usually the developer owns its model, and licenses rarely give suppliers ownership of model weights. What a supplier can control is scope: what the model may be used for, whether synthetic training sets can be reused for other models, and what happens to them at the end of the term.

Is synthetic data personal information under US privacy laws?

It depends on the facts. Synthetic records that cannot reasonably be linked to a person are generally treated differently from records that can. If outputs closely mirror real individuals, privacy laws may still apply. Testing outputs for re-identification risk is the practical safeguard, and counsel assesses each case.

How can we check that synthetic outputs are not copies of our records?

Ask the license to name who tests, when and by what method. Common checks are exact and near-duplicate matching against the source, searches for rare details such as unusual product names or one-off incidents, and human review of samples by someone who knows the originals. Agreeing with the buyer to include a few unique marker records in the source set lets you test later whether outputs reproduce them.

Can a buyer combine our synthetic data with other suppliers' data?

Only if the license allows it. Combination inside the buyer's training pipeline is often permitted, but a supplier can still require that its derived data stays traceable, is not packaged into a product sold to others, and can be deleted at the end of the term.

Should synthetic generation rights change the commercial terms?

Broader rights usually change the economics of a deal, so synthetic generation is worth discussing as a commercial point rather than as boilerplate. There is no standard price; value depends on the buyer, the records and the scope granted.

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify