Skip to content

Software companies

Can you license synthetic data generated from customer data?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

Synthetic data generated from customer data can be licensed only if the rights in the source data allow it, because generating new records does not erase customer contracts, confidentiality or privacy duties. Run three checks: may you use the source data to build the generator, does the output still reveal customers, and does the license disclose its origin.

Key takeaways

  • Rights in the source records carry into synthetic data derived from them.
  • Training or conditioning a generator on customer data is itself a use the customer contract must allow.
  • Synthetic records can still reveal real people or customer secrets, so they need privacy testing.
  • A licensed synthetic dataset should disclose that it was derived from customer records, and how.

Why does synthetic data not reset the rights question?#

Synthetic data does not reset the rights question because it is produced from the real records. A generator learns patterns from source data or is prompted with it, so building the generator is a use of that data, and the output carries traces of it.

Most B2B agreements limit customer data to providing the service, and confidentiality clauses protect the customer's business information whatever form it takes. A synthetic set that mirrors one customer's pricing, order mix or support issues can reveal exactly what that customer considered confidential, even with every name removed.

Treat this as general information, not legal advice: contracts and privacy laws differ, so these checks are assessed deal by deal with counsel.

Check one: may you use the source data to build the generator?#

The first check is whether your rights in the source records cover training or conditioning a generator, and then licensing what it produces. The answer depends mostly on whose records they are.

Aggregated data clauses rarely help here. They usually allow statistics that do not identify the customer, not a model that can produce realistic customer-like records on demand.

Check one: may you use the source data to build the generator?
Source recordsStarting position for synthetic generation
Company-owned internal records, such as engineering issues and internal docsUsually within your control, after personal data review
Support tickets from business customersMixed; the customer agreement and data processing addendum decide
Customer content stored in your productGenerally needs express customer permission
Data licensed to you by a third partyCheck whether the license allows derivatives and sublicensing
Public or scraped dataCheck source terms; public does not mean unrestricted

Check two: does the output still reveal customers?#

The second check is whether the synthetic output reproduces or points back to real records. Generators can memorize rare records, unusual combinations and verbatim phrases, especially in free text, and outliers are often the records most likely to identify someone.

Run the tests on the final output, not on a pilot batch. Changing generator settings, adding source records or increasing output volume can change what leaks, so repeat the tests whenever any of those change.

Record which privacy-enhancing techniques were used. The Data & Trust Alliance's Data Provenance Standards include a code list of such techniques, including differential privacy, k-anonymity, masking and pseudonymization, which gives buyers and sellers a shared vocabulary for describing the work.

  • Compare each synthetic record with its nearest real records and flag close matches.
  • Search the output for real names, email domains, account numbers and distinctive phrases.
  • Plant known canary records in the training data and check whether they reappear.
  • Test whether an outsider could tell if a specific record was in the training set.
  • Have a reviewer who knows the source customers read a sample.

Check three: does the license say what the buyer is getting?#

The third check is whether the license describes the dataset honestly. A buyer needs to know that the records are synthetic, that they were derived from customer data, under what permission, by what method, and what privacy testing showed.

Avoid warranting that synthetic data contains no personal data at all. Warrant the process instead: the generation method, the tests run and their results, with a commitment to remove anything identified later. Provenance standards help here too, since the Use group of the Data Provenance Standards includes elements such as license to use, intended data use and privacy-enhancing technologies applied.

How can a seller confirm the buyer stayed within the license?#

A seller confirms permitted use mostly through contract mechanisms, because no one can inspect a trained model and see which records went into it. That applies to synthetic and real data alike, and plain guidance on it is scarce.

Common tools include a written attestation that the data was used only for named models or purposes, a deletion certificate when the term ends, records of which training runs used the data, a ban on redistribution and resale, and audit rights exercised through an independent third party. None is perfect alone, so the license should combine several and set a remedy if a breach is found.

For synthetic data, add one more term: the buyer should not try to reverse the generation or re-identify source records, and should report any real record it finds in the set.

When is synthetic data the right tool?#

Synthetic data is the right tool when it solves a specific gap that real records cannot fill, not when it is used to route around restrictions on customer data. Used that way, it rarely survives a buyer's diligence.

Good uses tend to share a feature: the source rights are already clear. A company can generate rare edge cases from its own engineering or operations records to round out an evaluation set, produce a structurally faithful sample that shows a buyer field layouts and record types without exposing any real record, or test its own de-identification pipeline before running it on the real archive.

Weak uses look different. Generating customer-like records from data you could not license directly, then calling the output new, carries the original restriction forward and adds a disclosure problem on top.

Illustrative: a returns management software company weighs a synthetic set#

Illustrative: a fictional returns management software company serves online retailers and wants to license synthetic return-and-refund cases built from its customers' records. The general counsel reads the master agreement and finds that customer data may be used only to provide the service, with an aggregated data clause limited to statistics.

The company decides not to train a generator on customer records without permission. Instead, it builds a smaller synthetic set from its own support tickets about the product and its internal operations playbooks, offers an opt-in addendum to customers who want to contribute, and runs memorization and nearest-match tests before any scoping. The license draft states the source, the method and the test results.

The resulting package is smaller than the first idea, but every record in it can be traced to a source the company controls or a customer that said yes, and the general counsel can explain that in one paragraph.

How SourceX views synthetic and derived datasets#

SourceX treats synthetic and other derived datasets as inheriting the restrictions of their source records. In the Rights step of the SourceX five-step transaction, every derived set is traced back to the records it came from and the permissions behind them.

The SourceX Evidence Packet then records provenance, including that the data is synthetic and how it was made, along with licensing rights, permitted use, the privacy record and release authorization. The company approves every step, and SourceX's dataset rights are set out in the signed supplier agreement.

Frequently asked questions

Is synthetic data personal data?

It can be. Privacy laws generally look at whether a person can be identified, directly or by combining data. Well-generated synthetic data may fall outside that test, but poorly generated data can reproduce real people. Test for identifiability, document the results and review the conclusion with counsel.

Can we license the generator instead of the data?

A generator is a different asset with greater risk. It encodes patterns from the source data, can sometimes be prompted to reproduce training records, and lets the buyer create unlimited output. The same source rights questions apply, with stronger controls on use.

Is synthetic data worth less than real operational records?

It depends on the buyer and the use. Many AI developers prize real operational records for realism, while synthetic data can add coverage of rare cases or fill gaps. No value is known until a buyer engages with a specific dataset.

Who owns a synthetic dataset we generate?

The company that creates it generally holds whatever rights exist in the output, but ownership does not override contract limits on the source data. Copyright protection for machine-generated records is uncertain, so rely on contract terms rather than ownership alone.

Can we generate synthetic data from our own internal records?

Yes, and the rights path is usually simpler. Internal engineering, operations and support records are company-controlled, though they still contain employee and customer details that need the same leakage tests.

Sources

  • The Data Provenance Standards' Privacy Enhancing Tools code list includes data anonymization, encryption, masking, minimization, redaction, differential privacy, k-anonymity, l-diversity, pseudonymization and tokenization, among others. Source
  • The Use group of the Data & Trust Alliance Data Provenance Standards includes elements for consent documentation location, privacy-enhancing technologies applied, license to use and intended data use. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify