Skip to content

Software companies

How to use customer data to build AI features in your SaaS product legally

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

To use customer data for AI features lawfully, match each use to the rights it needs. Features that use a customer's data only for that customer usually fit existing terms. Training a shared model across customers needs clear contract rights and notice. Licensing to outside developers needs express permission. Build on the lowest rung your feature allows.

Key takeaways

  • Customer data rights form a ladder: in-tenant use, a shared product model, then licensing to third parties.
  • Each rung up needs more explicit contract language, more notice and stronger technical controls.
  • Calling an outside model provider is usually a subprocessor question, not a licensing question.
  • Terms changes generally work best going forward, with notice, and rarely reach older data automatically.
  • Third-party licensing is usually limited to company-owned records unless customers expressly opt in.

What is the rights ladder for customer data?#

The rights ladder sorts AI uses of customer data by how far the data travels from the customer's own purpose. The lower the rung, the more likely existing contracts already cover it; each step up needs clearer permission.

This is general information, not legal advice. Which laws and contract terms apply depends on your customers, their users and where they are, so counsel should review each feature before it ships.

What is the rights ladder for customer data?
RungExample featureRights usually neededCommon pitfall
In-tenant useSummaries, search or drafts built from one customer's records for that customerExisting right to process data to provide the serviceRetrieval that leaks across tenants
Outside model providerSending prompts with customer data to a hosted model APISubprocessor terms, DPA notice, provider terms barring training on inputsProvider terms that let it retain or learn from inputs
Shared product modelA model trained on pooled data from many customersExpress right to use data to develop models, notice, often opt-out or opt-inRelying on an improvement clause written for analytics
Third-party licensingLicensing records to an outside AI developerExpress customer permission, or company-owned records onlyTreating de-identified customer content as free to license

Rung one: features inside a customer's own tenant#

In-tenant features use a customer's records only to serve that same customer, which is why existing terms usually cover them. Drafting a reply from a customer's own knowledge base or summarizing their open work orders processes their data on their instructions.

The legal work is light, but the engineering must be strict. Retrieval indexes, caches and prompt logs have to respect tenant boundaries, and evaluation sets built from production data need the same access controls as the data itself.

If the feature calls a hosted model, the provider is typically a subprocessor. Update your subprocessor list as your DPA requires, and confirm in the provider's terms whether inputs and outputs may be retained or used for training.

Rung two: training a shared model across customers#

Training a shared model across customers is a new purpose for each customer's data, and it needs contract language that says so. A clause allowing the vendor to use data to improve the service may or may not stretch that far, and enterprise riders often narrow or prohibit it.

Privacy roles add another layer. Where your DPA casts you as a processor or service provider, using personal data to build your own product model may fall outside the customer's instructions under laws such as the GDPR or US state privacy laws. De-identification reduces the risk but does not remove the need to check the contract.

Many vendors handle this rung with an explicit program: updated terms that apply going forward, advance notice, a customer-level opt-out or opt-in, and technical flags that keep excluded tenants out of training sets.

Rung three: licensing customer data to outside developers#

Licensing customer data to outside developers sits at the top of the ladder because the data leaves your control for someone else's purpose. Standard SaaS agreements rarely grant that right, so it generally requires express permission from each participating customer.

Most SaaS companies start instead with records they own: engineering history, product decisions, internal documentation and support resolutions with customer details removed. Those records can often be licensed without touching customer content at all.

How to change your terms the right way#

Changing terms the right way means telling customers what will change, when, and what they can do about it. A quiet or retroactive change is the riskiest route, both with regulators and with the customers who read your terms closely.

Expect your largest customers to negotiate. A standard clause that most accounts accept may be struck from enterprise paper, so the training pipeline has to cope with a mix of permissions rather than one rule.

  • Collect every version of your paper: click-through terms, MSAs, order forms, DPAs, the privacy policy and negotiated riders.
  • Draft a clause that names the purpose, such as developing machine learning features, rather than leaning on general improvement language.
  • Apply it going forward and give advance notice in the manner your agreements require.
  • Offer an opt-out or opt-in where contracts, laws or customer expectations call for it.
  • Update your trust center, subprocessor list and security questionnaire answers at the same time.
  • Record which terms applied to which data, so training sets can be filtered by consent state.

What engineering controls make the rights real?#

Engineering controls make contract rights real by making data flows match what the terms allow. A clause that excludes opted-out tenants means little if the training pipeline cannot tell which tenants opted out.

Automated tools help but are not enough alone. Presidio, an open-source SDK for PII identification and anonymization, warns in its own documentation that automated detection cannot guarantee it will find all sensitive information and that additional protections should be used.

  • Tenant-level consent flags carried through every data pipeline.
  • Lineage for each training set: sources, dates, filters and the terms in force.
  • De-identification with human review of samples, not automated detection alone.
  • Deletion that reaches derived datasets when a customer leaves or opts out.
  • Access controls on evaluation sets and prompt logs equal to those on production data.

Illustrative: a landscaping software vendor plans its first AI features#

Illustrative: a fictional vertical SaaS company sells scheduling and estimating software to landscaping businesses. Its product team wants an assistant that drafts estimates from a customer's past jobs, and later a shared model that suggests prices across the customer base.

The CTO and counsel launch the first feature as in-tenant use, add the model provider to the subprocessor list and confirm the provider does not train on inputs. For the shared pricing model, they draft new terms with advance notice and an opt-in, and build tenant consent flags before any training starts.

Separately, the company considers licensing data to outside developers. It limits that discussion to its own engineering history and de-identified support records, leaving customers' job and pricing data out of scope.

Where SourceX fits on the ladder#

SourceX works only on the top rung, licensing records to outside AI developers, and its own rights in a deidentified dataset are set out in the signed supplier agreement. In the Rights step of the SourceX five-step transaction, customer content is separated from company-owned records and carved out unless permission is clear.

Each package that proceeds carries a SourceX Evidence Packet recording provenance, licensing rights, permitted use, the privacy record and release authorization, which keeps external licensing consistent with what your terms tell customers.

Frequently asked questions

Does sending customer data to a model API count as sharing it?

It is usually treated as processing by a subprocessor on your behalf rather than sharing for the provider's own purposes, provided the provider's terms restrict retention and training. Check those terms, list the provider where your DPA requires, and update questionnaire answers to describe the flow accurately.

Is de-identified customer data free to use for training?

Not automatically. Contracts may define customer data to include anything derived from it, and confidentiality duties can cover business information that is not personal data. De-identification lowers privacy risk, but the contract still decides whether the use is permitted.

Can we train on data from customers who have churned?

Usually only to the extent your terms allowed during their contract and any post-termination clauses permit. Many agreements require deletion or return of customer data after termination, which would rule out later training. Check retention and deletion terms before including former customers.

Is an opt-out enough, or do we need opt-in?

It depends on the data, the customers, the laws that apply and what your contracts promise. Opt-in is generally the safer choice for sensitive data, enterprise customers and any third-party licensing, while opt-out may suit some product improvement uses. Counsel should decide feature by feature.

What should we tell customers who ask whether their data trains our models?

Tell them which features use their data, whether any shared model is trained on it, which providers process it and how to opt out or in. Keep the answer identical across your trust center, terms and questionnaires, and update all three before any change takes effect.

Sources

  • Presidio is an open-source, MIT-licensed SDK for PII identification and anonymization in text and images. Source
  • Presidio's documentation warns that because it uses automated detection mechanisms, there is no guarantee it will find all sensitive information, and additional systems and protections should be employed. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify