Skip to content

Rights and contracts

Can we use customer data to train our own AI features?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

A software company can often use customer data to train its own AI features, but only when its contracts permit that use, its data processing agreements allow it, and its privacy notices tell customers and their users plainly. Training a feature that serves the same customers is usually easier to justify than licensing the same data to an outside developer.

Key takeaways

  • Three checks decide the answer: contract permission, your privacy role and what you told customers and their users.
  • As a processor or service provider, a software company usually needs contract terms that let it use customer data beyond each customer's own account.
  • Retrieval within one account and per-customer models are easier to support than one model trained across all customers.
  • Negotiated no-AI-training clauses override general improvement rights for the customers who have them.
  • Licensing the same data to an outside AI developer is a separate question with a higher bar.

What decides whether you can train on customer data?#

Whether you can train on customer data is decided by three checks: what your contracts permit, what role privacy law gives you over the data, and what you told customers and their users. All three need to line up; a strong contract cannot fix a privacy notice that said the opposite.

The design of the AI feature matters as much as the data. Answering a customer's question from that customer's own help articles is a different act from tuning one model on every customer's support tickets and offering it to all of them, and the table shows how the usual fit changes.

What decides whether you can train on customer data?
ApproachWhat happens to customer dataUsual contract and privacy fit
Retrieval within one accountThe feature reads a customer's data to answer that customerOften within providing the service
Per-customer modelA model is tuned on one customer's data for that customer onlyUsually within the service if the DPA allows it
Pooled training across customersOne model learns from many customers' data and serves all of themNeeds express improvement or development rights and clear notice
Evaluation setsSamples are kept to test feature qualityTreated much like training; minimize and de-identify
Licensing to an outside developerRecords leave the company for someone else's modelsNeeds express rights or consent and fuller preparation

Which contract terms matter?#

The contract terms that matter are the license grant to customer data, any service improvement clause, the data processing agreement's limits on processing, and clauses individual customers negotiated. Read them in that order, and read the versions that applied when the data was collected.

  • License grant: does it allow use only to provide the service, or also to improve and develop it?
  • Improvement clause: does it reach new AI features, or only fixes and enhancements to existing functions?
  • Aggregated data clause: does it permit de-identified data to build models, and with what safeguards?
  • DPA instructions: is processing limited to the customer's documented instructions?
  • Negotiated terms: have enterprise customers added no-AI-training, data residency or deletion clauses?

How does your privacy role change the answer?#

Your privacy role changes the answer because a software company usually processes personal information inside customer accounts on each customer's behalf, as a processor or service provider. In that role, using the data for your own purposes, such as training a pooled model, needs a basis in the contract and in the privacy laws that may apply.

Some privacy laws give processors limited room to use data to build or improve their services, but the conditions vary and counsel should check each law that may apply. Data about your customers' own users, such as end-customer names and messages inside tickets, deserves the most care, because those people never dealt with you directly.

Minimize before training. Removing names, contact details and account identifiers from tickets and chats before they reach a training pipeline narrows the privacy question and limits what a model could repeat back.

What should you tell customers?#

Customers should be told which data trains which features, whether a model is shared across customers, how to opt out and what happens to their data if they leave. Short, specific language in the product, the privacy notice and the DPA is more defensible than a broad clause buried in the terms.

Enterprise buyers will ask regardless. A clear AI data-use page, a workspace-level switch and a standard answer for security questionnaires reduce the number of custom contract clauses you have to track later.

How do you build the pipeline so the answer holds up?#

A training pipeline holds up when the legal answer is enforced in the data itself, not just in a policy document. Engineering and counsel should agree the controls before the first model is trained, because adding them later means reconstructing which data went where.

The controls below are what an enterprise customer's security review or a future acquirer's diligence will usually ask to see. Each one turns a contract promise into something the company can demonstrate.

  • Account-level flags for opt-in status, no-AI-training clauses and residency limits, checked by every pipeline job.
  • Lineage records showing which accounts and date ranges fed each model version.
  • A de-identification step before training data leaves the production environment.
  • Retention rules for training copies that match the customer contract, including deletion after churn.
  • A documented way to exclude an account from future training runs when its status changes.

How is this different from licensing data to others?#

Licensing customer data to an outside AI developer is different because the records leave your service and support someone else's product. Customers gave you data to run and improve a service they use; an outside license serves neither purpose, so it usually needs express permission and fuller preparation.

That is why many software companies that do license data start with records they author themselves: resolved support cases with customer details removed, engineering issues, code reviews and release histories. Those records describe how your team works and carry a simpler rights chain.

Illustrative: a construction software company builds two features#

Illustrative: a fictional software company selling project management tools to general contractors wants a suggested-reply tool for its support team and an assistant that drafts RFI responses for its customers. Its terms permit customer data to be used to provide and improve the service, and several enterprise customers have no-AI-training clauses.

Counsel approves the support reply tool, trained on the company's own agent replies and de-identified tickets, because it is internal and within improvement rights. The customer-facing RFI assistant launches with retrieval inside each account only, and pooled training applies only to customers who switch on a new opt-in setting. Accounts with no-AI-training clauses carry a flag in the data warehouse that excludes them from every pipeline.

A separate request from a model developer to license the ticket archive goes through a different review and is limited to company-authored resolution records.

Where SourceX fits#

SourceX works on the outside-licensing question, not on internal feature training. When a software company decides to license records to model developers, SourceX runs the SourceX five-step transaction with the company approving each step, and documents the result in a SourceX Evidence Packet.

Much of the groundwork overlaps with internal AI governance: a dated map of contract versions, a list of customers with restrictive clauses and a record of how data was prepared. A company that builds that map once can use it for both decisions.

Frequently asked questions

Do we need consent from each customer to train our own features?

Not always. If your contracts already grant improvement or development rights that cover the feature, and your notices describe it, separate consent may not be needed. Where terms are narrow, or customers negotiated restrictions, an opt-in or amendment is the safer route. Counsel should confirm for your contracts.

Do we need to tell our customers' end users, or only our customers?

Your contract is usually with the customer, and the customer gives notice to its own users, but your DPA may require you to help. If a feature trains on data about end users, check that customers' notices could reasonably cover it, and offer wording they can reuse in their own policies.

Does training on de-identified tickets avoid the contract question?

Not entirely. De-identification helps with privacy, but the contract may still limit use of customer data and its derivatives. Some agreements expressly permit de-identified or aggregated use; others are silent, which leaves room for disagreement with customers.

What changes if our features use an outside model provider?

The provider becomes part of your data flow. Confirm that it does not retain or train on the data your feature sends, list it as a subprocessor where your DPA requires, and make sure customer disclosures describe that flow accurately.

What is the downside if we get this wrong?

The downsides include contract claims from customers, regulatory scrutiny of unfair or deceptive practices, and lost enterprise deals that require clean AI terms. Retroactive terms changes and ignored opt-outs are common sources of complaints, so avoid both.

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify