Software companies
No-AI-training clauses in customer contracts: what B2B companies are seeing
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
A no-AI-training clause in a customer contract bars the vendor from using that customer's data to train AI models, and B2B software companies now see it in MSA redlines, DPAs and AI riders. The working rule: any record containing a customer's data follows the strictest clause in that customer's paper, so company-owned records are the usual starting point.
Key takeaways
- No-AI-training language arrives through MSA redlines, DPAs, AI riders and security questionnaire answers attached to contracts.
- Most clauses fall into a few patterns: a flat prohibition, a third-party ban, a de-identified or aggregated carve-out, or use allowed only with written consent.
- The definitions of customer data, derived data and training decide how far a clause reaches.
- Company-owned engineering and operating records usually sit outside these clauses once embedded customer content is removed.
- Tag each customer account with its clause pattern so exports can filter by it.
Where are no-AI-training clauses showing up?#
No-AI-training clauses are showing up wherever enterprise customers negotiate paper with their software vendors: MSA redlines, data processing agreements, standalone AI riders and security addenda. Customers add them because many vendors now seek contract room to improve models with customer data, and procurement teams want that room closed or tightly controlled.
The clause rarely arrives alone. A typical enterprise renewal now carries an AI section in the security questionnaire, a sub-processor list that names model providers, and a rider limiting what the vendor and its sub-processors may do with customer content. When the questionnaire is attached as an exhibit, its answers can become contract terms too.
Vendor-side commitments show the same pressure from the other direction. Zoom's online Terms of Service (Section 10.2) state that it does not use audio, video, chat, screen sharing, attachments or other communications-like customer content to train Zoom or third-party AI models. Slack drew public backlash in May 2024 when its privacy principles were read to allow training on customer data unless an organization opted out by email. Procurement teams point to episodes like these when they ask for a clause of their own.
- Master service agreement: confidentiality, customer data use and license-back sections, often rewritten in redlines.
- Data processing agreement: processing instructions, sub-processor terms and any carve-out for aggregated or de-identified data.
- AI rider or addendum: a separate exhibit covering AI features, model providers, training and outputs.
- Order forms and statements of work: one-off restrictions negotiated for a single deal.
- Security questionnaire answers and trust center statements incorporated by reference.
What do the common clause patterns say?#
The common clause patterns sit on a short spectrum, from a flat prohibition to use allowed with documented consent. The table paraphrases each pattern and reads it from the supplier's side: what it typically leaves open if the company later wants to license records to an AI developer.
Treat the table as a sorting aid, not a legal conclusion. The same label can hide very different drafting, and a clause that looks permissive can be narrowed by a definition three pages earlier.
| Pattern | Typical wording (paraphrased) | Effect on records you could license |
|---|---|---|
| Flat prohibition | Vendor will not use customer data to train, fine-tune or improve any AI model | Records containing that customer's data stay out of scope, even after de-identification, unless the customer amends |
| Third-party ban | Vendor and sub-processors will not let any third party train on customer data | Licensing to an outside AI developer is the exact use prohibited; internal product features may still be allowed |
| De-identified or aggregated carve-out | Prohibition applies except to aggregated data that cannot identify the customer or any person | Statistics and benchmarks may fit; record-level tickets and threads usually do not |
| Improve-the-service carve-out | Vendor may use data to provide, maintain and improve the services | Purpose is tied to your own service, so a license to another company rarely fits |
| Consent or opt-in | No training use without the customer's prior written consent | Records from customers who sign a specific consent may be licensable within its terms |
| Derived data extension | Prohibition covers embeddings, models and other data derived from customer data | Derived artifacts carry the same restriction as the source records |
Which definitions decide how far the clause reaches?#
The definitions decide how far a no-AI-training clause reaches, often more than the operative sentence does. A ban on training with customer data means little until you know whether customer data includes support conversations, usage logs or bug reports that customers emailed to your engineers.
Read these terms in the version of the agreement that actually governs each customer, and note where the MSA and DPA define the same word differently. An order-of-precedence clause usually settles conflicts, and it often favors the DPA.
- Customer data or customer content: only what users upload, or also tickets, chats and files sent to support?
- Usage data, service data or telemetry: carved out of customer data and assigned to the vendor, or not?
- Derived data: are aggregates, embeddings and analytics treated as customer data?
- Training: does it cover fine-tuning, evaluation, benchmarking and retrieval indexes, or only building a model?
- AI model or machine learning: is the term broad enough to sweep in rules engines and simple classifiers?
- Third party: could a licensee, an affiliate or a future acquirer count as a permitted recipient?
- Survival: does the restriction continue after the agreement ends?
How do these clauses interact with the rest of your paper?#
A no-AI-training clause interacts with confidentiality terms, privacy statements and sub-processor commitments, so the clause alone rarely tells the whole story. Customer data is commonly defined as the customer's confidential information, which means a license to a third party can breach confidentiality even where the AI clause is silent.
Public statements matter as well. A trust center page or privacy notice saying the company never shares customer data for AI training is a commitment customers rely on, and a questionnaire answer saying the same can be pulled into a contract. Align these documents before any licensing conversation, and record which version each customer accepted.
| Document | What to look for | Common conflict |
|---|---|---|
| MSA | Confidentiality, data use grant, survival | Confidentiality bars a disclosure the AI clause seems to allow |
| DPA | Processing instructions, carve-outs, precedence | Processing is limited to the customer's documented instructions |
| AI rider | Training, outputs, model providers | The rider reaches further than the MSA it amends |
| Privacy notice and trust center | Public promises about AI and sharing | Website copy promises more than the contracts require |
| Questionnaire answers | Statements attached as exhibits | An old answer becomes a binding representation |
What records usually remain licensable?#
Records a company creates about its own work usually remain licensable even when every customer contract carries a no-training clause. For a B2B software company, that typically means Jira issues and engineering discussion, pull requests and code review threads, internal Confluence or Notion documentation, product decision records and release notes.
The hard part is embedded customer content. Engineering tickets quote customer bug reports, paste log lines with account identifiers and attach screenshots from customer environments. Under most definitions those fragments are customer data, so they are removed, or the record is excluded, before anything moves.
Support conversations sit in the middle. A Zendesk or Intercom thread is often customer data in full, so many companies consider only the internal side: macros, internal notes, escalation decisions and the linked engineering fix, with customer text removed.
Illustrative: an HR software company sorts accounts by clause#
Illustrative: a fictional HR software company sells to mid-market and enterprise employers. Its older click-through terms are silent on AI, its enterprise customers signed redlined MSAs, and many recent renewals added an AI rider with a flat training prohibition.
Counsel tags every Salesforce account with its clause pattern: silent, flat prohibition, carve-out or consent. Support mirrors the tag onto Zendesk organizations so any export can filter by it. Engineering history in Jira and GitHub is reviewed separately, and a script flags ticket text pasted from customer emails for removal.
The company scopes a first package to internal engineering records with customer fragments stripped, and defers support threads entirely. It then asks a handful of silent-contract customers whether they would sign a narrow consent addendum, rather than assuming that silence permits use.
How SourceX reviews AI clauses#
SourceX reviews customer contract patterns in the Rights step of the SourceX five-step transaction, after Supply and before Preparation, Approval and Delivery. The initial fit check uses metadata only, so counsel can describe contract categories and known restrictions without sharing agreements or records.
When a package proceeds, the conclusions go into the SourceX Evidence Packet under licensing rights and permitted use, next to provenance, the privacy record and release authorization. That mirrors provenance standards such as the Data and Trust Alliance Data Provenance Standards, whose Use metadata includes elements for license to use, intended data use and consent documentation location.
Frequently asked questions
Does a no-AI-training clause still bind us after the customer leaves?
Often it does. Many agreements say confidentiality and data-use restrictions survive termination, and retained backups or support history may still hold that customer's data. Read the survival clause and the deletion terms together. If the data should already have been deleted, it does not belong in a licensing package at all.
Can de-identification get records around a flat prohibition?
Usually not. A flat ban on using customer data for training typically applies whatever form the data takes, unless the contract defines de-identified data as falling outside customer data. Where the clause has an express de-identified or aggregated carve-out, the records must still meet that carve-out's definition and its stated purpose.
Should we ask customers to amend their contracts?
Sometimes. A short consent addendum that names the record types, the de-identification steps and the permitted recipients can create clear rights for future licensing. Ask selectively, explain what is and is not included, and expect some customers to decline. Record each refusal and respect it in the export filters.
Do these clauses cover AI features inside our own product?
That depends on the wording. Some clauses ban only training models that serve other customers or third parties, while allowing features that process a customer's data for that customer. Others ban any model training. Product and legal teams should read the clause against each AI feature before answering customer questions about it.
Should we offer a no-training commitment to win deals?
Many vendors do, and it can shorten procurement. Before offering one, decide whether it covers only customer data or company-owned records as well, and draft it narrowly. A promise that no data will ever be used for AI can later block licensing of your own engineering history.
Sources
- The Use group of the Data & Trust Alliance Data Provenance Standards includes elements for confidentiality classification, consent documentation location, privacy-enhancing technologies applied, allowed and excluded processing and storage geographies, license to use, intended data use, and copyright, patent and trademark status. Source
- Zoom's Terms of Service (Section 10.2) state that Zoom does not use audio, video, chat, screen sharing, attachments or other communications-like Customer Content to train Zoom or third-party AI models. Source
- TechCrunch reported on May 17, 2024 that Slack drew user backlash after its privacy principles were found to allow customer data to be used to train its machine-learning models unless an organization emailed Slack to opt out. Source
Related resources
- InsightWhat permitted uses should a code license allow: training, evaluation or RL environments?
- InsightMemorization and regurgitation clauses for licensed source code
- InsightLicense-back terms after a software carve-out, including AI training rights
- IndustrySoftware development agencies data
- IndustryFintech software data
- DataCode review records
See if your company qualifies
A short company assessment. No data uploads are needed.