Skip to content

Systems and records

Do your software vendors train AI on your data? How to check

By SourceX Editorial · Updated

Short answer

Some software vendors may use customer data to train or improve AI features, and the only reliable way to know is to read what governs your account: the master agreement, data processing addendum, AI-specific terms and admin settings. Ask five questions per vendor, starting with the systems that hold your support tickets, CRM history, chat and documents.

Key takeaways

  • A vendor's AI data use is usually spread across the master agreement, data processing addendum, AI terms, privacy notice and admin settings.
  • Customer content, usage data and aggregated data are treated differently in most contracts, so check each category separately.
  • An opt-out setting matters only if you know whether it is on, who can change it and whether it covers data already used.
  • Vendor AI terms can affect what you can later promise a licensee of your own records, including exclusivity and confidentiality.
  • Re-check at every renewal and whenever a vendor announces a new AI feature.

Why vendor AI terms are hard to find#

Vendor AI terms are hard to find because they rarely sit in one document. A typical SaaS account is governed by an order form, a master subscription agreement, a data processing addendum, product-specific or AI-specific terms, an acceptable use policy, a privacy notice and a set of admin settings, and AI language can appear in any of them.

Online terms are incorporated by reference and can change during your subscription. Many agreements let the vendor update online documents with notice, and an order of precedence clause decides which document wins when they conflict. A statement on a trust center page or in a blog post usually carries less weight than the agreement, so trace every answer back to a contract document.

Separate two kinds of AI language. Customer terms say what the vendor itself may do with your content. Developer or API terms say what outside apps connected to the vendor may do. Both have moved recently: HubSpot's developer changelog says its updated Developer Terms restrict using data accessed through HubSpot APIs to train or improve AI models, and Slack's API Terms bar providers of apps offered outside their own organization from training large language models on Slack API Data. A tool connected to your CRM or chat falls under the second kind.

The five-question checklist#

The five-question checklist turns a pile of documents into answers you can record per vendor. Ask the same questions of every system that holds business records, and write down the document and section each answer comes from.

Record unknowns as unknown. A vendor that cannot answer the scope or subprocessor question in writing is telling you something useful about its own controls.

  • Use: does the vendor use customer content to train, fine-tune or improve AI models, or only usage data and telemetry?
  • Scope: are models trained on your content used only for your account, or shared across the vendor's customers?
  • Control: is there an opt-in or opt-out, what is the default, who in your company can change it, and does it reach data already used?
  • Subprocessors: which outside model providers receive your content, and may they retain it or train on it?
  • Change: how will the vendor notify you of changes to AI terms, and can you object, opt out or terminate if you disagree?

Where AI terms usually live, by type of system#

The location of AI terms tends to follow the type of system, because vendors attach AI language to the features that use it. Start with the systems holding the richest operational history, since those carry the most confidential content.

Check the admin console as well as the paperwork. Many AI features are controlled by tenant or workspace settings, and the setting that applies to your account may differ from the default described in marketing material.

Where AI terms usually live, by type of system
System typeRecords at stakeWhere to look first
Help deskTickets, customer conversations, macros, satisfaction commentsAI agent or assistant add-on terms, data processing addendum, admin AI settings
CRMContacts, emails, call notes, deal historiesProduct AI terms, data usage settings, privacy notice
Team chatMessages, files, threadsAPI terms, AI feature terms, workspace admin settings
Docs and wikisInternal knowledge, specifications, playbooksAI assistant terms, workspace or tenant settings
Meeting and call toolsRecordings, transcripts, AI summariesRecording and summary settings, consent notices, product terms
Code hosting and issue trackingSource code, issues, code reviewsAI coding assistant terms, prompt and suggestion retention settings

Customer content, usage data and aggregated data are not the same#

Customer content, usage data and aggregated data are separate categories in most SaaS contracts, and AI rights often differ between them. A vendor may promise not to train on your content while reserving broad rights over how you use the product, or over data it has combined and de-identified.

Definitions decide the outcome. If usage data is defined to include anything you input, a narrow-looking clause can reach your content. Read the definitions section before the AI section.

Customer content, usage data and aggregated data are not the same
CategoryExamplesWhat to check
Customer contentTicket text, messages, documents, recordings, codeWhether any training or improvement right exists, and whether it needs your opt-in
Usage data and telemetryClicks, feature use, performance logsWhether usage data can include fragments of your content
Aggregated or de-identified dataBenchmarks, statistics, derived datasetsHow de-identification is defined and whether outputs could reveal your business
FeedbackSuggestions, ratings, bug reports you sendWhether feedback gives the vendor unlimited rights to what you submit

Why vendor AI use matters if you may license your own records#

Vendor AI use matters to a company that may license its records because it shapes what you can promise a licensee. If a help desk vendor has already used your ticket content to train shared models, an exclusive license to that history may be harder to offer, and a buyer may ask about it in diligence.

It also bears on duties to your own customers. Customer contracts may limit who processes their information and for what purpose, so a vendor's AI use is a confidentiality question as well as a licensing one. A vendor training on your content does not normally take ownership of it, but the license you granted still counts.

Software companies face the same question from the other side. If you sell SaaS, your own terms decide what you may do with customer content, usage data and aggregated data, and changing them generally requires proper notice and, for some uses, consent. Your own telemetry and internal engineering records are usually simpler to license than content your customers entered.

Illustrative: an engineering firm audits its vendors#

Illustrative: a fictional civil engineering firm runs project accounting in Deltek, document markup in Bluebeam, a cloud help desk for IT requests, a CRM for proposals and a meeting tool that records design reviews. Its managing principal asks the operations director to find out which vendors can use firm content for AI.

Using the five questions, the director builds a vendor register. Two vendors state they do not train on customer content. One reserves rights over aggregated data under a loose definition. The meeting tool has AI summaries switched on by default, with transcripts processed by an outside model provider.

The firm turns off AI summaries for client meetings, asks the meeting vendor to confirm in writing that transcripts are not retained for training, and files the answers. When it later considers licensing internal project review records, the register shows which archives are free of conflicting vendor rights.

How SourceX uses a vendor review#

SourceX uses a vendor review in the Rights step of the SourceX five-step transaction, after Supply and before Preparation. The review asks where each record family lives, which vendor terms govern exports and reuse, and whether any vendor right conflicts with the permitted use a buyer wants.

The findings go into the SourceX Evidence Packet under licensing rights and permitted use, next to provenance, the privacy record and release authorization. The initial fit check collects only descriptions such as system names, years of history and known restrictions, so no records are shared to start the review.

Frequently asked questions

If a vendor trained AI on our data, do we lose ownership of it?

Usually not. Training rights are normally granted as a license, and ownership stays with the customer under most SaaS agreements. The practical issue is that the license may already have been used, which can affect exclusivity and confidentiality. Read the ownership and license clauses together, and have counsel review anything unclear.

Does opting out remove our data from models already trained?

Often it does not. An opt-out commonly applies to future processing, and removing data from a trained model is technically difficult. Ask the vendor whether the setting covers past data, whether copies of your content are kept for training, and whether deletion on request is available.

Should we negotiate an AI addendum with key vendors?

For systems holding your most sensitive records, it is worth asking. Useful terms include no training on customer content without written opt-in, a named list of model subprocessors, notice before AI terms change and a right to terminate. Smaller contracts may only let you choose settings, so document those choices.

Who should own the vendor AI review?

Legal usually owns interpretation, IT or operations owns the settings, and each system owner confirms what records the tool holds. Keep the results in one register so renewals, new AI features and any licensing review all draw on the same answers rather than repeating the work.

How often should we re-check vendor AI terms?

At every renewal, whenever a vendor announces a new AI feature, and whenever a notice of updated terms arrives. AI terms change more often than most contract language, and a setting you turned off may come back with a new product. Assign a named owner for each critical system so notices are actually read.

Sources

  • HubSpot's developer changelog says its updated Developer Terms restrict using data accessed through HubSpot APIs to train, fine-tune or improve AI or machine learning models, with a carve-out for legitimate single-customer use cases. The same post says the terms now state that customer data belongs to the customer, not to HubSpot or developers. Source
  • Slack's API Terms of Service state that a provider of an application offered for use outside its own organization may not use API Data to train a large language model, and may not bulk export Slack message and file data except where an additional agreement expressly allows it. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify