Skip to content

Privacy and preparation

Customers asking whether you train AI on their data: answering security questionnaires

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

When a security questionnaire asks whether you use customer data to train AI, answer by data category: what you do with content customers store, which AI vendors process it and on what terms, and how your own operational records are treated. Commit only to what your contracts, systems and vendor terms can prove today.

Key takeaways

  • Questionnaire answers often become contract evidence, so each statement should match your DPA, trust center and actual systems.
  • Separate customer content, account and usage data, AI subprocessors and your own operational records before writing a single answer.
  • A firm no-training commitment for customer content is compatible with licensing your own engineering and internal records.
  • Absolute statements such as never using any data for AI are the answers most likely to become false later.

What are customers really asking?#

Customers asking whether you train AI on their data want to know three things: whether their content could end up inside a model, whether a model could expose it to someone else, and which outside AI vendors touch it. The question arrives in several wordings, and each deserves a precise answer rather than a pasted paragraph.

Procurement teams compare answers across vendors and keep them on file. If a dispute or incident arises later, the questionnaire is one of the first documents both sides reread.

  • Do you use customer data to train or fine-tune AI or machine learning models?
  • Do you use third-party AI services that process our data, and do those providers retain or train on it?
  • Can our data be used to improve models that serve other customers?
  • Can we switch off AI features or AI processing for our account?
  • Will you notify us before changing how our data is used for AI?

Sort your data into categories first#

Sorting data into categories is the step that makes an honest answer possible, because a SaaS company holds very different records under very different rights. A single yes or no cannot describe all of them accurately.

The last row is the one most vendors forget. Engineering history and internal documentation are company records, and an honest answer can say so, as long as customer content, secrets and customer names are kept out of anything you might license.

Sort your data into categories first
CategoryTypical examplesWhat you can usually commit to
Customer contentFiles, messages, tickets and records customers create or upload in your productNo training of your own or third-party models without the customer's written agreement
Account and usage dataLogins, feature usage events, performance metrics, billing recordsUse limited to running, securing and improving the service, as your agreement describes
AI features and subprocessorsSummaries, search and drafting assistants built on outside model APIsNamed providers, processing purpose, retention terms and whether providers may train
Support correspondenceTickets, chats and emails customer admins send your team, often with pasted exports or screenshotsA plain description of how it is used, with anything customers paste treated as customer content
Your own operational recordsEngineering issues, code reviews, internal documentation, your staff's process notesDescribed as company records that contain no customer content once reviewed

An answer template you can adapt#

An answer template works when each sentence maps to a control you can show a customer or an auditor. Adapt the lines below to your contracts and delete any line you cannot support.

Fill each bracket from a source document, not from memory. The retention term comes from the vendor contract, the settings location from the product, and the notice commitment from your master agreement.

  • Customer content: We do not use customer content to train or fine-tune AI models, ours or any third party's, unless the customer agrees in writing.
  • AI features: Our product offers [feature names], which send relevant customer content to [provider] for processing. Our agreement with [provider] prohibits training on that content and limits retention to [retention term].
  • Controls: Administrators can switch off AI features for their workspace in [settings location].
  • Usage data: We use account and usage data to operate, secure and improve the service, as described in our agreement and privacy notice.
  • Company records: Our own internal records, such as engineering issues and process documentation, are company records. Before any use outside the company, they are reviewed to remove customer content, customer names and secrets.
  • Changes: We will give customers notice before any change to these practices, in line with our agreement.

Statements to avoid unless you can prove them#

Broad statements are the riskiest answers because they tend to become false as your product and vendors change. Each one below has a narrower version that stays true.

Statements to avoid unless you can prove them
StatementWhy it causes troubleSafer approach
We never use any data for AI.Your engineers may use coding assistants and your helpdesk may have AI features switched on.Scope the commitment to customer content and name the AI features that exist.
All data is anonymized.Anonymized has a legal meaning, and free text rarely meets it without review.Describe the specific removal steps, or leave the claim out.
Our AI vendors never see your data.Any AI feature that processes content sends it somewhere.Name the provider, the purpose and the contractual training and retention limits.
This policy will never change.Products and laws change, and you will need room to adapt.Commit to advance notice, and to customer agreement where your contract requires it.

Where your own operational records fit#

Your own operational records sit outside a no-training commitment for customer content, which is why a SaaS company can make that commitment and still consider licensing its engineering history. The line holds only if the records you license are truly yours and free of customer material.

Engineering issues, code reviews, release notes and internal runbooks are usually company records. Support tickets sit closer to the line: replies your staff wrote belong to your support operation, but anything customers pasted into them is customer content. If tickets are ever in scope, your questionnaire answer must already allow for that, and preparation must strip customer content before anything leaves your systems.

Give one internal owner both the questionnaire library and any licensing program. Mismatched answers from sales and legal are a common cause of customer escalations.

Illustrative: a construction scheduling vendor answers a contractor#

Illustrative: a fictional construction scheduling SaaS company receives a questionnaire from a large general contractor. It asks whether customer project data trains AI models, which AI vendors process it, and whether the company plans to license data.

The company runs support in Intercom, engineering in Jira and GitHub, and a schedule-summary feature on an outside model API. Its answer commits to no training on customer project data, names the summary feature and its provider, states that the provider's terms bar training on submitted content, and explains that workspace admins can switch the feature off.

On licensing, the answer says the company may license its own engineering and internal process records, never customer project data, and that customers would be told before any change. The same wording goes into the trust center, so every later questionnaire draws on one approved text.

Keep answers accurate as things change#

Accurate answers depend on a list of review triggers, because the facts behind them change faster than annual questionnaire cycles. Version every approved answer and record who approved it, so a customer asking what changed gets a dated history.

  • A new AI feature ships, or an existing one moves to a different provider.
  • A model provider updates its retention or training terms.
  • The company starts or widens a data licensing program.
  • A major customer negotiates new DPA language on AI.
  • A state privacy law that may apply takes effect or is amended.

How SourceX treats customer commitments#

SourceX treats published customer commitments as binding inputs in the Rights step of the SourceX five-step transaction. Questionnaire answers, trust center statements and DPAs are reviewed alongside contracts, and any record family they cover is excluded unless the supplier holds written customer authorization.

The resulting permitted use is recorded in the SourceX Evidence Packet, so what a company tells its customers and what it licenses stay consistent.

Frequently asked questions

Should enterprise and self-serve customers get different answers?

The practice should be the same for both, but enterprise customers often negotiate stronger contract language, such as a written no-training clause or advance notice of new AI subprocessors. Keep one standard answer and track negotiated exceptions per account, so sales never promises something the product cannot honor.

Do AI coding assistants used by our engineers count?

They can if engineers paste customer data, logs or support exports into them. Most questionnaires focus on product data flows, but an accurate answer reflects internal tool policies too. Set a rule that customer content never goes into unapproved AI tools, and say so if asked.

What if a customer asks for a contractual no-training clause?

Many vendors agree to one for customer content because it matches their practice. Read the clause's definition of customer data carefully. A definition that sweeps in usage data, support correspondence or derived statistics can restrict more than you intended.

Can we say we train only on deidentified customer data?

Only if the deidentification meets the legal standard that may apply and your contracts permit the use. In many SaaS agreements the use limits apply to customer data whether or not names are removed, so the safer answer is often a plain no.

Who should own the answer?

A named owner, often the head of security or the general counsel, should approve the standard text, with product and engineering confirming the facts behind each sentence. Sales should draw from the approved library rather than writing new answers under deadline.

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify