Skip to content

AI uses for records

Is your software vendor training AI on your business records?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

Your software vendor may be allowed to train AI on your business records if its terms permit using customer data to improve services, build models or create aggregated data. Read the data license, service improvement, aggregated data and AI feature clauses together. A contract that never mentions training does not necessarily forbid it, so ask in writing.

Key takeaways

  • Training rights usually come from broad clauses about improving services or using aggregated data, not from a clause titled AI.
  • Published vendor promises are often scoped to one kind of model, plan or setting, so read them in full.
  • AI feature addenda and admin settings can change the answer for the same vendor under the same plan.
  • Terms can change after signing, so keep dated copies and check which version applied when your records were created.
  • Vendor training rights affect your confidentiality promises and any later license of your own records.

Can a software vendor train AI on your records?#

A software vendor can train AI on your records only if its contract with you grants that right, but the grant is often indirect. Few business terms say in plain words that the vendor trains models on your support tickets. The right more often sits inside clauses about operating and improving the service, using aggregated or de-identified data, or providing AI features.

Vendors take different approaches. Some exclude customer content from model training by default, some offer an opt-out, and some treat training as part of an AI feature you switch on. The same vendor can also treat a free plan, a standard plan and a negotiated enterprise agreement differently, so the answer belongs to your contract, not to the brand.

Start the review with vendors that hold free text, because free text is what language models learn from most readily: help desk and chat platforms, CRM notes and call logs, meeting notetakers, document and wiki tools, and code hosting. Systems that hold mainly structured transactions, such as payroll or accounting, still matter for privacy but are usually a lower priority for this question.

The clause checklist: what to read and how to read it#

The clause checklist below covers the provisions that, together, decide whether a vendor may learn from your records. Read each one in the master agreement, the data processing addendum, any AI-specific terms and the order form, then note which document wins when they conflict.

The clause checklist: what to read and how to read it
ClausePlain-language readingWhat to ask the vendor
License to customer dataThe rights you grant over your content; a grant limited to providing the services is narrow, a long list of purposes is broaderIs model training or model improvement among the purposes?
Service improvementPermission to use your data to improve the product, which can be read to include trainingDoes improve include training or fine-tuning any model?
Aggregated or de-identified dataData the vendor derives from yours and may treat as its ownCan derived data include text from our records, and who decides it is de-identified?
Usage data and telemetryRecords of how you use the product, often owned by the vendorDo prompts, AI outputs or our edits count as usage data?
AI features addendumSeparate terms for assistants, summaries or reply draftingIs the feature on by default, and does using it feed training?
SubprocessorsThird parties, including model providers, that handle your dataWhich model providers receive our content, and may they retain it?
Changes to termsHow the vendor may update terms after signingDo changes apply to data collected before the change?

What do vendors' published AI training positions look like?#

Vendors' published AI training positions differ in scope: some exclude customer content entirely, some exclude only certain kinds of model, and some depend on the plan, the contract or an admin setting. The examples below come from each vendor's own terms or help pages, or from press reports of them, and show why a headline promise needs reading in full.

Treat these as examples of drafting patterns, not as a current reading of any contract. Terms change, and the version that governs your records is the one attached to your plan and order form for the period the records were created.

What do vendors' published AI training positions look like?
VendorPublished positionWhat it shows for your review
SlackIn May 2024, users objected that its privacy principles let customer data train its machine-learning models unless an organization emailed to opt out; Slack said it does not train its generative AI language models on customer dataA promise about one kind of model can leave other models, such as search or recommendation models, outside it
HubSpotIts knowledge base describes an AI model training switch that controls whether HubSpot uses an account's customer data to train its AI modelsThe answer can sit in an admin setting, so record who set it and when
GitHubIts Terms of Service grant GitHub a license to use AI-feature inputs and outputs to train AI models, with an opt-out in settings; customers under a GitHub Customer Agreement or volume licensing agreement are excludedThe same product can carry different training rights depending on the agreement you signed
AsanaIts product terms say it does not use customer data to train the generative AI models behind Asana AI; its privacy statement says it uses metadata about a domain's use to train machine learning models when AI features are enabledCustomer content and usage metadata are often treated differently
ProcoreIts AI FAQ says customer data is not used to train Microsoft Azure OpenAI or other Microsoft products, and that Procore's internal models may use data from customer inputs and outputs to improve Procore AI accuracyA promise about a third-party model provider does not settle what the vendor's own models may learn
QuickBooks Desktop (2022 US license)Grants Intuit a non-exclusive license to host and use content, and says Intuit may use data to improve the software and develop new products under its privacy statementOlder improvement and product development wording can be broad and predates most AI addenda

Why does vendor training matter beyond privacy?#

Vendor training matters beyond privacy because it touches three obligations the company already carries. The first is confidentiality: your contracts with customers, partners and employees may restrict who handles their information and for what purpose, and a vendor training on it can sit awkwardly with those promises.

The second obligation is protecting trade secrets. The federal definition in 18 U.S.C. 1839(3) treats information as a trade secret only if the owner has taken reasonable measures to keep it secret. Pricing logic, engineering discussions and internal playbooks shared with a vendor under an unexamined training grant may make that showing harder.

The third is your own ability to license. A buyer licensing your records will ask whether anyone else holds rights in them or has already trained on them. If a vendor has, the answer shapes the warranties you can give and may narrow what is worth licensing.

How do vendor terms change after you sign?#

Vendor terms often change after signing, because many online agreements let the vendor update policies by posting a new version or sending notice. Negotiated enterprise agreements may lock key terms, while click-through plans usually follow whatever version is current. Changes can also be reversed: in July 2025, after user objections, WeTransfer revised planned terms to remove language about using uploaded content to improve machine learning models.

For records, the version in force when data was created can matter as much as today's terms. Keep dated copies of the terms, the data processing addendum and any AI addendum at each renewal. Where a vendor adds AI features mid-contract, record the date and the setting you chose; that history answers questions from customers, auditors and future buyers.

Retroactive changes also draw regulators' attention. In a February 2024 blog post, FTC staff warned that a company adopting more permissive data practices, such as using consumers' data for AI training, and disclosing them only through a surreptitious, retroactive change to its terms or privacy policy may be engaging in unfair or deceptive practices. The warning addresses consumers, and how far it reaches business customers is less settled. If a vendor expands its data rights, ask in writing whether the change reaches data collected earlier, whether you can opt out without losing the product, and whether opting out applies to every workspace on your account.

What can you do when the answer is yes?#

When a vendor's terms do allow training, the company still has options, and most are administrative rather than legal. Start with settings, then the contract, then the vendor relationship itself.

  • Check admin settings for AI training, model improvement or data sharing controls, and record the date you changed them.
  • Ask for written confirmation of what has already been used, and whether opting out affects past data.
  • Negotiate an AI addendum or data processing terms that exclude your content from training.
  • Ask counsel whether privacy laws such as the CCPA or GDPR limit how the vendor, acting as your service provider or processor, may use personal information.
  • Limit which teams paste sensitive material into AI features until the review is finished.
  • Add the vendor's position to your contract register and your data inventory.
  • Plan a migration if the vendor will not move and the records are sensitive.

Illustrative: a manufacturer reviews its vendors at renewal#

Illustrative: a fictional precision machining company runs a cloud ERP, a cloud quality management system, a help desk for distributor questions and a file-sharing platform that holds customer drawings. Its general counsel starts the review after the ERP vendor announces a new AI assistant.

Working through the clause checklist, she finds four different positions. The ERP's AI addendum excludes customer data from model training, but only if an admin accepts the addendum. The QMS vendor may use aggregated quality data to build industry benchmarks. The help desk's reply drafting feature is on by default, and its terms allow service improvement. The file platform limits use to providing the service, which matters because customer drawings are covered by confidentiality clauses in purchase agreements.

The company accepts the ERP addendum, asks the QMS vendor in writing how aggregated data is de-identified, switches off help desk drafting until the terms are clarified, and files dated copies of every document. When it later considers licensing nonconformance records, it can answer a buyer's question about prior vendor use on one page and keeps customer drawings out of scope.

How SourceX approaches vendor terms#

SourceX reviews vendor terms in the Rights step of the SourceX five-step transaction with two questions in mind: whether the supplier may export the records for licensing, and whether any vendor already holds rights that overlap with a proposed license. The findings shape scope before any data moves.

The outcome is written into the SourceX Evidence Packet, in its licensing rights and permitted use entries, so the supplier and the buyer read the same position. Applicable laws and contract terms are assessed deal by deal with the supplier's counsel. SourceX's dataset rights, including any training use, are set out in the signed supplier agreement, and the supplier approves every step.

Frequently asked questions

Does opting out remove our data from models already trained?

Usually not on its own. Opting out generally stops future use, while removing the influence of data from a trained model is technically hard and rarely promised. Ask the vendor in writing what the opt-out covers, whether past data leaves its training sets, and what happens to derived or aggregated data.

Is using an AI feature the same as the vendor training on our data?

No. An AI feature can process your records to produce a summary or a draft without the vendor training a model on them. The terms decide whether inputs and outputs are retained or used for improvement, so read the feature's own terms rather than assuming either way.

Who in the company should own this review?

The general counsel or privacy lead usually owns the reading, with IT or each system owner confirming settings and plan tiers. Finance helps by listing every active subscription, because tools bought on a company card are often the ones on the broadest terms.

Does de-identified data derived from our records still belong to us?

Many contracts let the vendor own data it has aggregated or de-identified, even when it started as yours. Whether that data is truly de-identified depends on the method and the law that applies. Ask how the vendor de-identifies free text and whether text from your records can appear in derived data.

Can a vendor's training rights stop us from licensing our own records?

Usually they do not, because a vendor's rights are normally non-exclusive. They still matter: a buyer may value records less if they have already been used for training elsewhere, and any license you sign must describe prior grants accurately.

Sources

  • TechCrunch reported on May 17, 2024 that Slack's privacy principles allowed customer data to be used to train Slack's machine-learning models unless an organization emailed to opt out, and that Slack said it does not use customer data to train its generative AI large language models. Source
  • HubSpot's knowledge base says customers can use an AI model training switch to control whether HubSpot uses their account's customer data to train its AI models. Source
  • GitHub's Terms of Service (Section J) grant GitHub a license to use AI-feature inputs and outputs to train AI models, with an opt-out in settings; customers under a GitHub Customer Agreement or volume licensing agreement are excluded. Source
  • Asana's Product-Specific Terms state that Asana does not use or permit third-party AI partners to use Customer Data to train generative AI models used to provide Asana AI. Source
  • Asana's Privacy Statement says that when features powered by Asana AI are enabled in a domain, Asana uses metadata related to that domain's use of Asana to train machine learning models. Source
  • Procore's AI FAQ says no customer data is used to train Microsoft Azure OpenAI or other Microsoft products, and that Procore's internal models may use data from customer inputs and outputs to improve Procore AI accuracy. Source
  • Intuit's 2022 US QuickBooks Desktop and Payroll license grants Intuit a non-exclusive license to host and use Content and says Intuit may use data to improve the Software and develop new products or services. Source
  • In July 2025, after user backlash, WeTransfer revised updated Terms of Service to remove language referring to using uploaded content to improve machine learning models. Source
  • On February 13, 2024, FTC staff warned that adopting more permissive data practices, such as using consumers' data for AI training, through a surreptitious, retroactive change to terms or a privacy policy may be unfair or deceptive. Source
  • Under 18 U.S.C. 1839(3), information qualifies as a trade secret only if the owner has taken reasonable measures to keep it secret and it derives independent economic value from not being generally known. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify