Definitions and comparisons
Vendor AI training vs licensing your own data: who captures the value?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
When a software vendor trains AI on your data, it relies on rights in its own terms, usually without paying you or telling you which records it used. When you license your own records, your company sets the scope, the price and the deletion terms. Read five clauses in each major vendor contract to know which situation you are in.
Key takeaways
- Vendor AI training rights usually come from improvement, usage data or aggregated data clauses rather than a clause labeled AI.
- A vendor that trains on your records captures the value; a license you negotiate returns both value and control to your company.
- Aggregated or de-identified data clauses can let a vendor keep using derived data even after you opt out of other uses.
- Whether a vendor already trained on your records affects how unique and exclusive those records are to a later buyer.
- Review vendor terms at each renewal and whenever a vendor switches on AI features or updates its online terms.
What does it mean when a vendor trains AI on your data?#
A software vendor training AI on your data means the vendor uses records your company created in its product, such as help desk conversations, CRM notes or dispatch logs, to build or improve models. The permission comes from the vendor's contract, which may be spread across terms of service, a data processing addendum and separate AI terms.
The rights rarely sit in one obvious clause. They usually come from definitions that separate customer data from usage or service data, a license letting the vendor use data to provide and improve its services, and a clause letting it keep aggregated or de-identified data indefinitely.
This is general information, not legal advice. Vendor contracts vary widely, so have counsel read the agreements that actually govern your systems.
Vendor training vs licensing your own records#
Vendor training and licensing your own records differ on consent, compensation, control and deletion. In the first case the vendor writes the terms and decides how to use what it collects; in the second your company decides what leaves, on what terms and for how long.
The comparison also shows why this is a question about value, not only privacy. Records that show how your team resolves freight exceptions, diagnoses equipment faults or handles escalations are the same records AI developers look to license. If a vendor already has rights to use them for its own models, part of that value has moved to the vendor without any negotiation.
| Question | Vendor trains on your data | You license your own records |
|---|---|---|
| Consent | Given by accepting terms, sometimes by default | Given in a signed license for a defined scope |
| Compensation | Usually none beyond the service you already pay for | License fees or other negotiated consideration |
| Control over scope | The vendor decides which records and which models | You choose record families, date range and allowed uses |
| Who benefits | The vendor's product and, indirectly, its other customers | Your company, with the buyer receiving defined rights |
| Deletion | Depends on vendor policy; trained models may not forget | Deletion and return terms negotiated in the license |
| Record of use | Rarely disclosed in detail | Documented, including what was removed before delivery |
| Personal data | Governed by the vendor's privacy terms | Personal and confidential details removed before release |
Where vendor AI rights usually hide in a contract#
Vendor AI rights usually hide in documents customers skim: online terms incorporated by reference, product-specific terms for AI features and the privacy policy. Check each of these, not only the signed order form.
Admins often switch on AI features in a help desk or CRM without legal review. Ask IT for a list of AI features enabled in each major system, because enabling a feature can bring in a separate set of terms.
Public examples show how much the governing document and your plan matter. On August 7, 2023, after criticism of earlier changes, Zoom added a sentence to Section 10.4 of its terms saying it would not use audio, video or chat Customer Content to train its AI models without consent. GitHub's Terms of Service, in Section J on AI features, grant GitHub a license to use AI-feature inputs and outputs to train models, with an opt-out in account settings, and exclude customers under a GitHub Customer Agreement or volume licensing agreement from that training license. Two companies on the same tool can sit under quite different terms.
- Online terms of service linked from the order form, which the vendor may update over time.
- The data processing addendum, which often defines what counts as customer personal data.
- AI feature terms or an AI addendum that applies once someone enables an assistant or copilot.
- Acceptable use and API terms, which can limit bulk export of your own records.
- Notice provisions explaining how changed terms take effect, sometimes through continued use.
The five-clause check#
The five-clause check is a short read of each vendor agreement that tells you whether the vendor can train on your records and whether you can still license them. Run it first for the systems holding your deepest history, such as the help desk, CRM, ERP or code host.
Record each answer in a simple register listing the system, the contract version and the clause text. The same register becomes evidence in any later rights review and saves counsel from reading the contracts twice.
| Clause | What to look for | Question for the vendor |
|---|---|---|
| Data definitions | How customer data, usage data and service data are defined, and who owns each | Are our support conversations and internal notes our customer data? |
| Improvement license | Rights to use data to improve or develop services, including new products | Does improving services include training models? |
| Aggregated or de-identified data | Rights to keep and use derived data once it is de-identified | Can de-identified records from our account train your models? |
| AI training and opt-out | Whether training is on by default and how to opt out | Does our opt-out cover models already trained? |
| Changes and exit | How terms change, and what is exported or deleted when you leave | Can we export full history, including attachments, before termination? |
Does vendor training limit your ability to license?#
Vendor training rarely transfers ownership, but it can still limit your ability to license in three ways. Some terms restrict exporting records in bulk or through the API, which matters when you prepare a delivery. A buyer will ask whether the same records already trained someone else's model, which affects how unique and exclusive they are. And a vendor may hold derived data you cannot recall.
Export limits deserve a close read. Slack's API Terms of Service, for example, say an application offered for use outside its own organization may not use API data to train a large language model or bulk export Slack message and file data unless an additional agreement expressly allows it. That clause governs third-party apps rather than your own admin export, but it shows why the export route should be confirmed before a delivery is promised.
None of this usually ends a licensing project, but it changes the scope. Records created before an AI training clause took effect, or in a system with no such clause, are often the cleaner part of the archive. Your rights review should note which systems carry vendor training rights and from what date.
Illustrative: a freight brokerage reads its help desk terms#
Illustrative: a fictional regional freight brokerage runs loads through a transportation management system and handles shipper and carrier questions in a cloud help desk. While scoping its records for licensing, the COO asks counsel to run the five-clause check on both vendors.
The TMS terms define load records as customer data and give the vendor no training rights. The help desk's online terms include an improvement license and an aggregated data clause, and the support team had switched on an AI reply assistant, which brought in AI feature terms with training enabled by default.
The brokerage turns the setting off, asks the vendor to confirm in writing whether the opt-out covers past data, and exports its full ticket history with attachments. In the licensing scope it discloses the help desk's training rights to buyers and leads with TMS records, which carry a cleaner rights history.
How SourceX treats vendor terms#
SourceX reviews vendor terms during the Rights step of the SourceX five-step transaction, alongside customer contracts and employee notices. The findings feed the licensing rights and permitted use sections of the SourceX Evidence Packet, so a buyer can see which system each record family came from and what that vendor's terms allow.
The fit check itself needs only metadata, such as system names and years of history, so a company can learn whether its records merit a review before asking counsel to read every vendor contract.
Frequently asked questions
Can we opt out after a vendor has already trained on our records?
Usually you can opt out going forward, but models already trained are hard to unwind. Ask the vendor in writing whether the opt-out applies to past data, whether derived data is deleted and whether any model trained on it is retired. Keep the reply with the contract file.
Does an AI assistant feature mean the vendor is training on our data?
Not always. Many assistants use your records only at the moment of a request, to retrieve context, without adding them to a model. Others keep inputs to improve the service. The AI feature terms and vendor documentation say which applies, so read them before switching a feature on.
Is de-identified data still ours?
That depends on the contract. Many vendor agreements say de-identified or aggregated data belongs to the vendor or may be used freely. Privacy laws may also treat de-identified data differently from personal data. If the clause is broad, negotiate limits at renewal, especially on model training.
Who in the company should own vendor AI reviews?
Usually the general counsel or contracts lead, working with the COO and IT. Legal reads the terms, IT knows which features are enabled and which systems hold the history, and operations decides which records matter most. Repeat the review at each renewal.
Should we renegotiate vendor terms before licensing our own records?
It often helps. Ask for confirmation that your records are customer data you own, a narrower aggregated data clause, training off by default and a full export right at exit. Cleaner vendor terms make your records easier to license and easier to explain to a buyer.
Sources
- On August 7, 2023, after backlash over March 2023 changes to its terms, Zoom added to Section 10.4 of its Terms of Service a sentence saying it will not use audio, video or chat Customer Content to train its AI models without consent. Source
- GitHub's Terms of Service (Section J, AI features) grant GitHub a license to use AI-feature inputs and outputs to train AI models, which users can opt out of in account settings; customers under a GitHub Customer Agreement or volume licensing agreement are excluded. Source
- Slack's API Terms of Service state that a provider of an application offered outside its own organization may not use API Data to train a large language model, and may not bulk export Slack message and file data except where an additional agreement expressly allows it. Source
Related resources
- InsightWhat permitted uses should a code license allow: training, evaluation or RL environments?
- QuestionShould companies sell or license their data?
- QuestionDo AI labs buy code?
- InsightOpt-in vs opt-out for AI training in B2B SaaS contracts
- InsightCan a distributor license its pricing and quote history?
- SolutionWhat is AI evaluation data?
See if your company qualifies
A short company assessment. No data uploads are needed.