Private equity and portfolios
Software acquisition due diligence: AI and data rights questions to add
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
Software acquisition due diligence should add six AI and data rights question groups to the usual technical review: AI features shipped, training on customer data, model vendors, data exports, prior data licenses and customer AI clauses. The rule is simple: count the target's records as an asset only where its contracts allow the use your plan assumes.
Key takeaways
- Standard technical diligence checks code, security and open source, but rarely asks what the target may do with customer records.
- Training on customer data is the question most likely to change price or structure, because the answer sits in contracts, not code.
- Model vendor settings decide whether customer content sent to an AI feature can be retained or reused by the vendor.
- Prior data licenses can quietly grant exclusivity or survival rights that limit what the group can do after closing.
- Ask for documents, not descriptions: signed contracts, vendor account settings, subprocessor lists and export test results.
Why standard software diligence misses AI and data rights#
Standard software diligence focuses on architecture, code quality, open source licenses, security and scalability, and it rarely asks what the target is allowed to do with the records its customers create. For a buy-and-hold software acquirer that gap matters, because the thesis often assumes the group can use support history, usage records and engineering history in new AI features.
The answers sit in places a code review never touches: the master subscription agreement template and its negotiated versions, the data processing addendum, model vendor terms and any side agreements the founder signed with partners. A short, separate question set gets those documents into the data room early.
The six question groups to add#
The six question groups cover what the target already does with AI and what it is permitted to do with its records. Each group has one core question and a document that answers it, so the deal team can track gaps instead of collecting narrative answers.
| Group | Core question | Document that answers it |
|---|---|---|
| AI features shipped | Which product features call a model, and what customer content do they send? | Feature inventory, architecture notes and release notes |
| Training on customer data | Has the company trained or tuned any model on customer content, and under which terms? | Customer agreements, data processing addendum and training logs |
| Model vendors | Which model providers process customer content, and may they retain or reuse it? | Vendor terms, account data settings and subprocessor list |
| Data exports | Can the company export its own issue, support and code history with links intact? | Export test results and system admin settings |
| Prior data licenses | Has anyone already received rights to the company's records? | Signed licenses, benchmarking and partner agreements |
| Customer AI clauses | Which customers restrict AI use of their data? | Contract abstract with clause references |
AI features and model vendors: what to ask#
AI features and model vendors should be reviewed together, because a feature's risk depends on where customer content goes after the user clicks. A support summarizer that sends ticket text to a hosted model raises different questions from a forecasting feature that runs on the target's own aggregated tables.
Retention and training settings often differ between a provider's consumer, team and enterprise plans, so ask for screenshots or exports of the actual account configuration rather than a link to the provider's public policy page.
- List every feature that calls a model, internal or third party, with the release in which it shipped.
- For each feature, name the customer content sent, such as ticket text, attachments, call transcripts or records from the customer's own database.
- Confirm the provider's data settings: whether inputs and outputs are retained and whether the provider may use them to improve its models.
- Check that the subprocessor list in the data processing addendum names each model provider, and whether customers were notified when it changed.
- Ask whether any enterprise customer signed an AI addendum or opted out of AI features, and how the product enforces that choice.
Reading customer contracts for training and AI clauses#
Customer contracts decide whether the target may train on customer records, and most software companies have several generations of terms in force at once. Early customers may sit on a short click-through, while later enterprise customers negotiated their own data language.
Ask the target for a contract abstract that lists each customer's governing terms version and any data or AI clause, with references. Then spot-check the abstract against the signed contracts for the largest accounts, since abstracts prepared for a sale tend to smooth over exceptions.
| Language found | What it may support | Follow-up |
|---|---|---|
| Silent on AI and aggregated data | Often read as limiting use to providing the service | Treat training as unsupported until counsel reviews |
| Aggregated or de-identified data clause | Analytics and product improvement, sometimes more | Check whether disclosure to third parties is allowed |
| Explicit no-training or no-AI clause | No training on that customer's content | Schedule the customer for exclusion |
| Opt-in AI program | Use limited to customers who opted in | Request the opt-in records and their dates |
| Customer owns data and outputs | Narrow vendor rights | Confirm any license back to the vendor |
Data exports and prior licenses#
Data exports matter because records the company cannot extract with their links intact are worth less, whether for internal AI or a later license. Ask the target to test an export of issue history from Jira or Linear, pull requests and review comments from GitHub or GitLab, and tickets from Zendesk or Intercom, and to confirm that issue keys and ticket references survive.
Prior licenses are the quieter risk. Founders sometimes sign benchmarking agreements with industry associations, data partnerships with channel partners or early data deals with AI developers. Read each one for exclusivity, field of use, term, survival after termination and whether the counterparty may keep models trained on delivered records.
Illustrative: a software holding company reviews a landscaping software target#
Illustrative: a fictional vertical software holding company is buying a scheduling and billing platform used by commercial landscaping firms. Standard diligence comes back clean, so the head of M&A adds the six question groups to the request list.
The contract abstract shows three generations of terms. The oldest customers accepted an aggregated data clause, the middle group's terms are silent, and a handful of large customers negotiated no-AI addenda. The support summarizer sends ticket text to a hosted model under a team plan with limited retention, and the founder signed a non-exclusive benchmarking agreement with a trade association covering aggregated job data.
The acquirer adds a representation listing all prior data grants, a schedule of customers with AI restrictions and a pre-closing covenant to move the summarizer to an enterprise account with stricter data settings. After closing, only records from customers whose terms support the use enter any AI or licensing plan.
Red flags that should change price or structure#
Red flags in AI and data rights diligence rarely end a deal, but they should change the price, the representations or the post-closing plan. Raise each one with deal counsel as soon as it appears.
- The target trained a model on customer content under terms that are silent on that use.
- A prior data license grants exclusivity, survives termination or lets the counterparty keep trained models.
- No one can produce the model provider's actual account settings.
- The customer contract abstract does not exist or cannot be reconciled to signed agreements.
- Issue, support or code history was lost in a migration and cannot be exported with links.
How SourceX uses diligence answers after closing#
SourceX picks up where diligence leaves off. When a software company in a group decides to license engineering or support records, the SourceX five-step transaction starts with Supply and Rights, and the contract abstract and prior-license review from diligence shorten both steps.
Each package that proceeds gets a SourceX Evidence Packet covering provenance, licensing rights, permitted use, the privacy record and release authorization, so the customers and prior grants excluded at signing stay excluded at delivery.
Frequently asked questions
Should AI questions sit in technical or legal diligence?
Both, with one owner. Technical diligence answers what the product does with models and where content flows. Legal diligence answers what contracts permit. Assign one person to reconcile the two, because the most common gap is a feature that technical reviewers approve and contracts do not support.
How far back should the contract review go?
Back to the oldest terms still governing an active customer, plus any former customer whose records the company still holds. Some click-through terms let the vendor update them, but whether an update reaches records created under earlier terms is a question for counsel, so record which version governed each customer and when.
What if the target is too small to have an AI policy?
Many smaller software companies have none. Ask instead for a list of AI tools staff use, the accounts they use them under and what content goes in. That answers the practical question and gives the integration team a starting point for the group policy.
Do these questions apply to a carve-out of a product line?
Yes, with an extra step. In a carve-out, confirm which entity holds the customer contracts, system accounts and historical records for the product, and whether the transition services agreement lets the buyer export full history before the seller's systems are separated.
Related resources
- QuestionCan SaaS data be licensed?
- QuestionDo AI labs buy legal documents?
- InsightHow do I de-identify contracts and legal documents for AI training?
- InsightIndemnification in data licenses: who covers which claims
- InsightCode copyright and AI training: where the law stands in 2026
- IndustryConstruction data
See if your company qualifies
A short company assessment. No data uploads are needed.