Skip to content

Private equity and portfolios

AI due diligence checklist for private equity deal teams

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

An AI due diligence checklist for private equity covers four groups: the target's exposure to AI, the AI it already uses, its rights to use records for AI training, and its governance. Run it beside commercial, financial and legal diligence, and close with 12 written questions the target answers with documents, not management slides.

Key takeaways

  • AI exposure asks whether AI will compress the target's revenue, not whether the target uses AI.
  • Unapproved AI tools that receive customer content are a frequent finding in the AI in use group.
  • Data rights for AI training depend on contracts and notices, so the answer sits with counsel, not engineers.
  • Governance should be proportional: a named owner, a tool inventory and a short policy matter more than a long framework.
  • Every answer should come with a document the deal team can verify.

What AI diligence adds to a standard deal process#

AI diligence adds three questions that commercial, financial and legal diligence do not ask on their own: will AI change the target's market, is the target's own AI use creating liabilities, and do its records give it an advantage it can legally use. Each needs evidence from a different part of the company.

Many deal teams fold AI into technology diligence and get back a list of tools. A checklist organized by the four groups below produces findings the investment committee can act on, such as a revenue line at risk, an exposure in vendor terms or a set of rights-cleared records.

Group 1: AI exposure#

AI exposure measures how much of the target's revenue depends on work that AI tools may let customers do faster, cheaper or themselves. It belongs in commercial diligence, but the AI checklist makes sure someone asks the question directly.

Exposure findings belong in the base case, not an appendix. If a meaningful share of revenue comes from work customers could start doing with general AI tools, the model should show what happens to that line, and the value creation plan should say how the company responds.

Group 1: AI exposure
SignalHigher exposureLower exposure
Type of work soldRoutine document, research or content workOn-site physical work or licensed sign-off
Pricing modelBilled by hours or seats for repeatable tasksPriced by job, outcome or asset uptime
Customer switchingEasy to replace with a self-serve toolEmbedded in customer operations
Proprietary recordsFew records beyond invoicesYears of linked operating history
Competitor behaviorRivals already selling AI-led versionsLittle AI activity in the niche

Group 2: AI in use#

The AI in use group inventories every model the target relies on, in its back office and in any product it sells, and follows customer content into each one. Unapproved tools used by staff are the usual gap, because they rarely appear in IT's system list or the software budget.

Account type is the detail that matters most. The same tool can carry very different data terms on a personal plan and a business plan, so ask which accounts staff actually log in with, not which tool the company intended them to use.

For a target that sells software, product AI features deserve their own review of contracts and provider settings. For services, distribution and trades businesses, the bigger exposure is usually AI switched on inside systems the company already runs.

  • Embedded features: AI switched on inside the help desk, CRM, field service platform or ERP, and which records those features read.
  • Staff tools: drafting, coding and chat assistants, with the account type each person actually uses.
  • Meeting notetakers and call recorders: which customer calls they capture, and whether callers are told.
  • Product features, if the target sells software: the provider behind each one and the customer content it sends.
  • Provider data settings: retention, training on inputs and data location for each account.
  • Human review and incidents: where a person checks outputs, and any wrong outputs, data exposures or complaints.

Group 3: Data rights for AI training#

Data rights for AI training decide whether the target can use its records for its own models or license them to AI developers. The answers live in customer agreements, vendor terms, privacy notices, employee policies and any prior data licenses, so legal diligence should own this group.

Ask the target to describe each major record family by where it came from, how it was handled and what use is permitted. Industry groups have begun to standardize that description: the Data & Trust Alliance's Data Provenance Standards organize dataset metadata into three groups, Source, Provenance and Use, which makes a practical structure for the request even if the target has never used it.

Look for three findings in particular: customer contracts that prohibit AI use, records the target holds but does not own, such as client deliverables, and prior licenses that granted exclusivity to someone else.

Group 4: Governance#

Governance diligence checks whether someone at the target owns AI decisions and whether the company can show what it did. For a company with a few hundred employees, a named owner, a current tool inventory and a short written policy are a reasonable bar.

Check board minutes for AI discussions, vendor approval records and how the company answers customer questionnaires about AI. State privacy laws, sector rules or the EU AI Act may apply depending on where the target operates and sells, and counsel assesses which ones matter deal by deal.

A small target with no policy is not a red flag on its own. What matters is whether management can name its AI tools, say what content goes into them and show a willingness to adopt the platform's standards after closing.

The 12 questions to send the target#

The 12 questions below close the checklist. Send them in writing alongside the main request list and ask for a document behind every answer, so findings rest on contracts and settings rather than recollection.

  • Which product features use AI models, and which providers power them?
  • What customer content does each feature send to a provider?
  • Which AI tools do employees use for work, and under which accounts?
  • What retention and training settings apply to each provider account?
  • Has any model been trained, fine-tuned or grounded on customer, employee or operating records?
  • Which customer contracts restrict AI use, training or disclosure of data?
  • Has the company licensed, shared or sold any records to a third party?
  • Which records does the company hold that belong to clients or partners?
  • Which systems hold the main operating records, and how far back can they be exported?
  • Who owns AI decisions, and is there a written AI policy?
  • Have there been any AI-related incidents, complaints or regulatory inquiries?
  • Which AI projects are planned, and which records do they depend on?

Illustrative: a deal team reviews a commercial HVAC services target#

Illustrative: a fictional deal team is reviewing a commercial HVAC service company that runs on ServiceTitan and maintains equipment for property managers under service agreements. The exposure review is reassuring, because the work is on-site and priced by contract.

The AI in use review finds that customer service staff paste service agreement details into a consumer chat tool to draft emails. The data rights review finds years of linked service calls, quotes, work orders and callbacks, but also property manager contracts that treat building information as confidential.

The team treats neither finding as a reason to walk away. It adds a covenant to move staff onto an approved business account, and the post-closing plan includes a rights review of the service records before any internal AI or licensing project begins.

Where SourceX fits after closing#

SourceX's role starts when a portfolio company decides to license records, and the Group 3 findings feed straight into that work. The SourceX five-step transaction, Supply, Rights, Preparation, Approval and Delivery, starts from the diligence record rather than repeating it, and the company approves every step.

For each package that proceeds, the SourceX Evidence Packet carries forward provenance, licensing rights, permitted use, the privacy record and release authorization, so restrictions found in diligence stay attached to the records.

Frequently asked questions

Who should run AI diligence on a mid-market deal?

Split it by group. The commercial adviser covers exposure, the technology adviser covers AI in use, and deal counsel covers data rights and governance. One deal team member should own the checklist and reconcile the findings, because the groups overlap and gaps hide between advisers.

What if the target says it does not use AI at all?

Verify it. Ask for the list of software subscriptions and expense reports for AI tools, then check the admin settings of core systems such as the help desk, CRM or field service platform. Some business software now ships with AI features switched on by default, so the purchase list alone can miss them.

Should data rights findings change valuation?

They change structure more often than price. Restricted records usually mean a narrower plan for AI or licensing, which matters only if the thesis depended on those records. Where it did, findings may justify specific representations, an indemnity or a price conversation.

Do the same questions work for add-on acquisitions?

Yes, in a lighter form. For a small add-on, the 12 questions often fit into one call and a short document request. Data rights still deserve attention, because add-ons bring legacy systems and customer contracts that the platform inherits.

How do AI findings carry into the post-closing plan?

Turn each finding into an owner and a next step. Exposure findings shape the commercial plan, AI in use findings become policy and account changes, and data rights findings become a rights review of the records most likely to support internal AI or licensing. Track them like any other integration item.

Sources

  • The Data & Trust Alliance's Data Provenance Standards (version 1.0.0 specification) define dataset metadata in three groups: Source, Provenance and Use. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify