Skip to content

Private equity and portfolios

What data has already left an acquired company through AI tools?

By SourceX Editorial · Updated

Short answer

Data can leave an acquired company through AI tools by four channels: AI features switched on in its SaaS systems, vendor terms that allow training on customer records, browser and personal AI tools used by staff, and bulk exports into outside services. Audit all four soon after closing, because earlier disclosure can limit later licensing and any offer of exclusivity.

Key takeaways

  • Audit four channels: SaaS AI features, vendor training terms, browser and personal AI tools, and bulk exports or integrations.
  • Admin consoles and vendor contracts show what was switched on; short interviews show what staff actually did.
  • Stop new exposure first, then document what already happened, without trying to reconstruct every prompt.
  • Records given to an AI vendor under training-permissive terms can weaken a later offer of licensing exclusivity.
  • Findings belong in the integration file and, where personal data or customer confidentiality is involved, in front of counsel.

Where does data leave a company through AI tools?#

Data leaves a company through AI tools by four main channels: AI features inside the business software it already pays for, vendor terms that let those vendors train on customer records, browser extensions and personal AI accounts used by staff, and bulk exports or integrations that move records into outside AI services.

None of these requires bad intent. A support lead turns on reply suggestions in the help desk; a sales rep installs a notetaker; an analyst uploads a customer export to summarize churn reasons. For an acquirer, the questions are what happened before closing and what is still happening now.

The audit checklist#

The audit checklist covers each channel with where to look and who owns the check. Run it soon after closing, alongside the system register, so both draw on the same list of systems.

The audit checklist
ChannelWhat to checkWhere to lookOwner
SaaS AI featuresWhich AI features are enabled in each system holding customer or employee recordsAdmin consoles of the help desk, CRM, phone system, field service platform and ERPIT lead
Vendor training termsWhether each vendor may use customer data to train or improve its models, and any opt-outs takenVendor contracts, data processing terms, admin privacy settingsCounsel with the IT lead
Browser and personal AI toolsExtensions and personal AI accounts used with company recordsManaged browser reports, expense claims, interviews with department headsIT lead with department heads
AI notetakers and recordersTools that join calls, and where their transcripts are storedCalendar integrations and meeting platform app listsIT lead
Bulk exportsLarge exports of customer, ticket or code records, and where they wentSystem audit logs, export histories, file sharing logsIT lead
Integrations and API keysAutomations and keys that send records to AI servicesIntegration platforms, developer accounts, code repositoriesCTO or engineering lead

SaaS AI features and vendor training settings#

SaaS AI features are often the largest channel because they process records at volume and run quietly once enabled. Many help desk, CRM and call platforms now offer AI summaries, reply drafting or call analysis, and whether the vendor may use that data to improve its models varies by vendor, contract and plan.

Check both the feature switch and the training setting, and note the date of each change if the console shows a history. Where it does not, record the current state and the date you checked. Counsel should read the vendor's data processing terms, since marketing pages are not a reliable guide to what the contract permits.

Browser tools, personal accounts and notetakers#

Browser tools and personal accounts are the hardest channel to see because they sit outside the company's systems. Managed browser reports show extensions on company devices, expense claims show AI subscriptions bought on company cards, and short interviews with department heads fill in the rest.

Ask concrete questions rather than general ones. Which tools did you use to draft customer emails? Did anyone paste ticket threads, contracts or code into a chat tool? Did a notetaker join customer calls, and where did transcripts go? People answer specific questions honestly when the tone is fact-finding rather than disciplinary.

Code needs a separate check. Developers may have used AI coding assistants on personal accounts, sending repository content to a vendor under consumer terms the company never reviewed. GitHub's Copilot documentation, for example, says that from April 24, 2026, interactions on individual Copilot plans (Free, Pro, Pro+ and Max) may be used to train and improve AI models unless the user turns off the 'Allow GitHub to use my data for AI model training' setting. Record which plan each developer used and how that setting was configured.

Why earlier disclosure matters for licensing#

Earlier disclosure matters because a later license is easier to value when the company can say exactly who already holds the records. If an AI vendor received support tickets under terms that allowed training, a future licensee of those tickets may ask whether it is paying for something another developer has already used.

Exclusivity is the sharpest issue. A company cannot credibly offer exclusive rights to records it has already disclosed under training-permissive terms, and an acquirer or licensee may ask for representations on exactly that point. Disclosure does not always erase value, since prepared, linked and documented records differ from raw uploads, but it has to be known and disclosed.

What to do with what you find#

Act on findings in a fixed order, so the audit stops exposure without stalling the business.

  • Stop new exposure: switch off unreviewed AI features and training settings, and pause notetakers and integrations pending review.
  • Document the past: record each finding with the system, dates, record types and the terms in force at the time.
  • Assess with counsel: decide whether personal data, customer confidentiality or contract obligations are affected, and whether any notice is required.
  • Request deletion where the vendor's terms provide for it, and keep the written confirmation.
  • Update the AI use policy and approved tool list, and give staff a sanctioned alternative.
  • Carry the findings into the integration file and any later licensing or exit diligence.

Illustrative: a portfolio CTO audits a transportation software company#

Illustrative: a fictional sponsor acquires a transportation management software company, and the portfolio CTO runs the checklist soon after closing. The company uses Zendesk for support, HubSpot for sales, a call recording platform and GitHub for code.

The audit finds AI reply suggestions enabled in the help desk with the vendor's model improvement setting left at its default. Two sales reps used a personal notetaker on customer calls. An engineer connected a personal AI coding assistant to a private repository. A customer success manager uploaded a ticket export to a consumer chat tool to draft a quarterly review.

The CTO switches off the training setting, ends personal notetaker use, moves coding assistants to a company plan with reviewed terms, and asks counsel to assess the ticket upload. When the sponsor later considers licensing the ticket history, the file lets the rights review state which records were exposed, and the affected accounts are left out of scope.

How SourceX treats prior AI tool exposure#

SourceX asks about prior exposure in the Rights step of the SourceX five-step transaction, because buyers ask about it too. A documented audit lets the company state what was disclosed, to whom and under what terms, and scope any package around it. The audit findings can be discussed as metadata, since no records move during the initial assessment.

For each package that proceeds, the SourceX Evidence Packet records provenance and licensing rights, including any known prior disclosure that bears on exclusivity.

Frequently asked questions

Is an AI tool audit the same as a security audit?

No. A security audit looks for unauthorized access and vulnerabilities. An AI tool audit looks at authorized tools and settings that may have sent records to vendors under terms nobody reviewed. The two overlap on exports and integrations, so run them together where you can.

Can we ask an AI vendor to delete data it already received?

You can ask, and some vendors' terms provide a deletion process for customer data. Whether data already used to improve a model can be removed from that model is a separate question with a less certain answer. Keep written confirmations of any deletion and note exactly what they cover.

Should we tell customers what the audit found?

It depends on what was disclosed and on what contracts, privacy notices and laws require. Some findings may call for notice; many will not. Counsel should decide from the documented findings, which is one more reason to record the system, dates and record types for each item.

How do we stop this from happening again?

Adopt an AI use policy with an approved tool list, put vendor AI settings on the quarterly review, route new AI tool purchases through IT, and give staff a company-approved alternative for the tasks they were doing with personal tools. Bans without alternatives tend to push use back into personal accounts.

Sources

  • GitHub's Copilot documentation states that starting April 24, 2026, interactions on individual Copilot plans (Free, Pro, Pro+ and Max) may be used to train and improve AI models, and users can turn this off with the 'Allow GitHub to use my data for AI model training' setting. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify