Private equity and portfolios
AI features in acquired products vs licensing records out: a holdco rule
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
A SaaS company can use customer data to train AI only where its customer contracts and terms allow it, and acquired products often carry terms that are silent or restrictive. Licensing the vendor's own operational records, such as code, engineering issues, support replies and release notes, follows a separate rights path. Holdcos should decide the two questions separately.
Key takeaways
- In-product AI built on customer content needs permission in customer contracts, which acquired products may lack.
- The vendor's own code, engineering issues, release notes and internal discussions follow a separate, usually simpler rights path.
- Support tickets are mixed records: agent replies are the vendor's work, while customer messages may contain customer data.
- Whether new terms can reach data collected under older terms is a question for counsel, product by product.
Can a SaaS company use customer data to train AI?#
A SaaS company can use customer data to train AI only to the extent its customer contracts, terms of service and privacy commitments permit it. For an acquired product, that means reading the terms that applied when the data was collected, not just the current version, plus any negotiated agreements with larger customers.
Acquired vertical software products often carry terms written before anyone considered model training. Some are silent, some limit data use to providing the service, and some include aggregate or de-identified data clauses whose reach is unclear. Silence is not permission, so treat each product as its own rights question.
That question is separate from a second one that holdcos often merge with it: whether the vendor may license its own operational records to outside AI developers. The two have different owners, different permissions and different risks, and mixing them slows both.
The holdco rule in one table#
The holdco rule is simple: decide in-product AI by customer permission, decide licensing by ownership of the vendor's own records, and never use one decision to justify the other. The table shows how the two tracks differ on every point a group COO will be asked about.
| Question | In-product AI on customer data | Licensing the vendor's own records |
|---|---|---|
| Whose records? | Data customers enter, upload or generate in the product | Code, engineering issues, code reviews, release notes, internal docs, support replies |
| Where permission comes from | Customer contracts, terms of service, DPAs, privacy notices | Company ownership, employee and contractor agreements, vendor terms, open source licenses |
| Who decides | Product leadership with counsel, often with customer notice or consent | The supplier entity's signer, with sponsor and lender consents where required |
| Typical preparation | Contract review, opt-in or opt-out design, customer communication | Secret scanning, removal of personal and customer details, scope review |
| Main risk | Breaching customer terms or privacy commitments | Leaking customer data embedded in engineering or support records |
| Where value shows up | Product features, retention and pricing | A licensing revenue line inside the hold |
Which records belong to the vendor and which to customers?#
Records belong to the vendor when its own people created them while running the business: source code the company wrote, issues and pull request discussions, release notes, runbooks, internal wikis and the replies support agents wrote. Records are customer-controlled when they consist of what customers put into the product.
The hard cases sit in between. A support ticket combines the customer's description of a problem, which may include their own data, with the vendor's diagnosis and fix. Bug reports may attach customer files. Repositories may include third-party and open source components under their own licenses.
- Clearly vendor-owned: internal engineering discussions, code reviews, release notes, internal documentation.
- Vendor-owned with checks: source code, after reviewing open source and third-party licenses and scanning for secrets.
- Mixed: support tickets and bug reports, where customer details are removed or replaced during preparation.
- Customer-controlled: data in the product's database, customer uploads and customer-configured content.
What to check in an acquired product's contracts#
Checking an acquired product's contracts for AI use takes a structured read, because the relevant language is spread across several documents and several versions of each. Start with the items below, and record for each product which documents govern which customers and periods.
- Every terms of service version, the dates each applied and how changes were notified.
- Master agreements with larger customers that override the standard terms.
- Data processing agreements that limit processing to providing the service.
- Aggregate, de-identified or usage data clauses, and how they define those terms.
- Confidentiality clauses that treat customer data as the customer's confidential information.
- AI-specific clauses and any security questionnaires where the product described its data use.
Changing terms to allow in-product AI#
Changing terms to allow in-product AI is possible but slow, and it rarely reaches backward cleanly. Product teams typically combine updated terms, notice to customers, and an opt-in or opt-out for training on customer content, with enterprise customers handled through individual amendments.
Whether updated terms can cover data collected under older terms depends on the original wording, how changes were notified and which privacy laws may apply. Counsel assesses that product by product. Meanwhile, licensing the vendor's own records does not have to wait for any of it.
Mistakes holdcos make when the two tracks blur#
Holdcos make predictable mistakes when the two tracks blur, and each one either stalls a legitimate licensing program or exposes the group to a customer dispute. Keeping a separate register for each track, per product, prevents most of them.
| Mistake | Why it causes trouble | Better approach |
|---|---|---|
| Treating a licensing review as cover for in-product training | Rights over vendor records say nothing about customer content | Run a separate customer-permission review for each AI feature |
| Pausing licensing until terms are updated | Vendor-owned records do not depend on customer terms changes | Scope licensing of vendor records in parallel |
| Applying one product's terms across the group | Each acquisition brought its own contract history | Keep a dated terms register per product |
| Licensing raw support tickets | Customer messages may contain customer data and confidential details | Remove or replace customer details and keep the agent's reasoning |
| Assuming new terms reach old data | Coverage depends on the original wording and how notice was given | Ask counsel to map which data each terms version covers |
Illustrative: a holdco with three vertical software products#
Illustrative: a fictional software holding company owns a property management product, a towing dispatch product and a collision repair estimating product, each acquired with its original terms. The group COO wants AI features in all three and asks whether the same records could also be licensed.
The rights review finds that the property management terms limit data use to providing the service, the towing product's terms are silent, and the estimating product's larger customers signed agreements treating all their data as confidential. In-product AI on customer data goes to counsel and product leadership, starting with an opt-in program for the towing product.
Separately, all three companies hold years of Jira issues, GitHub code reviews, release notes and Zendesk agent replies. Those records move into a licensing review on their own track, with customer details stripped from tickets and secrets scanned out of repositories. Neither track waits on the other.
How SourceX handles the licensing track#
SourceX works on the licensing track: the vendor's own operational records. Each product company runs the SourceX five-step transaction, Supply, Rights, Preparation, Approval and Delivery, and the rights step separates vendor-owned records from customer-controlled content before anything is prepared.
Every approved package carries a SourceX Evidence Packet with provenance, licensing rights, permitted use, the privacy record and release authorization. The records are licensed, not sold, so each product company keeps ownership.
Frequently asked questions
Does an aggregate or de-identified data clause allow model training?
Not necessarily. Those clauses were often written with benchmarking or analytics in mind, and whether they extend to training models depends on how they define the data and the permitted purposes. Counsel should read each clause in context before anyone relies on it.
Do we need customer consent to license our own engineering records?
Customer consent is generally not the governing question for records the company created itself, such as internal issues and code reviews. It becomes relevant when those records contain customer data, for example a customer file attached to a bug report, which is removed during preparation or excluded.
Can licensing engineering records expose our product to competitors?
It can if scoped carelessly, which is why licenses use field-of-use limits and exclude core proprietary code where needed. A company can license issue histories, code reviews and support reasoning while keeping product-specific source code out of the package.
Should each acquired product be decided separately?
Yes. Each product has its own contract history, customer base and data practices, so a decision that is sound for one may not hold for another. A holdco can still use one review template and one approval policy across all of them.
What about open source code inside our repositories?
Open source components carry their own licenses, which may set conditions on redistribution. Repositories are usually scoped to code the company wrote, with vendored third-party code excluded or reviewed, and scanned for secrets before preparation.
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.