Skip to content

Software companies

Customer DPAs and de-identified data: what secondary use is allowed?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

A customer DPA allows secondary use of de-identified data only through an express carve-out, and only within that carve-out's definition, purpose and permitted recipients. The decision rule: data qualifies only if it meets the DPA's own de-identification standard, the intended use matches the stated purpose, the recipient is allowed, and no other document forbids it.

Key takeaways

  • A DPA starts from processing on the customer's instructions, so any secondary use needs an express carve-out.
  • The common carve-outs cover aggregated data, de-identified or anonymized data, and use to improve the services, and each reaches different things.
  • Removing names rarely meets a contractual de-identification standard for free-text records such as tickets and chats.
  • A carve-out that permits internal improvement usually does not permit disclosure to an outside AI developer.
  • Apply four tests for every governing agreement: definition, purpose, recipient and no conflicting document.

What does a customer DPA allow beyond running the service?#

A customer DPA allows very little beyond running the service unless it says so expressly. Most DPAs cast the vendor as a processor or service provider that handles personal data only on the customer's documented instructions, so any secondary use, including building datasets or licensing records, has to come from a carve-out in the text.

Carve-outs exist because vendors need some room to secure the platform, produce benchmarks and improve the product. Most were drafted with analytics in mind rather than AI training data, which is why they need a careful read before anyone relies on them.

Look in four places: the definitions, the clause on the customer's processing instructions, the annex describing the nature and purpose of processing, and the clause on what the vendor may keep after termination. Carve-outs are often split across them, and an annex that lists purposes narrowly can undercut a broader carve-out in the body.

How to read the common DPA carve-outs#

The common DPA carve-outs each answer a narrower question than their labels suggest. The table paraphrases typical wording and shows what each one usually reaches and where it usually stops.

Watch the structure as well as the words. Some DPAs say anonymized data falls outside the agreement entirely, which leaves the definition doing all the work. Others keep it inside the agreement but grant the vendor a license to use it, which adds purpose limits on top.

How to read the common DPA carve-outs
Carve-outTypical wording (paraphrased)Usually permitsUsually does not permit
Aggregated dataVendor may compile data combined across customers so no customer or individual is identifiedStatistics, benchmarks and product metricsRecord-level tickets, threads or documents
De-identified or anonymized dataData that no longer identifies an individual is outside the DPA, or may be used by the vendorUse of records that truly meet the definition, within any stated purposeRecords with only names removed, or re-identifiable free text
Improve the servicesVendor may process data to maintain, secure and improve its servicesInternal analytics, quality and security work, sometimes internal modelsDisclosure to an outside company for that company's own models
Vendor's own operationsVendor may process data for billing, security and legal complianceAccount management, fraud prevention, legal holdsCommercial reuse or licensing

When is data de-identified enough for the carve-out?#

Data is de-identified enough for a carve-out only when it meets the definition in that specific DPA, and many definitions require that the data cannot reasonably be linked back to an individual or a customer. Stripping names and email addresses from a support ticket rarely meets that bar alone, because free text carries job titles, locations, order numbers and distinctive events.

Privacy teams measure re-identification risk rather than assume it away, and mainstream tooling reflects that: Google's Sensitive Data Protection API offers k-anonymity, l-diversity, k-map and delta-presence risk analysis. Automated detection has limits too. The open-source Presidio project warns that its automated detection cannot guarantee finding all sensitive information, so human review stays part of any defensible method.

Customer identity matters as well. A ticket can be anonymous about people and still reveal which customer it came from through product names, configuration details or unusual workflows, and many MSAs protect the customer's confidential information alongside personal data.

State privacy laws may set their own floor as well. Under the California Consumer Privacy Act as amended by the CPRA, information counts as deidentified only if the business takes reasonable measures so it cannot be associated with a consumer or household, publicly commits to keep it deidentified and not attempt reidentification, and contractually obligates any recipients to do the same. Where a law like that may apply, a license would need to carry the no-reidentification obligation to the licensee.

A four-part decision rule before any secondary use#

A four-part decision rule keeps the analysis honest. Each part must pass under every customer agreement that governs the records, and a failure on any part ends the inquiry for that customer's data.

Where all four pass, document the reasoning and the version of each agreement relied on. Where any part fails, the usual route is consent: a short addendum naming the record types, the method and the permitted recipient.

  • Definition: do the actual records meet the DPA's own definition of aggregated, de-identified or anonymized data?
  • Purpose: does the intended use fit the purpose the carve-out names, such as improving the services, or go beyond it?
  • Recipient: does the carve-out allow disclosure to a third party, or only use by the vendor itself?
  • Conflict: do the MSA, an AI rider, the privacy notice or a negotiated side letter forbid what the DPA seems to allow?

Why processor status narrows what you can do#

Processor status narrows what a vendor can do because the role is defined by acting for the customer. Under GDPR and similar frameworks, a processor that starts using personal data for its own purposes may be treated as a controller for that use, with obligations of its own, and some US state privacy laws draw a comparable line for service providers.

That is why licensing at SaaS companies usually starts with company-owned records, such as engineering history, internal documentation and product decisions, rather than customer data. Those records are not processed on anyone's instructions, although they still need personal and customer details removed before anything moves.

Data that is not personal data at all, such as system telemetry without user identifiers, may sit outside the DPA, but the MSA usually still governs it as customer data or usage data. Check both agreements before treating any category as free to use.

Illustrative: a returns management software company reads three DPA versions#

Illustrative: a fictional returns management software company serves online retailers. Over the years it has used three DPA templates: the first says anonymized data falls outside the agreement, the second lets the vendor use de-identified data to improve the services, and the third, adopted for enterprise deals, adds a flat ban on AI training.

An AI developer shows interest in support conversations about return disputes. Counsel runs the four-part test per template. Under the third template the answer is no. Under the second, purpose and recipient fail, because licensing to another company is not improving the vendor's services. Under the first, the definition is the obstacle: the conversations hold shopper names, order numbers and retailer-specific details that removal scripts miss.

The company decides not to license customer conversations. It offers instead its internal escalation notes and the linked engineering fixes, with retailer and shopper details removed, and records why none of the DPA templates reaches them.

How SourceX approaches DPA carve-outs#

SourceX reviews DPA templates and negotiated versions in the Rights step of the SourceX five-step transaction, and the Preparation step removes personal and confidential details from whatever remains in scope. The fit check needs only metadata, such as which templates exist and roughly how many customers sit on each.

When a package proceeds, the SourceX Evidence Packet records the rights analysis, permitted use, the privacy record of what was removed and how, and release authorization, so the reasoning behind each carve-out decision stays attached to the data.

Frequently asked questions

Is pseudonymized data de-identified under a DPA?

Usually not. Replacing names with consistent tokens lets anyone holding the key link records back, and GDPR generally treats pseudonymized data as personal data. Some DPAs define de-identified data more loosely, so read the definition, but expect pseudonymized records to stay inside the agreement's protections.

Does an improve-the-services clause allow training our own AI features?

Sometimes, if the model serves the vendor's own service and the DPA, MSA and privacy notice support it. Many customers now restrict this in AI riders. Even where internal training is allowed, the clause rarely stretches to handing the data, or a model trained on it, to another company.

What if customers signed different DPA versions?

Map each customer to the version that governs, using signature dates, click-through acceptance records and negotiated amendments. Records then inherit the rules of their customer's version. The oldest and newest templates often point in opposite directions, so a single company-wide answer is rarely correct.

Can we update our DPA template to allow licensing?

You can change the template for new customers and renewals, subject to negotiation, but changes generally do not reach records collected under earlier versions without the customer's agreement. Any update should be specific about record types, methods and recipients, and reviewed with counsel.

Who decides whether data meets the DPA's de-identification standard?

The vendor makes the call and carries the risk, usually through counsel and a privacy lead, with documented testing on samples. Some companies also bring in an outside expert. Write down the method, the residual risk and who approved it, because a customer or regulator may ask later.

Sources

  • Google's Sensitive Data Protection API offers four re-identification risk-analysis metrics: k-anonymity, l-diversity, k-map estimation and delta-presence estimation. Source
  • Post-CPRA, Cal. Civ. Code 1798.140(m) treats information as deidentified only if the business takes reasonable measures to ensure it cannot be associated with a consumer or household, publicly commits not to reidentify it, and contractually obligates any recipients to comply. Source
  • Presidio's documentation warns that because it uses automated detection mechanisms, there is no guarantee that Presidio will find all sensitive information, so additional systems and protections should be employed. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify