Skip to content

Definitions and comparisons

Aggregated data vs de-identified data: what's the difference?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

Aggregated data is a summary across many records, such as counts, averages or rates, while de-identified data keeps individual records but removes or transforms the details that identify people. AI developers usually want de-identified record-level data, and an aggregated data clause in a SaaS customer contract rarely covers licensing it.

Key takeaways

  • Aggregated data describes groups; de-identified data still describes individual events, conversations or transactions.
  • Record-level de-identified data is what most AI training and evaluation work needs, because models learn from the detail of each case.
  • Aggregates over small groups can still point to one customer or person, so aggregation is not automatically safe.
  • Aggregated data clauses in SaaS contracts usually permit statistics and benchmarks, not disclosure of record-level customer content.
  • Under US state privacy laws, both categories can fall outside personal information, but each has its own conditions.

What is aggregated data?#

Aggregated data is information combined across many records so that the output describes a group rather than any single record. In a SaaS company that means numbers such as weekly ticket volume by product area, median time to first response, feature adoption by customer segment or churn rate by plan.

Aggregates are useful for dashboards, benchmarks and board reporting. They are far less useful for training AI systems, because the individual conversation, the reasoning in a code review or the sequence of steps that fixed a bug has been collapsed into a figure.

What is de-identified data?#

De-identified data is record-level data from which the details that identify a person, and often a customer company, have been removed or transformed. Each support ticket, CRM activity or Jira issue still exists as its own record, with names, emails, account numbers and other identifiers stripped or replaced.

The method matters as much as the label. Several US state privacy laws, the CCPA among them, treat data as deidentified only when it cannot reasonably be linked to a person and the business backs that up with technical measures, a public commitment not to re-identify, and contract terms that bind recipients. Counsel confirms which definitions may apply.

Some packages sit between the two. Records can be generalized, so exact dates become months, cities become regions and contract values become bands, while each ticket or issue remains its own record. Generalization lowers re-identification risk without collapsing the record into a statistic, which is why it is a common step in preparing business text.

Aggregated vs de-identified data, side by side#

Aggregated and de-identified data differ in the unit of data, their usefulness for AI and the contract language that usually covers them. The table compares the two for a typical B2B software company.

Aggregated vs de-identified data, side by side
DimensionAggregated dataDe-identified data
UnitA group, such as all tickets in a monthA single record, such as one ticket thread
SaaS exampleMedian resolution time by product moduleA ticket thread with customer names and emails replaced
Usefulness for AI trainingLow; patterns are already summarizedHigh; each case shows language, reasoning and outcome
Usefulness for evaluationLimited to benchmarks and baselinesStrong; real cases can test model answers
Re-identification riskLow for large groups, real for small onesDepends on method, free text and quasi-identifiers
Typical legal treatmentOften outside personal information when not linkableOutside personal information only if the conditions are met
Typical contract wordingAggregated data, usage statistics, benchmarkingAnonymized or de-identified customer data, often undefined

Which form fits which use?#

The right form depends on what the recipient will do with the data. Matching the form to the use keeps risk proportional and avoids preparing records nobody can use.

Which form fits which use?
UseForm that usually fitsWhy
Board and investor reportingAggregatedDecisions rest on trends, not individual cases
Published industry benchmarksAggregated with minimum group sizesReaders compare themselves with a group
Training an AI support agentDe-identified record-levelThe model needs full conversations and resolutions
Evaluating a coding assistantDe-identified record-levelReal issues and fixes test whether answers hold up
Describing a dataset to a prospective buyerAggregated metadataCounts and date ranges describe scope without sharing records

Can aggregated data still identify someone?#

Aggregated data can still identify someone when a group is small or unusual. A report of average contract value for customers in one state and one industry may describe a single customer, and a count of support escalations for an enterprise tier with few accounts can expose that account's problems.

Privacy teams handle this with minimum group sizes, suppression of small cells and measures such as k-anonymity, which checks that each combination of attributes is shared by enough records. These measures are built into mainstream tooling: Google's Sensitive Data Protection API offers k-anonymity, l-diversity, k-map and delta-presence risk analysis.

What aggregated and anonymized data clauses usually allow#

Aggregated and anonymized data clauses usually allow a SaaS vendor to compile statistics from customer use of the service and use them to improve the product, publish benchmarks or report trends. Whether they allow licensing record-level customer content to an AI developer is a different question, and the answer is usually no unless the wording clearly says so.

Read the clause word by word before relying on it. Small drafting choices change its reach.

  • Conjunction: does it say aggregated and anonymized, which suggests both conditions, or aggregated or anonymized?
  • Definition: is anonymized defined, and does the definition cover customer company identity as well as individuals?
  • Purpose: is use limited to improving the service, or does it extend to any lawful purpose?
  • Disclosure: may the vendor share the data with third parties, or only use it internally?
  • Source: does the clause cover customer content such as tickets and messages, or only usage and telemetry data?
  • AI wording: do newer versions of the agreement address model training, and which customers signed which version?

Illustrative: a construction software vendor checks its contracts#

Illustrative: a fictional vendor of bid management software keeps customer support conversations in Intercom, engineering work in Jira and GitHub, and product analytics in a data warehouse. Its customer agreement lets it create aggregated data from use of the service for product improvement and benchmarking. The CEO asks whether that clause covers licensing de-identified support conversations to an AI developer.

The general counsel concludes that it does not. The clause speaks of statistics derived from use, and support conversations are defined as customer content. The company instead scopes a license around its own engineering records, Jira issues and code reviews, which the customer agreement does not restrict, and updates its agreement template so future customers can opt in to de-identified use of support content.

How SourceX treats aggregated and de-identified records#

SourceX focuses on record-level data prepared to a documented de-identification method, because that is what AI developers can use. Aggregates may describe a dataset during the fit check, but they are rarely the product.

In the SourceX Enterprise Data Value Framework, human-generated signal and AI utility raise value, while privacy burden and preparation cost reduce net value; aggregation removes most of the signal along with the risk. The privacy record in each SourceX Evidence Packet names the method used, so a buyer knows exactly which category it is receiving.

Frequently asked questions

Is aggregated data personal information under the CCPA?

The CCPA treats aggregate consumer information separately from personal information: information about a group or category of consumers, with individual identities removed, that is not linked or reasonably linkable to any consumer or household. Small or narrow groups can fail that test, so counsel reviews how aggregates were built.

Can we license aggregated benchmarks to AI developers?

You can offer them, and some buyers use benchmarks to calibrate or evaluate systems. Demand is limited, though, because training work needs individual records. Aggregates are more often useful as a description of what a record-level package contains. If you do license benchmarks, apply minimum group sizes so no single customer can be singled out.

Does de-identifying customer data make it ours to license?

Not by itself. De-identification addresses privacy, while ownership and permission come from contracts. If a customer agreement says customer content belongs to the customer or may be used only to provide the service, de-identification does not override that term.

Is pseudonymized data the same as de-identified data?

No. Pseudonymized data replaces names with codes while a key or mapping still exists somewhere, so the records can be relinked. Many laws continue to treat pseudonymized data as personal data. De-identification aims for records that cannot reasonably be linked back.

Which is safer to license?

Aggregated data is usually lower risk but rarely what a buyer wants. De-identified record-level data carries more risk but much more value, and that risk is managed through the method, a sampled human review of the output and contract terms that prohibit re-identification and onward sharing.

Sources

  • Google's Sensitive Data Protection API offers four re-identification risk-analysis metrics: k-anonymity, l-diversity, k-map estimation and delta-presence estimation. Source
  • Post-CPRA, Cal. Civ. Code 1798.140(m) treats information as deidentified only if the business takes reasonable measures to ensure it cannot be associated with a consumer or household, publicly commits not to reidentify it, and contractually obligates recipients to comply. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify