Definitions and comparisons
Aggregated data vs de-identified data: what's the difference?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
Aggregated data is a summary across many records, such as counts, averages or rates, while de-identified data keeps individual records but removes or transforms the details that identify people. AI developers usually want de-identified record-level data, and an aggregated data clause in a SaaS customer contract rarely covers licensing it.
Key takeaways
- Aggregated data describes groups; de-identified data still describes individual events, conversations or transactions.
- Record-level de-identified data is what most AI training and evaluation work needs, because models learn from the detail of each case.
- Aggregates over small groups can still point to one customer or person, so aggregation is not automatically safe.
- Aggregated data clauses in SaaS contracts usually permit statistics and benchmarks, not disclosure of record-level customer content.
- Under US state privacy laws, both categories can fall outside personal information, but each has its own conditions.
What is aggregated data?#
Aggregated data is information combined across many records so that the output describes a group rather than any single record. In a SaaS company that means numbers such as weekly ticket volume by product area, median time to first response, feature adoption by customer segment or churn rate by plan.
Aggregates are useful for dashboards, benchmarks and board reporting. They are far less useful for training AI systems, because the individual conversation, the reasoning in a code review or the sequence of steps that fixed a bug has been collapsed into a figure.
What is de-identified data?#
De-identified data is record-level data from which the details that identify a person, and often a customer company, have been removed or transformed. Each support ticket, CRM activity or Jira issue still exists as its own record, with names, emails, account numbers and other identifiers stripped or replaced.
The method matters as much as the label. Several US state privacy laws, the CCPA among them, treat data as deidentified only when it cannot reasonably be linked to a person and the business backs that up with technical measures, a public commitment not to re-identify, and contract terms that bind recipients. Counsel confirms which definitions may apply.
Some packages sit between the two. Records can be generalized, so exact dates become months, cities become regions and contract values become bands, while each ticket or issue remains its own record. Generalization lowers re-identification risk without collapsing the record into a statistic, which is why it is a common step in preparing business text.
Aggregated vs de-identified data, side by side#
Aggregated and de-identified data differ in the unit of data, their usefulness for AI and the contract language that usually covers them. The table compares the two for a typical B2B software company.
| Dimension | Aggregated data | De-identified data |
|---|---|---|
| Unit | A group, such as all tickets in a month | A single record, such as one ticket thread |
| SaaS example | Median resolution time by product module | A ticket thread with customer names and emails replaced |
| Usefulness for AI training | Low; patterns are already summarized | High; each case shows language, reasoning and outcome |
| Usefulness for evaluation | Limited to benchmarks and baselines | Strong; real cases can test model answers |
| Re-identification risk | Low for large groups, real for small ones | Depends on method, free text and quasi-identifiers |
| Typical legal treatment | Often outside personal information when not linkable | Outside personal information only if the conditions are met |
| Typical contract wording | Aggregated data, usage statistics, benchmarking | Anonymized or de-identified customer data, often undefined |
Which form fits which use?#
The right form depends on what the recipient will do with the data. Matching the form to the use keeps risk proportional and avoids preparing records nobody can use.
| Use | Form that usually fits | Why |
|---|---|---|
| Board and investor reporting | Aggregated | Decisions rest on trends, not individual cases |
| Published industry benchmarks | Aggregated with minimum group sizes | Readers compare themselves with a group |
| Training an AI support agent | De-identified record-level | The model needs full conversations and resolutions |
| Evaluating a coding assistant | De-identified record-level | Real issues and fixes test whether answers hold up |
| Describing a dataset to a prospective buyer | Aggregated metadata | Counts and date ranges describe scope without sharing records |
Can aggregated data still identify someone?#
Aggregated data can still identify someone when a group is small or unusual. A report of average contract value for customers in one state and one industry may describe a single customer, and a count of support escalations for an enterprise tier with few accounts can expose that account's problems.
Privacy teams handle this with minimum group sizes, suppression of small cells and measures such as k-anonymity, which checks that each combination of attributes is shared by enough records. These measures are built into mainstream tooling: Google's Sensitive Data Protection API offers k-anonymity, l-diversity, k-map and delta-presence risk analysis.
What aggregated and anonymized data clauses usually allow#
Aggregated and anonymized data clauses usually allow a SaaS vendor to compile statistics from customer use of the service and use them to improve the product, publish benchmarks or report trends. Whether they allow licensing record-level customer content to an AI developer is a different question, and the answer is usually no unless the wording clearly says so.
Read the clause word by word before relying on it. Small drafting choices change its reach.
- Conjunction: does it say aggregated and anonymized, which suggests both conditions, or aggregated or anonymized?
- Definition: is anonymized defined, and does the definition cover customer company identity as well as individuals?
- Purpose: is use limited to improving the service, or does it extend to any lawful purpose?
- Disclosure: may the vendor share the data with third parties, or only use it internally?
- Source: does the clause cover customer content such as tickets and messages, or only usage and telemetry data?
- AI wording: do newer versions of the agreement address model training, and which customers signed which version?
Illustrative: a construction software vendor checks its contracts#
Illustrative: a fictional vendor of bid management software keeps customer support conversations in Intercom, engineering work in Jira and GitHub, and product analytics in a data warehouse. Its customer agreement lets it create aggregated data from use of the service for product improvement and benchmarking. The CEO asks whether that clause covers licensing de-identified support conversations to an AI developer.
The general counsel concludes that it does not. The clause speaks of statistics derived from use, and support conversations are defined as customer content. The company instead scopes a license around its own engineering records, Jira issues and code reviews, which the customer agreement does not restrict, and updates its agreement template so future customers can opt in to de-identified use of support content.
How SourceX treats aggregated and de-identified records#
SourceX focuses on record-level data prepared to a documented de-identification method, because that is what AI developers can use. Aggregates may describe a dataset during the fit check, but they are rarely the product.
In the SourceX Enterprise Data Value Framework, human-generated signal and AI utility raise value, while privacy burden and preparation cost reduce net value; aggregation removes most of the signal along with the risk. The privacy record in each SourceX Evidence Packet names the method used, so a buyer knows exactly which category it is receiving.
Frequently asked questions
Is aggregated data personal information under the CCPA?
The CCPA treats aggregate consumer information separately from personal information: information about a group or category of consumers, with individual identities removed, that is not linked or reasonably linkable to any consumer or household. Small or narrow groups can fail that test, so counsel reviews how aggregates were built.
Can we license aggregated benchmarks to AI developers?
You can offer them, and some buyers use benchmarks to calibrate or evaluate systems. Demand is limited, though, because training work needs individual records. Aggregates are more often useful as a description of what a record-level package contains. If you do license benchmarks, apply minimum group sizes so no single customer can be singled out.
Does de-identifying customer data make it ours to license?
Not by itself. De-identification addresses privacy, while ownership and permission come from contracts. If a customer agreement says customer content belongs to the customer or may be used only to provide the service, de-identification does not override that term.
Is pseudonymized data the same as de-identified data?
No. Pseudonymized data replaces names with codes while a key or mapping still exists somewhere, so the records can be relinked. Many laws continue to treat pseudonymized data as personal data. De-identification aims for records that cannot reasonably be linked back.
Which is safer to license?
Aggregated data is usually lower risk but rarely what a buyer wants. De-identified record-level data carries more risk but much more value, and that risk is managed through the method, a sampled human review of the output and contract terms that prohibit re-identification and onward sharing.
Sources
- Google's Sensitive Data Protection API offers four re-identification risk-analysis metrics: k-anonymity, l-diversity, k-map estimation and delta-presence estimation. Source
- Post-CPRA, Cal. Civ. Code 1798.140(m) treats information as deidentified only if the business takes reasonable measures to ensure it cannot be associated with a consumer or household, publicly commits not to reidentify it, and contractually obligates recipients to comply. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.