Skip to content

Privacy and preparation

What de-identification standard should a data license specify?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

A data license should specify a de-identification standard that names the legal definition it relies on, the methods used for each record family, an agreed limit on residual identifiers with the quality checks that test it, the licensee's promise not to re-identify, and audit and remedy terms. A bare promise of anonymized data settles none of these.

Key takeaways

  • Reference the definition in the privacy laws that apply as a floor, then add contract terms that make it measurable.
  • Name the methods per record family, because structured fields and free text fail in different ways.
  • Set an acceptance test: an agreed sampling method, human review and a residual limit for identifiers found.
  • The licensee's no re-identification covenant, passed on to anyone it shares with, is part of the standard rather than an add-on.
  • Audit rights and a defined response to a found identifier keep the standard enforceable on both sides.

Why a license needs its own de-identification standard#

A license needs its own de-identification standard because anonymized and de-identified mean different things in different laws, tools and buyer policies. Without a written standard, the supplier cannot show it delivered what was promised, and the licensee cannot show what it agreed not to do.

The standard also carries legal weight. Many state privacy laws may treat data as de-identified only when the holder takes reasonable measures, commits publicly not to re-identify it and requires recipients, by contract, to make the same promise. The license is where that contractual element lives, so a vague clause can leave a package counted as personal data.

A complete clause has six parts: the definition referenced, the methods, a residual limit, quality assurance, the licensee's no re-identification covenant, and audit with remedies. The sections below take them in that order.

Which definition should the clause reference?#

The clause should reference the definition in the privacy laws that apply to the records as a floor, then add contract terms that make it testable. Statutory definitions describe an outcome, whether data can reasonably be linked to a person, but rarely say how to measure it.

Most licenses combine the first two options below with a reasonableness test for free text. Expert determination borrows the structure of the HIPAA Privacy Rule at 45 CFR 164.514(b), where a qualified expert finds re-identification risk very small and documents the methods; outside healthcare it is sometimes adopted by contract when a record family is unusually sensitive.

Which definition should the clause reference?
Reference optionHow it worksBest useWatch for
Applicable law's definitionIncorporates the statutory test of each relevant stateThe legal floor in every licenseDefinitions differ by state and change through amendments
Enumerated identifier listNames the fields that must be removed or transformedStructured CRM, ERP or billing exportsLists miss identifiers hidden in free text
Expert determination modelA qualified expert certifies that re-identification risk is very smallSensitive or unusual record familiesCost, and the expert's assumptions must be written down
Measured risk thresholdApplies metrics such as k-anonymity to quasi-identifiersTables with dates, locations and categoriesHard to apply to tickets, emails and notes

Naming the methods for each record family#

Methods should be named per record family because a support ticket, an order line and a code review carry identifiers in different places. A shared vocabulary helps: the Data & Trust Alliance's Data Provenance Standards list privacy-enhancing tools that include anonymization, masking, redaction, pseudonymization, tokenization and k-anonymity, which gives both sides common terms.

For a neutral reference on what each technique does, ISO/IEC 20889:2018 sets out de-identification terminology and classifies techniques by how well they reduce re-identification risk. Citing it in a schedule lets both sides use the same names without arguing over definitions.

Add one sentence that rules out a common misunderstanding: pseudonymized data, where anyone keeps a way to reverse the tokens, does not meet the standard. That line alone settles many later disputes about what was promised.

  • Direct identifiers stripped, such as personal names, contact details, home addresses and customer account numbers.
  • Tokenization where records must stay linked, with the token key destroyed before release.
  • Generalization of dates, locations and rare categories that could single out a person.
  • Free-text redaction of names, signatures and contact details inside tickets, emails and notes.
  • Suppression of records too unusual to protect, such as a one-off incident at a named site.
  • Removal of secrets and credentials from code, configuration files and logs.

Setting a residual limit and the checks that test it#

A residual limit is the agreed tolerance for identifiers that survive preparation, and it is what makes the standard testable. No method catches everything in free text, so the honest commitment is a sampling process, a limit on what a sample may contain and a fix-and-retest rule when a sample fails.

Automated detection is part of quality assurance, not all of it. Presidio, an open-source toolkit for spotting personal information in text, is candid about this in its own documentation: automated detection carries no guarantee, and other protections should sit alongside it. A clause that relies on scanning alone inherits that gap, so human review of samples belongs in it.

For structured fields, measured risk is a practical test. Google's Sensitive Data Protection API is one example of cloud tooling that ships k-anonymity, l-diversity, k-map and delta-presence metrics, and such measures suit dates, postal codes and job titles far better than free text.

Setting a residual limit and the checks that test it
Quality elementWhat the clause specifies
SamplingHow samples are drawn from each record family and who draws them
ReviewAutomated scanning plus human reviewers working from written instructions
Residual limitThe agreed tolerance for identifiers found in a sample, set per record family
Failure ruleRe-prepare the affected batch and draw a fresh sample
Risk metricsFor tabular data, which measures are run and what result is acceptable
RecordResults kept with the privacy record for each delivery

The licensee's side: no re-identification and flow-down#

The licensee's commitments complete the standard, because de-identification can be undone by anyone who links the package to other data. The clause should bar attempts to re-identify individuals and any linking with other datasets for that purpose, and require every recipient the licensee may share with to accept the same terms.

Add a duty to report and remove. If the licensee finds a direct identifier, it should stop using the affected records, tell the supplier promptly and delete them or accept a corrected batch, so a licensee that spots a miss becomes part of quality control rather than a liability.

Model outputs deserve a line too. A covenant to take reasonable steps so models do not reproduce identifying details from the package addresses the way training data most often resurfaces.

Audit, remedies and change control#

Audit and remedy terms make the standard enforceable for both parties. The supplier should be able to request evidence that the licensee honors the no re-identification covenant, and the licensee should be able to see the supplier's quality record for each delivery.

Change control matters over a long term. If a law's definition changes, a new record family is added or a re-identification technique becomes practical, the clause should require the parties to review the standard and apply updates to future deliveries, with counsel deciding whether earlier deliveries need action.

Illustrative clause wording#

Illustrative: a fictional precision parts manufacturer is licensing nonconformance reports, CAPA records and maintenance work orders from its QMS and ERP, and its counsel drafts the wording below. It shows how the six parts fit together and is not a template to sign without counsel.

De-identified means data that cannot reasonably be used to identify, or be associated with, a particular individual, and that meets, at minimum, the definition in each privacy law applicable to the Licensed Data. Supplier will remove or transform the identifiers listed in Schedule B, redact identifiers in free text, and destroy any token key before Delivery.

Supplier will review samples from each record family using automated scanning and human review, and will re-prepare any batch whose sample exceeds the residual limit in Schedule B. Licensee will not attempt to re-identify any individual or link the Licensed Data with other data for that purpose, will notify Supplier promptly of any identifier found, and will bind each permitted recipient to the same terms.

How SourceX approaches the de-identification standard#

SourceX settles the de-identification standard while a package is in Preparation, the third stage of the SourceX five-step transaction, once Rights has fixed the licensable scope. The supplier approves the methods, the residual limit and the sample results before the package moves to Approval and Delivery.

Inside the SourceX Evidence Packet, the privacy record holds the definition referenced, the methods by record family and the quality results, so the clause in the license and the evidence of what was done line up. Counsel for each party confirms the legal definition deal by deal.

Frequently asked questions

Should the supplier or the licensee perform de-identification?

Usually the supplier, before data leaves its systems, because the supplier knows where identifiers hide and keeps control of the raw records. Some licenses let the licensee de-identify inside a controlled environment, but identifiable data has then already been disclosed, which changes the legal analysis and the contract terms needed.

Can the standard differ between record families in one license?

Yes, and it often should. Structured CRM fields can meet a measured threshold, while support transcripts need redaction plus sampled human review. A schedule per record family keeps the main clause short and lets each family carry the method that suits it.

Is a stricter standard always better for the supplier?

Not always. Stricter standards lower privacy risk but can strip the context that makes records useful, such as sequences, timing and product detail. The aim is the least removal that still meets the legal test and the agreed residual limit, decided by counsel and the preparation team together.

Does the license cover the public commitment some laws expect?

Not on its own. Several state laws expect the holder of de-identified data to commit publicly not to re-identify it, and that statement usually sits in the privacy notice rather than the license. Confirm the notice says it before the first delivery.

What happens if an identifier is found after delivery?

The clause should answer this in advance: the licensee stops using the affected records, notifies the supplier and deletes them or takes a corrected batch. The supplier then checks whether the same miss affects other records and documents the fix in the privacy record.

Sources

  • Presidio's documentation warns that because it uses automated detection mechanisms, there is no guarantee it will find all sensitive information, and additional systems and protections should be employed. Source
  • Google's Sensitive Data Protection API offers four re-identification risk-analysis metrics: k-anonymity, l-diversity, k-map estimation and delta-presence estimation. Source
  • The Data Provenance Standards' Privacy Enhancing Tools code list includes data anonymization, encryption, masking, minimization, redaction, differential privacy, k-anonymity, l-diversity, pseudonymization, tokenization and other techniques. Source
  • Under 45 CFR 164.514(b), data can be de-identified by Expert Determination, in which a qualified expert finds re-identification risk very small, or by Safe Harbor. Source
  • ISO/IEC 20889:2018 sets out de-identification terminology and a classification of techniques, including how well each reduces re-identification risk. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify