Skip to content

Privacy, de-identification and sensitive data

Combining licensed datasets without re-identifying anyone: linkage and mosaic risk

Quick answer

Joining de-identified datasets can quietly void their de-identification. Each dataset was assessed against a specific recipient and a specific set of "reasonably available" outside information; when you add a second licensed corpus, a public file or your own CRM, the quasi-identifiers combine and records that were safely ambiguous can become unique. Treat every join as a new disclosure: re-assess risk on the combined table, check each license for linkage restrictions, and record the decision before engineers merge anything.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why a join can undo de-identification

A join undoes de-identification when it supplies the outside information the original assessment assumed an attacker did not have. The HIPAA Expert Determination standard is explicit about this: the expert must conclude that the risk is very small that the information could be used, "alone or in combination with other reasonably available information," by an anticipated recipient to identify an individual [1]. Change the recipient's available information and the premise of that determination changes with it.

Other regimes frame the test the same way. FERPA allows release of de-identified education records only after a reasonable determination that a student is not identifiable, taking other reasonably available information into account [9]. The UK ICO's anonymisation guidance ties identifiability to the other information available and the context of the release, and uses a "motivated intruder" test [6]. None of these treat de-identification as a permanent property of a file. For the legal status of each label, see de-identified vs anonymized vs pseudonymized.

The research record supports the caution. Sweeney's k-anonymity work showed that ZIP code, birth date and sex, linked against outside records such as voter files, can single out people in data stripped of names [4]. Rocher, Hendrickx and de Montjoye estimated with a generative model that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes; that figure is a model estimate, not an observed attack rate [5]. Legal scholarship continues to document re-identification through auxiliary information [7].

The mosaic effect in AI data pipelines

The mosaic effect is the case where every input passes its own review but the combination does not. It shows up in AI work in ways that look like ordinary data engineering, not like an attack.

Common patterns buyers should recognize:

  • Feature enrichment. A de-identified support-ticket corpus carries account_tier, region and signup_month. Your team joins it to an internal firmographics table on a shared hashed account_id. Each table alone had large equivalence classes; together, "enterprise tier, Montana, signed up March 2021" may match one customer.
  • Pseudonym collisions. Two suppliers both used SHA-256 of a lowercased email as a stable key without a secret salt. The keys match across suppliers and against any email list you hold, so the "pseudonym" is a join key to the real world.
  • Temporal stitching. Exact timestamps in an engineering incident log and a separate on-call chat export let you align events minute by minute, which re-attaches a named on-call engineer to otherwise de-identified tickets.
  • Free-text leakage. A finance workflow corpus has structured fields generalized, but the memo column still mentions a vendor name and invoice amount that also appear in a public procurement record.
  • RAG co-retrieval. No table is ever merged, but a retriever returns chunks from two corpora in the same context window, and the model synthesizes them. Linkage happens at inference time. De-identifying a RAG corpus without breaking retrieval covers the retrieval side.

Language models make the last two patterns worse, because they can match writing style and contextual detail that rule-based checks ignore; LLM-assisted re-identification discusses that threat model.

What each de-identification method assumed

Each method carries an assumption a join can break, and buyers should know which one they are relying on. HHS describes two HIPAA routes, Expert Determination and Safe Harbor, and notes that neither removes all risk [2].

MethodWhat it assumed about outside dataHow a join breaks it
HIPAA Safe Harbor (164.514(b)(2))18 identifier types removed and no actual knowledge that remaining data could identify someone [1]Your enrichment table restores precision (full dates, 5-digit ZIP) or gives you actual knowledge of identity
HIPAA Expert Determination (164.514(b)(1))An anticipated recipient and the reasonably available information the expert considered, often tied to a stated environment [1]New recipient, new environment or new auxiliary data outside the expert's scope
HIPAA limited data set (164.514(e))Still PHI; protected by a data use agreement [1]Join with identified data, or use outside the DUA's permitted purposes
k-anonymity or generalizationEquivalence classes of size k on a declared quasi-identifier set [4]The joined table adds columns that were not in the quasi-identifier set
Differential privacy on released statisticsFormal bound that holds regardless of auxiliary data [3]Privacy budget exhausted if you combine many DP releases of the same population

NIST SP 800-188 makes the same contrast: traditional techniques have inherent limitations against linkage, while formal methods such as differential privacy give guarantees that do not depend on what an attacker already knows [3]. For health data, the method choice is covered in Safe Harbor vs Expert Determination for AI training, and how to read the scope section of a report in how to review an Expert Determination report.

License terms that restrict linkage

Licenses and data use agreements for de-identified or limited data commonly restrict re-identification and sometimes restrict linkage, so read them before the join, not after. Data use agreements for HIPAA limited data sets commonly prohibit the recipient from identifying the information or contacting the individuals; the Harris Health template is one public example [8], and 164.514(e) requires that DUAs include such terms [1].

Clauses to look for, and what they mean for a merge:

  • No re-identification. Bars attempts to identify individuals. A deliberate join on quasi-identifiers can count as an attempt even if your purpose is feature engineering.
  • No linkage without consent. Bars combining the data with other datasets, sometimes listing allowed joins. Treat silence differently from permission; ask.
  • Recipient and environment limits. Names the entity, team or environment where data may sit. Moving it into a shared feature store can breach this without any join.
  • Notification duty. Requires you to report suspected or accidental re-identification to the licensor.
  • Onward transfer. Restricts passing the data, or derived tables, to affiliates or vendors.

Re-identification prohibition clauses in data licenses covers drafting and negotiation. If a planned join is central to your use case, say so when requesting the data so the supplier's assessment can include it. When you describe a data request to SourceX, state the datasets you intend to combine and the join keys; every dataset SourceX sources is rights-reviewed and delivered under a license that defines records, uses, term and delivery.

A join approval workflow for AI teams

A join approval is a short, recorded decision made before any code merges two sources that contain data about people. It works best as a gate in the same tool where engineers request data access, so it cannot be skipped.

  1. Register every dataset. Keep an inventory entry per source: license ID, de-identification method, the assessment's stated recipient and environment, declared quasi-identifiers, and linkage clauses.
  2. Declare the join. The requester lists join keys, columns retained and the downstream use (training mix, eval set, RAG index).
  3. Check license compatibility. Confirm each license permits the combination and the destination environment.
  4. Re-measure risk on the combined table. Compute equivalence-class sizes on the union of quasi-identifiers, check for unique rows, and run a motivated-intruder style review of free text [6].
  5. Apply mitigations. Generalize or drop columns, coarsen timestamps, re-key with a fresh secret salt per project, or move the join into a controlled environment such as a data clean room.
  6. Record and re-assess. Store the decision, and re-run step 4 when a new source joins or a source is refreshed.

Illustrative example: invented to show structure; it does not describe an available dataset.

join_request:
  id: JR-2026-041
  requested_by: ml-platform/feature-eng
  purpose: "fine-tuning mix for support-agent model"
  sources:
    - id: DS-ticket-corpus-A
      license_ref: LIC-0192
      deid_method: expert_determination
      assessment_scope: { recipient: "buyer ML team", env: "isolated VPC" }
      quasi_identifiers: [account_tier, region, signup_month]
      linkage_clause: "no linkage to identified data; other de-identified data permitted with notice"
    - id: INT-firmographics
      license_ref: internal
      deid_method: none
      quasi_identifiers: [industry, employee_band, state]
  join_keys: [account_id_hash]
  key_salt_shared_across_sources: true   # red flag
  combined_checks:
    min_equivalence_class: 1             # fails threshold of 10
    unique_rows_pct: 4.2
    free_text_review: "memo field mentions vendor names"
  decision: rejected
  mitigations_required:
    - "drop state; band region to 4 census regions"
    - "re-key with per-project secret salt"
    - "send notice to licensor per LIC-0192"
  reassess_on: [new_source_added, source_refresh]

The threshold of 10 and the field names here are placeholders; set them with your privacy team and, for health data, with the expert who wrote the original determination.

Measuring linkage risk on a combined dataset

Measure risk on the joined output, using the attributes an outsider could plausibly know, not on each input separately. Useful checks include equivalence-class size across the merged quasi-identifier set, the share of records unique on that set, and population uniqueness estimates rather than sample uniqueness, since Rocher and colleagues showed that releasing only a sample does not by itself make records safe [5].

For text-heavy corpora, structured metrics miss most of the risk. Run named-entity and rare-phrase scans across the combined corpus, look for strings that occur in only one record but also appear in public sources, and sample records for manual intruder review. Re-identification risk assessment for licensed datasets lists the methods a buyer can ask a supplier to run, and those same methods apply to your combined table.

Training adds one more channel. A model fine-tuned on a joined corpus can memorize and emit rare combinations, so linkage risk carries into outputs; see training-data extraction and memorization risk.

How SourceX handles requests that involve combining datasets

SourceX sources operational datasets from US companies on request and runs a Find, Assess, Agree, Transact and Manage process, where Assess covers the data and its licensing permissions. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Describe the data you need and the sources you plan to combine it with at sourcex.si/buyers.

For the wider cluster, start at the privacy and de-identification buyer's guide or the AI data hub, and see the re-identification glossary entry.

Frequently asked questions

Does joining two HIPAA de-identified datasets create PHI?

It can. If the combined data no longer meets 164.514(b), for example because the expert's scope did not cover the new auxiliary data or because you now have actual knowledge of identity, the de-identified status is in doubt [1]. Ask counsel and, where possible, the original expert before relying on it.

Is hashing identifiers enough to make cross-dataset joins safe?

No. Unsalted or shared-salt hashes of emails, phone numbers or account numbers are deterministic, so anyone with the original values can recompute them. Use a secret, project-specific key and keep it out of the training environment.

Does linkage risk apply to data we never physically merge?

Yes. Co-retrieval in a RAG index, a model trained on several corpora, or analysts querying two tables side by side can all link records without a SQL join.

Sources

  1. Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  2. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  3. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  4. Data Privacy Lab (Latanya Sweeney), "k-Anonymity: A Model for Protecting Privacy" (2002). https://dataprivacylab.org/projects/kanonymity/
  5. Nature Communications (Rocher, Hendrickx, de Montjoye), via PubMed Central, "Estimating the success of re-identifications in incomplete datasets using generative models" (2019). https://pmc.ncbi.nlm.nih.gov/articles/PMC6650473
  6. UK Information Commissioner's Office, "How do we ensure anonymisation is effective?". https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/
  7. Michigan Journal of Environmental & Administrative Law, "Clarke, Spring 2026" (2026). https://www.mjeal-online.org/2026/04/12/clarke-spring-2026/
  8. Harris Health System, "Limited Data Set Use Agreement". https://www.harrishealth.org/SiteCollectionDocuments/Limited-data-sets-use-agreement.pdf
  9. U.S. Government Publishing Office (govinfo), "34 CFR 99.31 - Under what conditions is prior consent not required to disclose information?" (2026). https://www.ecfr.gov/current/title-34/subtitle-A/part-99/subpart-D

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data