Privacy, de-identification and sensitive data
De-identifying tabular and transactional data for machine learning: generalization, suppression and free-text columns
Quick answer
To de-identify tabular data for machine learning, classify every column as a direct identifier, quasi-identifier, sensitive attribute, key or free text, then treat each class differently. Drop or tokenize direct identifiers, generalize or suppress quasi-identifier combinations until small groups disappear, replace keys with consistent keyed pseudonyms so joins survive, and run text redaction on every free-text column. Then measure two things: residual re-identification risk and how far the joint distributions your model needs have moved.
By SourceX Editorial · Updated
This guide is written for data engineers specifying deliveries of CRM, ERP, billing, claims and ticketing tables for tabular foundation models, forecasting, text-to-SQL and agents. It sits under the privacy and de-identification hub and complements the broader guide to tabular, time-series and transactional data.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why removing names is not enough for record-level tables
Removing names, emails and account numbers leaves most structured records re-identifiable, because combinations of ordinary columns single people out. NIST's survey of the field documents repeated re-identification of data that had been released as de-identified [2]. Rocher, Hendrickx and de Montjoye built a generative model that estimates how likely a specific person is to be unique in a dataset from a handful of demographic attributes, and it held up across 210 populations even when the released data was only a sample [5]. Their widely quoted near-certain re-identification figure is a model estimate, not a count of actual attacks, but the direction is clear.
Transactional data makes this worse. A customer's sequence of merchant categories, timestamps and amounts is often more distinctive than their demographics, and every additional row adds signal. A de-identification plan that only looks at one row at a time misses the trajectory.
Classify every column before choosing a technique
The first deliverable is a column inventory that assigns each field one role, because the role determines the treatment. NIST SP 800-188 frames de-identification as a governed process that starts with understanding the data and the release model, not with a tool [1]. Use five roles:
- Direct identifiers: names, emails, phone numbers, SSN or tax IDs, card PANs, IBANs, street addresses, IP addresses, device IDs. Remove or replace.
- Quasi-identifiers: date of birth, ZIP or postal code, gender, job title, employer, store location, account open date, rare product codes. Generalize, suppress or perturb.
- Sensitive attributes: diagnosis codes, claim amounts, credit limits, complaint category. Usually kept, since they are often the prediction target, but checked for homogeneity within groups.
- Keys: customer_id, order_id, invoice_id, ticket_id, foreign keys across tables. Pseudonymize consistently.
- Free text: notes, memo, description, subject, comment, address_line_2 fields misused as notes. Route through text redaction.
Misclassification is the usual failure. ERP and CRM tables carry custom fields (Salesforce __c fields, SAP Z-fields, NetSuite custom segments) that hold phone numbers or names despite innocuous labels, so profile values, not just column names.
Generalization and suppression: what they do to joins and distributions
Generalization coarsens values and suppression removes rare cells or rows; both reduce re-identification risk at a measurable cost to model signal. Typical generalizations are date of birth to birth year or five-year age bands, five-digit ZIP to three-digit ZIP, timestamps to day or week, and exact amounts to bins. HIPAA Safe Harbor encodes one fixed version of this for health data: among 18 identifier categories, it requires removing all date elements except year, aggregating ages over 89 into one category, and removing geography below the state level except the first three ZIP digits where that area holds more than 20,000 people [3][4].
The goal of these operations is usually stated as k-anonymity: every combination of quasi-identifier values is shared by at least k records. For ML, the cost lands in three places:
- Joint distributions: coarsening age and ZIP independently can preserve each marginal yet flatten the age-by-region interaction a forecasting model depends on.
- Temporal resolution: shifting timestamps to weekly buckets destroys intraday seasonality and event ordering, which agents and sequence models need.
- Tail behavior: suppression removes exactly the rare combinations (high-value accounts, unusual products, small branches) where anomaly and fraud models learn the most.
Perturbation (noise addition, swapping, date shifting per individual) keeps resolution but changes values. NIST SP 800-188 distinguishes these traditional techniques from formal privacy methods such as differential privacy and cautions that traditional methods offer weaker guarantees [1]. Per-patient or per-customer date shifting, where one random offset applies to all of an individual's rows, is a common compromise: intervals between events are preserved while absolute dates are hidden.
Keyed pseudonyms keep referential integrity across tables
Replace every key with a deterministic keyed pseudonym, applied identically in every table, so that foreign keys still join after de-identification. The standard construction is an HMAC (for example HMAC-SHA-256) over the original ID with a secret key held by the data holder and never delivered. Plain unsalted hashes of customer numbers, emails or phone numbers are reversible by enumeration because the input space is small.
Three rules prevent the most common breakage:
- One key per release scope. The same secret across a customer table, an orders table and a support-ticket table preserves joins; a different secret per delivery prevents linking across separate licenses.
- Normalize before hashing. Trim, lowercase and canonicalize formats, or
ACME-0042andacme-0042become two customers. - Pseudonymize IDs embedded in text. Order numbers inside ticket bodies or invoice memos must map to the same pseudonym as the key column, or the join to the text is lost and the raw ID leaks.
Keep in mind that a consistent pseudonym also makes each person's full history linkable, which raises the trajectory risk described above. That trade-off is one reason linkage and mosaic risk across combined datasets deserves its own review.
Free-text columns need text redaction, not column rules
Free-text columns inside structured tables carry the same identifiers as documents and must be redacted with NER and pattern detection, not handled by column-level generalization. Support ticket bodies, claim adjuster notes, CRM activity descriptions and invoice memo lines routinely contain names, phone numbers, account numbers and street addresses. Tools such as Microsoft Presidio combine pattern recognizers with trained NER models, and the project itself warns that it cannot guarantee finding all sensitive information [6].
Practical handling for text inside tables:
- Replace detected entities with typed surrogates (
[PERSON_1],[ACCOUNT_3]) that stay consistent within a record, so a text-to-SQL or agent model can still follow references. - Reconcile text surrogates with key pseudonyms where the entity is also a key.
- Measure recall on a hand-labeled sample of each text column, not on a generic benchmark, because memo fields and ticket bodies have very different entity mixes.
- Check for indirect identifiers that NER misses, such as rare job titles or small-town incidents; see indirect identifiers in business text.
The PII redaction pipeline guide covers tool choice and evaluation in depth.
Record-level versus aggregate releases
Record-level releases support model training; aggregate releases rarely do, so most ML buyers need record-level data with stronger controls rather than tables of counts. NIST SP 800-188 describes multiple data-sharing models, from publishing de-identified records to synthetic data and controlled access, each with different risk and utility [1]. Aggregates (counts, sums by segment) suit dashboards and benchmarks but remove the per-row variation that tabular foundation models, forecasting models and text-to-SQL evaluation require.
For health data, the HIPAA rule offers a middle path: a limited data set may keep dates and some geography but excludes 16 direct identifier categories and can only be shared under a data use agreement [4]. Commercial data has no single equivalent, which is why contractual controls such as re-identification prohibitions travel with record-level licenses; see re-identification prohibition clauses.
Preserving what tabular and text-to-SQL models actually learn from
Specify which statistical properties must survive de-identification before the supplier starts, because every technique trades some of them away. TabDPT's authors found that incorporating real tables during pretraining led to faster training and better downstream generalization [8], so over-sanitized tables can erase the advantage of licensing real data. Spider 2.0 builds its 632 enterprise text-to-SQL problems on real application databases with large, messy schemas [7]; a de-identification pass that renames columns, drops nullable fields or collapses category codes makes delivered tables less representative of that environment.
Properties worth naming in the request:
- Schema fidelity: original table and column names, types, nullability and foreign keys retained.
- Cardinality: distinct-value counts of categorical columns within an agreed tolerance after generalization.
- Pairwise and conditional distributions for named column pairs (for example region by product, tenure by churn).
- Event ordering and inter-event intervals per entity.
- Null and sentinel patterns (
0000-00-00,-1,N/A), which models must learn to handle.
Column-level de-identification specification
The most useful single artifact is a column specification that both parties sign off before extraction, because it makes every treatment and its utility cost explicit.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Table.column | Role | Treatment | Join / utility effect |
|---|---|---|---|
| customers.customer_id | Key | HMAC-SHA-256, release-scoped secret | Joins preserved across all tables |
| customers.full_name, email, phone | Direct identifier | Dropped | None for modeling |
| customers.date_of_birth | Quasi-identifier | Birth year; age 90+ top-coded | Age resolution reduced to years |
| customers.postal_code | Quasi-identifier | First 3 digits; suppressed where group < k | Small regions merged |
| customers.signup_ts | Quasi-identifier | Per-customer date shift (one offset per customer) | Intervals preserved, calendar seasonality blurred |
| orders.customer_id | Key (FK) | Same HMAC as customers | Referential integrity intact |
| orders.amount | Sensitive / target | Kept; top-coded above 99.9th percentile | Tail compressed |
| orders.store_id | Quasi-identifier | Kept; stores with < k customers merged to "other" | Small-store signal lost |
| tickets.body | Free text | NER + regex redaction, typed surrogates, sampled recall check | Coreference preserved within ticket |
| tickets.order_ref_in_text | Embedded key | Mapped to orders pseudonym | Text-to-table joins preserved |
Pair the specification with a delivery format that preserves types; see CSV delivery done right and Parquet vs JSONL for deliveries.
Evidence to request from a supplier
Ask for evidence that shows both the residual risk and the utility loss, not a one-line statement that data is anonymized. A reasonable package includes the column specification, the quasi-identifier set and the k or equivalent threshold used, suppression rates per column, redaction recall measured on a labeled sample of each text column, and before/after statistics for the properties you named. For health data, ask which HIPAA method was used and, for Expert Determination, the expert's report [3].
The de-identification evidence package checklist and the re-identification risk assessment guide give fuller lists. Sector-specific walkthroughs exist for financial transaction data and spreadsheets and financial models.
SourceX sources operational datasets, including support and sales histories and finance workflows, from US companies on request; personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Buyers can describe the tables they need on the SourceX buyer page.
Request de-identified tabular data through SourceX
SourceX sources record-level operational data from US businesses on request, with every dataset rights-reviewed and delivered under a license that defines records, uses, term and delivery. Health records require HIPAA de-identification through Safe Harbor or Expert Determination, and every release is approved by the supplying company. Describe your tables, columns and required properties at sourcex.si/buyers.
Sources
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- National Institute of Standards and Technology, "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Electronic Code of Federal Regulations, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- Nature Communications (Rocher, Hendrickx, de Montjoye), "Estimating the success of re-identifications in incomplete datasets using generative models" (2019). https://pmc.ncbi.nlm.nih.gov/articles/PMC6650473
- Microsoft (presidio project), "Presidio - Data Protection and De-identification SDK". https://microsoft.github.io/presidio/
- arXiv, "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows" (2024). https://www.arxiv.org/pdf/2411.07763
- arXiv, "TabDPT: Scaling Tabular Foundation Models on Real Data" (2024). https://arxiv.org/pdf/2410.18164
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.