Evaluation and benchmarking datasets
De-identifying evaluation data without breaking the test
Quick answer
To de-identify evaluation data without breaking the test, first name the property each item measures (coreference, field formatting, retrieval cues, arithmetic, policy logic), then replace personal details with consistent, type-correct surrogates rather than a blanket REDACTED token. Transform the answer key in the same pass, re-run the item against a reference model to confirm the original failure still reproduces, and keep the raw traces in a restricted system. Blanket redaction often silently turns a hard eval item into an easy or meaningless one [1].
By SourceX Editorial · Updated
Why blanket redaction breaks evaluation items
Blanket redaction fails because it removes the very signals many eval items exist to test. A support transcript where the agent must work out that "she," "Ms. Alvarez" and "the account holder" are one person becomes trivial once all three collapse into "[REDACTED]," and a formatting test about invoice numbers has nothing left to format. The OneUptime guide on building golden sets from production failures makes the same point: swapping every entity for one token can destroy coreference, formatting or retrieval behavior and make the test meaningless [1].
The damage is usually invisible in aggregate scores. A redacted item still parses, still has a gold answer, and still produces a pass or fail, so the suite looks healthy while measuring something different. Typical failure modes:
- Coreference collapse: two different customers both become "[NAME]," so the model cannot be wrong about who is who.
- Format loss: an IBAN, policy number or ICD-10 code replaced by "XXXX" removes the check-digit or pattern logic the item was probing.
- Retrieval cue loss: the query mentions "the Hendricks account" and the target document contains the same string; redact one side differently and retrieval recall drops for reasons unrelated to the model.
- Answer-key drift: the gold answer still contains the original name or a character offset that no longer lines up after replacement.
- Distribution shift: placeholder tokens are rare in pretraining text, so the model behaves differently on them than on real names.
Start from the property under test, not the PII list
The correct transform for a field depends on what the item measures, so write that down before choosing a tool. A PII detector such as Microsoft Presidio finds entity spans and offers operators (replace, mask, hash, encrypt), but it does not know which spans carry the test [4]. That knowledge lives in your evaluation dataset specification, which should already list slices, labels and acceptance criteria per item.
Tag each item with one or more properties: identity tracking, format validity, temporal reasoning, numeric reasoning, retrieval matching, policy or eligibility logic, or tone. Then map each property to a transform family. Items with no identity-dependent property can take aggressive redaction; items that test identity tracking need surrogates that preserve distinctness and gender or number agreement where the text depends on it.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Property under test | Fields at risk | Transform that preserves it | Transform that breaks it |
|---|---|---|---|
| Coreference / identity tracking | Person and org names, pronouns, titles | Consistent per-entity surrogates (PERSON_A → "Dana Whitfield" everywhere in the item and its documents) | Single [NAME] token for all people |
| Format validity | Account, policy, card, IBAN, phone numbers | Format-preserving surrogates with valid check digits (Luhn, IBAN mod-97) | "XXXX" masks or random strings of the wrong length |
| Temporal reasoning | Dates, timestamps, durations | Per-subject consistent date shift that keeps intervals and weekday logic | Year-only truncation or independent random dates |
| Retrieval matching | Names and IDs shared across query and corpus | One keyed mapping applied to query, corpus and gold citations | Separate redaction runs per file |
| Numeric reasoning | Amounts, balances, quantities | Leave untouched if not identifying, or scale all related values by one factor | Rounding or bucketing only some of the figures |
| Policy or eligibility logic | Age, state, plan type | Generalize only to the granularity the policy uses (age band matching the rule's threshold) | Generalization that crosses the policy threshold |
Build consistent surrogates with a keyed mapping
Consistent surrogates come from a deterministic, keyed mapping: the same original value always yields the same replacement within a scope, and the key never leaves the restricted environment. A common pattern is HMAC-SHA256 over the normalized value with a secret key, used to index into a surrogate name or ID pool, so "J. Alvarez," "Jane Alvarez" and "jane.alvarez@" resolve to one surrogate identity once you normalize aliases. HIPAA's rules allow a re-identification code only if it is not derived from information about the individual and the mechanism is not disclosed, which is a useful design constraint even outside health data [3].
Decide the mapping scope deliberately. Item-level scope limits linkage across the suite but breaks multi-turn or multi-document items that span items; dataset-level scope keeps cross-item consistency but makes the whole suite one linkable unit. For RAG suites the mapping must cover the query, every corpus chunk and the gold citations together; the retrieval-preserving de-identification guide covers corpus-side details.
Surrogate pools need their own care. Draw names from a pool whose length, script and gender distribution resemble the originals so tokenization and agreement do not shift, and exclude surrogates that collide with real customers, staff or public figures mentioned elsewhere in the data. Keep surrogate emails on reserved domains such as example.com and phone numbers in ranges reserved for fiction, so a leaked item never points at a live person.
Transform the answer key in the same pass
The gold answer, rubric and any span offsets must go through the same mapping as the input, or the test grades against data that no longer exists. Exact-match and F1 graders fail when the expected answer still says "Jane Alvarez" while the prompt now says "Dana Whitfield." Span-based labels (NER, extractive QA, citation offsets) need recomputation, because surrogate lengths rarely match originals.
Rubric text written by domain experts is a frequent leak path. Reviewers often paste the customer name or ticket ID into the rationale, so run the detector over rubrics, grader prompts and judge few-shot examples, not just the item body. Store the mapping version alongside each item so a later re-run can reproduce exactly which surrogates were used.
Verify that the failure still reproduces
The acceptance test for de-identified eval data is behavioral: an item that exposed a model failure before transformation should expose the same failure after it. Run a fixed reference model (and ideally a second one) on the original and transformed versions inside the restricted environment, then compare outcomes item by item. Items that flip from fail to pass, or whose grader score moves beyond a tolerance you set, need manual review; this is a working check rather than a published standard.
Illustrative example: invented to show structure; it does not describe an available dataset.
Equivalence check record (one per item)
item_id: supp-coref-0142
properties: [identity_tracking, policy_logic]
mapping_scope: item
mapping_version: map-v3
transforms:
person_names: consistent_surrogate
account_number: format_preserving_luhn
dates: subject_date_shift
reference_model_runs:
original: {verdict: fail, score: 0.25}
transformed: {verdict: fail, score: 0.25}
delta_flag: false
residual_pii_scan: {detector: presidio, hits: 0, manual_sample_checked: true}
answer_key_transformed: true
reviewer: eval-ops-2
Aggregate the flags into a suite-level report: share of items whose verdict flipped, broken down by property and slice. A flip rate concentrated in one property (often coreference or retrieval) points at a transform bug rather than noise. Also recheck difficulty statistics, since an easier suite will show a jump in pass rate across every model you test.
Keep residual-risk checks proportionate to what the item reveals
Detection tools miss things, so plan a residual-risk check rather than trusting one pass. Presidio's own project notes that because it relies on trained models, there is no guarantee it finds all sensitive information [4], and research on LLM-based redaction reports measurable gaps and trade-offs as well [5]. Free-text fields in eval items, such as agent notes, quoted emails and stack traces with usernames, are where misses cluster.
Surrogates also do not remove quasi-identifiers. A rare diagnosis plus a small town plus an exact date can point to one person even with every name replaced, which is why NIST SP 800-188 stresses the limits of traditional de-identification and the governance around sharing [6]. If the eval suite will leave your environment or go to a vendor, test it against the stronger attacks described in our page on LLM-assisted re-identification risk.
Map the transform to the legal standard that applies
The legal bar is set by the data type and jurisdiction, not by the eval use case. For US health data, HIPAA de-identification means either Safe Harbor, which removes 18 listed identifiers including most date elements, or Expert Determination by a qualified expert [2][3]. Per-subject date shifting and some format-preserving surrogates do not fit Safe Harbor's list cleanly, so items that depend on exact intervals usually need Expert Determination.
Under the GDPR, Recital 26 places anonymous information outside the regulation only when the person is no longer identifiable by means reasonably likely to be used, and pseudonymized data with a retained key remains personal data [7]. A keyed surrogate mapping is pseudonymization in that sense for whoever holds the key. The full legal and methods playbook sits in de-identifying company data for AI.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Where raw traces and keys should live
Raw, un-transformed traces and the mapping key belong in a restricted system with its own access list, separate from where engineers run the eval suite [1]. That system is also where the equivalence runs happen, because they need both versions side by side. Log every access, and keep the key out of eval repos, notebooks and CI secrets that general staff can read.
When the source data comes from another company, the supplier may not let raw traces leave at all. In that case consider running the equivalence check, or the whole evaluation, on the supplier side; see evaluating models on data you can't take. If you are tempted to replace real items with generated ones instead, read the limits of synthetic evaluation data first, since synthetic items often lose the same properties redaction does.
A buyer checklist for de-identified eval sets from suppliers
When you license evaluation items from a supplier, ask for the de-identification method in writing and check it against the properties your suite needs. Useful requests:
- The transform per field type, and whether surrogates are consistent within an item, a document set or the full delivery.
- Whether answer keys, rubrics and citation spans were transformed in the same pass.
- The detector and version used, the manual sample size checked, and known gaps.
- Which legal standard the supplier applied for regulated fields (for example HIPAA Safe Harbor or Expert Determination).
- Whether a behavioral equivalence check was run, and the flip rate by property if so.
SourceX sources operational records from US companies, such as support histories, engineering records and finance or legal workflows, on request rather than from stock, and a request does not guarantee a match. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded, and a sample is checked, though no method is perfect. Buyers who need to specify which properties must survive can describe that in a buyer request. For building the suite itself, start from the evaluation datasets hub and the guide to golden datasets from business records.
Sourcing de-identified evaluation data
SourceX sources operational datasets from US companies on request and runs the commercial process from Find and Assess through Agree, Transact and Manage, with nothing contracted until a supplier agrees. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Describe the evaluation data and the properties it must preserve at sourcex.si/buyers.
Frequently asked questions
Is hashing names good enough for an eval set?
Usually not. A raw hash such as "a3f9c2" keeps consistency but produces tokens unlike real names, which shifts model behavior and can be brute-forced if the input space is small. A keyed mapping into realistic surrogates keeps consistency without those problems.
Should we de-identify the model outputs we store from eval runs?
Yes, if outputs were generated on raw items. Responses to un-transformed prompts often echo names and IDs, and they end up in dashboards and bug trackers with wider access than the restricted trace store.
Can we mix redacted and surrogate items in one suite?
You can, but report them as separate slices. Pass rates on redacted items are not comparable to surrogate items for identity-dependent properties.
Sources
- OneUptime, "How to Build a Golden Evaluation Dataset from Real LLM Production Failures" (2026). https://oneuptime.com/blog/post/2026-08-31-build-golden-evaluation-dataset-production-failures/markdown
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- Microsoft (microsoft/presidio), via pkg.go.dev, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- arXiv, "PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction" (2025). https://arxiv.org/pdf/2508.05545
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.