Data quality, coverage and contamination
Using a Licensed Real-Data Holdout to Validate Synthetic Training Data
Quick answer
To validate synthetic training data, train your model on the synthetic set and test it on a small, licensed set of real records that the generator has never seen. This "train synthetic, test real" (TSTR) check is the most direct test of whether the synthetic data transfers to the distribution you deploy into [1]. Size the real holdout by the slices you must report on rather than as a share of the synthetic corpus. Keep it firewalled from the generator, and audit its labels before you trust any score it produces.
By SourceX Editorial · Updated
Why distribution-match scores are not enough
Fidelity metrics tell you the synthetic data looks like its seed; only a real holdout tells you a model trained on it works. Practitioner guides split synthetic data quality into fidelity, utility and privacy, and they test utility by training on synthetic data and scoring on real data [1]. Libraries such as SDMetrics compare column shapes, pair trends and correlations between a synthetic table and the real table it was modeled on [3]. Those scores are useful, but they are computed against the generator's own seed, so a generator that copies its seed's blind spots still scores well.
Recent security research makes the same point from another angle: utility scores can mislead when there is no trustworthy real reference to anchor them [2]. The practical failure modes are familiar to anyone who has shipped a synthetic-trained model:
- Missing tails. The generator under-samples rare ticket types, unusual invoice layouts or edge-case transactions, and the model never learns them.
- Over-clean text. Generated support transcripts lack typos, pasted stack traces, mixed languages and half-finished sentences.
- Smoothed correlations. Joint distributions look right on two-way plots but break on three-way interactions such as region by product by season.
- Label-process drift. Synthetic labels follow a written policy; real labels follow what agents and adjudicators actually did.
The tabular foundation-model literature shows why the real anchor matters. TabPFN was pretrained only on synthetically generated tables, and the Real-TabPFN follow-up reported gains from adding a continued-pretraining stage on real-world tables [4]. For metric definitions, see our companion guide to assessing synthetic data quality; this page is about the purchase decision for the real reference set.
What a real holdout must contain
A useful real holdout mirrors your deployment distribution, carries trustworthy labels, and covers the slices you will be judged on. "Real" alone is not enough. A holdout drawn from one region, one product line or one quarter will tell you the synthetic-trained model works there and nowhere else.
Write the specification around four properties:
- Population match. Records come from the same kind of operation you deploy into: the same ticketing system (Zendesk, ServiceNow, Salesforce Service Cloud), the same document types, the same transaction channels.
- Slice coverage. Each segment you must report on, such as language, customer tier, document template or failure category, has enough records to estimate performance on its own. Use coverage gap analysis to list them.
- Label provenance. Outcome fields come from a recorded business decision (resolution code, approved or denied, final GL account), not from post-hoc annotation where avoidable. See verifying outcome labels in operational records.
- Time window. Records postdate anything the generator or its seed data saw, so you can split by time.
Label quality deserves its own budget. An audit of ten widely used test sets estimated an average label error rate of at least 3.3%, enough to reorder model rankings [5]. If your real holdout has a similar error rate and your synthetic-trained model is within a point or two of a baseline, the comparison is noise. Plan a double-review pass on a sample, and use inter-annotator agreement metrics where humans add labels.
How much real data you need
The holdout size is set by the narrowest slice you must measure and the precision you need, not by the size of the synthetic corpus. A binary metric's standard error is roughly the square root of p(1−p)/n. At 80% accuracy, 400 records in a slice give a standard error near 2 points, and 100 records give about 4 points. If you need to detect a 3-point gap between a synthetic-trained model and a real-trained baseline on a slice, 100 records will not do it.
That arithmetic turns the purchase into a slice budget. Multiply the records needed per slice by the number of slices, then add a margin for records that fail validation or de-identification checks. Rare slices dominate the total, which is why the long-tail and edge-case coverage question often decides the size.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Slice (support-ticket routing model) | Expected share in production | Target precision (±, 1 SE) | Records needed | Notes |
|---|---|---|---|---|
| English, standard tier | 55% | 2 pts | 400 | Easy to source |
| English, enterprise tier | 15% | 2 pts | 400 | Longer threads, attachments |
| Spanish, all tiers | 12% | 3 pts | 180 | Code-switching common |
| Billing disputes | 8% | 3 pts | 180 | Outcome = refund issued Y/N |
| Security incidents | 2% | 4 pts | 100 | Rare; oversample deliberately |
| Escalated to engineering | 8% | 3 pts | 180 | Linked Jira issue required |
| Total before 15% attrition margin | 1,440 | ~1,660 with margin |
The table shows the shape of the decision: about 1,700 real records can anchor a model trained on hundreds of thousands of synthetic tickets, but only because each slice was sized on purpose. If your budget cannot cover a slice, say so in the evaluation report rather than pooling it into an aggregate score.
Keeping the holdout out of the generator
A real holdout is worthless the moment any of its records, or near-copies of them, reach the synthetic generator. Leakage paths are mundane: the same export seeds both the generator and the test set, a prompt template quotes real examples, or a later generator refresh pulls from a shared bucket. Temporal leakage is the time-series variant, where future information is used to predict the past, and time-based splits are the standard defense [6].
Controls that work in practice:
- Separate storage and access. Put the holdout in its own bucket or project with a distinct IAM role; the generator pipeline's service account has no read permission.
- Hash registry. Store record IDs and content hashes for the holdout; check every generator seed and every synthetic output batch against them, including near-duplicate checks with MinHash and LSH.
- Time split. Seed the generator only with records before a cutoff date; draw the holdout only from after it.
- Memorization probe. Search synthetic outputs for long verbatim spans from the holdout; any hit means the firewall failed.
- Freeze and version. Version the holdout (for example,
holdout_v1_2026-09) and never edit it in place, so scores remain comparable across generator releases.
Our guide to contamination through synthetic data covers the reverse problem, where generated examples mirror public benchmarks.
Licensing a reference set rather than a training set
A real holdout is a narrower purchase than a training corpus, and the license should say exactly that. Generic buyer advice: ask for terms that name evaluation and validation as the permitted use, state whether the records may also seed or condition a generator, and specify retention and deletion. Whether a supplier offers evaluation-only terms, and at what price, is negotiated deal by deal; see can I license data for evaluation only?. Teams that need operational records from US businesses for this purpose can describe the reference set on the SourceX buyer page.
Decide up front whether the holdout may ever become training data. If it does, you lose your independent measure and need a fresh holdout; some teams therefore license a second, later tranche for that purpose. Record the decision in a training data use register so nobody repurposes the set by accident.
Illustrative example: invented to show structure; it does not describe an available dataset.
reference_set_request:
purpose: "TSTR validation of synthetic-trained ticket routing model"
permitted_use: [evaluation, error_analysis]
prohibited_use: [model_training, generator_seeding, generator_conditioning]
records: 1700
record_unit: "closed support ticket with full thread"
date_window: "2026-01-01 to 2026-06-30" # after generator seed cutoff
required_fields: [ticket_id, created_at, channel, language, tier,
thread_text, final_category, resolution_code, escalated_to]
slices_min_counts: {security_incident: 100, billing_dispute: 180}
label_source: "final_category as set at ticket close"
deidentification: "names, emails, phones, account numbers removed or replaced; method documented"
delivery: "access-controlled transfer; no email attachments"
Privacy and regulatory checks on the real side
The real holdout carries the privacy and compliance obligations that the synthetic set was meant to avoid, so plan for them. Plan for personal details to be removed or replaced before delivery, and health records need HIPAA de-identification under Safe Harbor or Expert Determination [9]. De-identification can also change the text your model sees, for example by replacing names with tokens, so run your synthetic generator's placeholder conventions through the same transformation before comparing.
For high-risk systems under the EU AI Act, Article 10 requires training, validation and testing data sets to meet quality criteria, and a documented real test set can help show that your testing data meets them [7]. As of October 2026, Regulation (EU) 2026/1744 has reportedly moved the high-risk application dates to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems, and also amends Article 10 [10]. ISO/IEC 5259-3 gives a data quality management process frame for documenting how you assess and maintain data quality across both synthetic and real sets [8]. See applying EU AI Act data quality criteria to purchased datasets for the full mapping.
Running and reading the TSTR comparison
Run three models against the same real holdout, so the synthetic result has something to be compared with. Train model A on synthetic data only, model B on whatever real data you already hold (or a small real training split), and model C on synthetic plus that real data. Score all three on the frozen real holdout, slice by slice.
Reading the results:
- A close to B on every slice: the synthetic data transfers; keep generating.
- A close to B overall but far behind on one slice: the generator misses that slice; target it with conditioned generation or real records.
- C clearly above both: the real data adds signal the generator lacks; this is the case for combining licensed and synthetic data.
- A above B: check for leakage before celebrating.
Report confidence intervals per slice, not just point estimates, and rerun the comparison every time the generator, its seed data or its prompts change.
Sourcing a real reference set for synthetic data validation
SourceX sources operational datasets, such as support and sales histories, engineering records, documents and finance workflows, from US companies on request, so you describe the real reference data you need rather than picking from stock. Every dataset is rights-reviewed, personal details are removed or replaced before delivery, and the license defines the records, permitted uses, term and delivery; nothing is contracted until the supplying company agrees. Describe the real holdout you need to validate your synthetic data.
More guides on data quality, coverage and contamination are in the cluster hub, and the full library is at AI data for buyers.
Frequently asked questions
Can I use the same real data to seed the generator and test it?
No. If the generator saw the records, the test measures memorization, not transfer. Split by time or by entity before seeding, and keep the test portion in separate storage.
Is a real holdout still needed if fidelity scores are high?
Yes. Fidelity scores compare synthetic data to its own seed [3], so they cannot reveal gaps the seed shares. Only performance on independent real records shows whether the model works in deployment [1].
How often should the real holdout be refreshed?
Refresh it when your production distribution shifts, for example after a product launch, a new ticket taxonomy or a code-set revision. Keep the old version for trend comparisons and see label definition changes in multi-year datasets.
Sources
- Amazon Web Services (AWS Machine Learning Blog), "How to evaluate the quality of the synthetic data: measuring from the perspective of fidelity, utility, and privacy". https://aws.amazon.com/blogs/machine-learning/how-to-evaluate-the-quality-of-the-synthetic-data-measuring-from-the-perspective-of-fidelity-utility-and-privacy
- IEEE Symposium on Security and Privacy 2026, "IEEE S&P 2026 poster on synthetic data utility evaluation" (2026). https://www.ieee-security.org/TC/SP2026/downloads/posters/sp2026posters-final90.pdf
- DataCebo / The Synthetic Data Vault, "SDMetrics". https://docs.sdv.dev/sdmetrics
- arXiv, "Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data" (2025). https://arxiv.org/pdf/2507.03971
- Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Hyndman and Athanasopoulos, "Forecasting: Principles and Practice (3rd Edition)". https://otexts.com/fpp3/tscv.html
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-3:2024 Data quality for analytics and machine learning, Part 3: Data quality management requirements and guidelines" (2024). https://www.iso.org/standard/81092.html
- eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- European Parliament and Council of the European Union (EUR-Lex), "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.