Skip to content

Data quality, coverage and contamination

Assessing Synthetic Data Quality for AI Training: Fidelity, Diversity, Utility and Privacy

Quick answer

Evaluate synthetic data on four separate axes, never one blended score: fidelity (does it match the real distribution), diversity (does it cover the space or collapse onto the generator's favorite outputs), utility (does a model trained on it perform on held-out real data), and privacy (does it leak the seed records). Each axis needs its own test against a real reference set, and the axes trade off, so a set that maxes one usually sacrifices another [1][4].

By SourceX Editorial · Updated

This page is the modality-spanning framework, with extra depth on LLM-generated text and instruction data. For table-specific checks (column shapes, pair trends, SDMetrics reports), see evaluating synthetic tabular data. It sits inside the data quality, coverage and contamination hub.

Why a single synthetic quality score misleads

A single score misleads because fidelity, utility and privacy pull against each other, and an aggregate hides which one failed [4]. A generator that copies its seed rows will score near-perfect fidelity and fail privacy; a heavily noised generator passes privacy and loses the correlations a model needs. Research comparing 8 synthesizers on 12 datasets found that several widely used metrics behave poorly and proposed alternative fidelity, privacy and utility measures [2].

Utility scores carry a specific trap. The SynEval work warns that weak generators can produce inflated utility numbers and adds diversity as an explicit dimension [3]. One common cause is a downstream task easy enough that almost any data clears it. Treat a vendor's single "quality: 0.92" figure as a prompt to ask for the per-axis breakdown and the reference data it was computed against.

Fidelity: comparing synthetic and real distributions

Fidelity is measured by comparing the synthetic set to a real reference set, column by column and relationship by relationship. For structured fields, the standard checks are marginal distributions (Kolmogorov-Smirnov statistic for continuous fields, total variation distance for categoricals), category proportions and cardinalities, and pairwise correlations or contingency tables [1][5]. Libraries such as SDMetrics run these comparisons regardless of which generator produced the data [5].

For text, fidelity moves to properties you can count or classify:

  • Length and structure: token-length distribution, turn counts in dialogues, field presence in JSON outputs, schema validity rate.
  • Label and intent mix: distribution of intents, topics or ticket categories versus the real source, using the same classifier on both.
  • Domain vocabulary: rate of product codes, error strings, SKU formats or legal citations that look plausible but do not exist in the reference corpus.
  • Embedding distribution: distance between real and synthetic embedding clouds (for example, Fréchet-style distance or a real-versus-synthetic classifier; if a simple classifier separates them easily, fidelity is low).

The real-versus-synthetic classifier is the most practical single fidelity test for text. Train a logistic regression or small encoder to tell the two apart; an AUC near 0.5 means indistinguishable, and the misclassified examples show you what the generator gets right.

Diversity and the predictability trap in LLM-generated text

Synthetic text fails most often on diversity, because a generator samples from its own high-probability regions and under-produces rare cases. Repeated prompting of the same model yields near-paraphrases, a narrow set of openings ("Certainly, here is"), and uniform reasoning patterns. This is the predictability trap: each example looks fine, but the set is far more homogeneous than real data, and the model you train inherits that narrowness.

The long-run version of this is usually called model collapse: repeated training on unfiltered model output across generations erodes the tails of the original distribution. Even a single generation of synthetic data tends to over-represent the head of the distribution, which is exactly where your production edge cases are not.

Measure diversity on several levels, as covered in more depth in measuring dataset diversity:

  • Lexical: distinct-n (unique n-grams over total n-grams) and Self-BLEU across sampled pairs.
  • Semantic: mean pairwise embedding distance and cluster count at a fixed similarity threshold; near-duplicate rate using MinHash and LSH.
  • Task coverage: counts per intent, difficulty tier, tool, or language against a target taxonomy; empty cells are coverage gaps.
  • Tail coverage: share of examples in the real data's rarest 5% of clusters, compared with the real data's own share.

Utility: train-synthetic, test-real

Utility is measured by training on synthetic data and testing on held-out real data (TSTR), then comparing against a model trained and tested on real data (TRTR) [1]. The gap between TSTR and TRTR is the utility cost of using synthetic data. If you have no real held-out set, you cannot measure utility honestly; see using a licensed real-data holdout to validate synthetic data.

For LLM fine-tuning, run three arms: base model, base plus real SFT data, and base plus synthetic SFT data (and optionally a mix). Score all three on a real evaluation set that was never shown to the generator. Work on synthetic data for tool-using LLMs argues that generated instruction data needs explicit quality evaluation before training, because data quality shapes downstream behavior [6].

Two failure modes inflate utility. First, the evaluation set may itself be synthetic or generated by the same model, which rewards stylistic match; see when synthetic evaluation data misleads. Second, the generator may have seen benchmark items, so synthetic examples mirror test questions; check n-gram overlap against your evals as described in contamination through synthetic data.

Privacy: distance to closest record and membership inference

Privacy is assessed by asking whether synthetic records reveal specific seed records, using distance-based checks and membership-inference attacks [1]. Distance to closest record (DCR) measures how near each synthetic row sits to its nearest real training row; compare that distribution with DCR from a real holdout set to the training set. If synthetic rows are systematically closer than genuinely new real rows, the generator is memorizing.

Membership inference goes further: an attacker model tries to tell whether a given real record was in the generator's training data, using the synthetic output. Report attack AUC or advantage over a 0.5 baseline, on both typical records and outliers, because outliers leak first. For text, add three checks:

  • Verbatim overlap: longest common substring and long n-gram matches (for example 50+ tokens) between synthetic text and seed documents.
  • PII scan: run the same detectors you use on real data (names, emails, phone numbers, account numbers) on the synthetic output; LLM generators can reproduce identifiers from seeds or invent realistic ones.
  • Canaries: insert unique strings into the seed set before generation and count how often they surface.

Differential privacy changes the question from "did we detect leakage" to "what is the formal bound." DP-SGD fine-tuning of a generator, followed by sampling, can produce synthetic text with a stated epsilon, at a utility cost you should measure with TSTR [7]. An empirical DCR pass does not prove privacy; it only fails to find a specific leak. For memorization risk in the downstream model, see training-data extraction and memorization.

Factual consistency and correctness for synthetic text

Generated text needs a correctness check that tabular fidelity metrics do not provide. Instruction and Q&A data can be fluent, diverse and private while being wrong, and a model fine-tuned on confident errors learns them. Data-centric LLM training practice combines metric checks (format, duplication, diversity, fluency, factual accuracy) with LLM-based scoring and human review of samples [9].

Use verifiers wherever the domain allows them: unit tests for generated code, executed SQL against a real schema, recomputed arithmetic, and schema validation for tool-call arguments. Where no verifier exists, use an NLI model or an LLM judge to check each answer against its grounding document, then validate the judge on a human-labeled sample before trusting its pass rate.

Synthetic data acceptance scorecard

Illustrative example: invented to show structure; it does not describe an available dataset.

AxisTestReference setExample thresholdResultPass
FidelityReal-vs-synthetic classifier AUC5,000 real support tickets<= 0.700.64Yes
FidelityIntent distribution, total variation distanceSame<= 0.100.18No
DiversityDistinct-3 ratio, synthetic / realSame>= 0.850.61No
DiversityShare in real-data tail clustersSame>= 0.5x real share0.2xNo
UtilityTSTR macro-F1 gap vs TRTRHeld-out real set, never seen by generator<= 3 points2.1Yes
PrivacyDCR 5th percentile, synthetic vs real holdoutTraining seedsSynthetic not lowerLowerNo
PrivacyMembership-inference AUC (outliers)Training seeds<= 0.550.58No
PrivacyVerbatim 50-token matchesTraining seeds03No
CorrectnessVerifier or judge pass rate300 human-checked items>= 95%91%No

Read a table like this as a diagnosis, not a grade: here the set is useful on the main metric but homogeneous, skewed in intent mix and leaking near-copies, so the fix is generation-side (prompt diversification, seed rotation, deduplication against seeds) rather than more volume. Singapore GovTech's practice guidance similarly treats fidelity, utility and privacy as separate evaluations with their own methods [8].

What to request from a synthetic data vendor or pipeline owner

Ask for evidence per axis, computed against real data you can name. A minimal request:

  1. The generator model, version, prompts or templates, sampling settings (temperature, top-p) and the seed data description.
  2. Per-axis metrics with the reference set used, not an aggregate score.
  3. Deduplication method and thresholds, both within the synthetic set and against seeds.
  4. Privacy method: DP parameters if any, DCR and membership-inference results, PII scan results.
  5. Verification method and pass rate for correctness, with the human-checked sample size.
  6. Confirmation that no evaluation benchmark content was in the generator's prompts or seeds.

The weak point in most synthetic programs is the real reference data: without licensed real records you cannot measure fidelity, tail coverage, TSTR or DCR. The trade-offs between sources are covered in licensed vs synthetic vs scraped data and combining licensed and synthetic data; for definitions see the synthetic data glossary entry and synthetic data vs real business data. Teams that need a real seed or holdout set can describe the operational data they need to SourceX.

Source real data to anchor your synthetic evaluation

SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records and documents, and manages the licensing process; a request does not guarantee a match. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, with personal details removed or replaced before delivery. Describe the real reference data your synthetic pipeline needs.

Sources

  1. Amazon Web Services, "How to evaluate the quality of the synthetic data – measuring from the perspective of fidelity, utility, and privacy". https://aws.amazon.com/blogs/machine-learning/how-to-evaluate-the-quality-of-the-synthetic-data-measuring-from-the-perspective-of-fidelity-utility-and-privacy
  2. arXiv, "Synthetic data evaluation: critique of common metrics and new fidelity, privacy and utility measures (arXiv 2402.06806)" (2024). https://arxiv.org/html/2402.06806v1
  3. IEEE Computer Society Technical Community on Security and Privacy, "SynEval poster (IEEE Symposium on Security and Privacy 2026 posters)" (2026). https://www.ieee-security.org/TC/SP2026/downloads/posters/sp2026posters-final90.pdf
  4. ApX Machine Learning, "Taxonomy of evaluation metrics (Evaluating Synthetic Data Quality course)". https://www.apxml.com/courses/evaluating-synthetic-data-quality/chapter-1-foundations-synthetic-data-evaluation/taxonomy-evaluation-metrics
  5. DataCebo / Synthetic Data Vault (SDV), "SDMetrics documentation". https://docs.sdv.dev/sdmetrics
  6. arXiv, "Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs" (2024). https://arxiv.org/html/2409.16341v2
  7. arXiv, "Differentially Private Knowledge Distillation via Synthetic Text Generation" (2024). https://arxiv.org/html/2403.00932v2
  8. Government Technology Agency of Singapore (GovTech), "4. Quality Evaluation of Synthetic Data". https://public.practice.gto.tech.gov.sg/data/content/publications/synthetic-data/chapters/quality-evaluation/
  9. arXiv, "Towards Next-Generation LLM Training: From the Data-Centric Perspective" (2026). https://arxiv.org/pdf/2603.14712

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data