Skip to content

Tables, time series and transactional data

Evaluating Synthetic Tabular Data: Fidelity, Utility, Privacy and the Real Seed Tables Behind It

Quick answer

Evaluate synthetic tabular data on three axes, each against real data you control. Fidelity compares column distributions, column-pair relationships and, for multi-table data, keys and cardinalities. Utility trains a model on synthetic rows and tests it on a held-out real table (TSTR). Privacy checks, at minimum, that no synthetic row copies a real one. A single aggregate score hides failures, so acceptance should rest on per-column, per-table and task-level results, plus seed tables that actually contain the segments you need generated.

By SourceX Editorial · Updated

This page covers tabular and relational specifics. For synthetic data quality across modalities, see assessing synthetic data quality for AI training; for the wider cluster, start at the tabular, time-series and transactional data buyer's guide.

What a fidelity score actually measures

A fidelity score summarizes how closely synthetic marginals and pairwise relationships match the real table, and nothing more. SDMetrics, the open-source evaluation library from the Synthetic Data Vault project, is model-agnostic and groups its metrics into statistical, detection, efficacy (ML utility), likelihood and privacy families, applied at column, column-pair, single-table, multi-table and time-series granularity [1]. Its Quality Report rolls this into two properties, Column Shapes and Column Pair Trends, and averages them into one percentage [2].

Read the breakdown, not the headline. A 90% overall score can coexist with a badly reproduced transaction_amount tail or a broken region by product_line contingency, because averages across dozens of columns dilute single failures [2][3]. The Singapore government's synthetic-data guide makes the same point from the other direction: evaluation results must be interpretable by the people approving use, not just a number from a tool [5].

Three limits matter for buyers:

  • Pairwise is not structural. Column Pair Trends compares correlations (numeric) and contingency tables (categorical) two columns at a time. Research on tabular generators shows that matching marginals and pairs does not mean a generator has learned the dependency structure of the data, such as conditional relationships across three or more columns [4].
  • Ordered data is out of scope. The Quality Report is not optimized for sequential data; event logs, ledgers and per-customer histories need separate checks on ordering, gaps and inter-event times [2].
  • Business rules are invisible. A marginal score will not flag ship_date < order_date, negative quantities on non-return lines, or totals that no longer equal the sum of line items. Encode these as explicit constraints and count violations.

Utility: train on synthetic, test on real

Utility is measured by training the downstream model on synthetic data and scoring it on a real holdout that the generator never saw. This train-on-synthetic, test-on-real (TSTR) design is the efficacy family in SDMetrics terms [1], and the general fidelity-utility-privacy framing appears across practitioner guidance [6].

Run it as a comparison, not an absolute number:

  1. Split the real data before generation: a seed partition the generator learns from, and a holdout partition locked away. Split by entity (customer, account, site) and by time where the task is temporal, or leakage inflates both arms.
  2. Train the same model family with the same hyperparameter budget on (a) real seed rows (TRTR baseline) and (b) synthetic rows of equal count (TSTR).
  3. Score both on the real holdout using the task metric you will ship on: AUC-PR for rare-event classification, MAE or pinball loss for forecasting, calibration error where probabilities drive decisions.
  4. Report the gap per segment, not only overall. A small overall gap with a large gap on the minority class is a failure for fraud, churn or defect models.

The practical consequence: you cannot evaluate utility without licensed real data. If a vendor sells only synthetic output and cannot show TSTR against a holdout you trust, the utility claim is unverifiable. For building that holdout from operational records, see building a golden evaluation dataset from real business records, and for leakage pitfalls in the split itself, target leakage in licensed tabular data.

Privacy: copy detection is the floor, not the ceiling

The minimum privacy test is confirming that synthetic rows are not copies of real rows. SDMetrics' NewRowSynthesis metric measures the share of synthetic rows that do not exactly match a real row, with tolerance for numeric values [1]. Treat any exact or near-exact match on quasi-identifiers (ZIP3, birth year, job title, rare product codes) as a defect to investigate, even when the overall rate looks small.

Copy detection does not prove privacy. A generator can reproduce a rare combination, such as the only customer in a county with a given contract size, without copying a whole row, and that combination can be enough to re-identify someone. Add distance-to-closest-record comparisons against both the seed and the holdout set; if synthetic rows sit much closer to seed rows than holdout rows do, the generator has memorized. Where the seed tables held protected health information, remember that synthetic generation is not itself one of the two HIPAA de-identification methods (Expert Determination or Safe Harbor), and neither method removes all risk [7].

Multi-table checks: keys, cardinality and time

Relational synthetic data fails most often at the joins, so multi-table evaluation must test structure before statistics. SDMetrics exposes multi-table metrics alongside single-table ones [1], but buyers should insist on these checks regardless of tool:

  • Referential integrity. Every foreign key resolves to an existing parent (order_lines.order_id to orders.order_id); orphan rate should be zero unless the real data has orphans you chose to preserve.
  • Key validity. Primary keys are unique; composite keys hold; surrogate keys do not leak real identifiers.
  • Cardinality distributions. The distribution of children per parent (lines per order, tickets per account, visits per patient) matches the real shape, including the long tail of very large parents.
  • Cross-table consistency. Parent attributes agree with child aggregates: invoice_total equals the sum of line amounts within tolerance; account open_date precedes all its transactions.
  • Temporal ordering. Event sequences per entity keep valid order, plausible inter-arrival times and realistic state transitions (created, assigned, resolved, never resolved before created).

For what these relational tables look like in real enterprise systems, see multi-table relational data for relational deep learning and ERP transaction and master data for AI training.

Acceptance scorecard for a synthetic table or generator

An acceptance scorecard turns the metrics above into pass, fail or investigate decisions agreed before the data arrives. Thresholds depend on the task; the structure below is what matters.

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckGranularityMethodExample thresholdResult
Column shapesPer columnKS or total-variation complement vs. seedNo critical column below 0.85Investigate: refund_amount 0.71
Column pair trendsPer pairCorrelation and contingency similarityNamed pairs (e.g., plan_tier x churned) above 0.85Pass
Business constraintsPer rowRule set (dates, signs, totals)0 violations on hard rulesFail: 312 ship_date < order_date
Copy detectionPer rowNewRowSynthesis plus quasi-identifier match100% new rows; 0 unique QI matchesPass
MemorizationPer rowDistance to closest record, seed vs. holdoutSeed DCR not lower than holdout DCRPass
Referential integrityMulti-tableOrphan foreign keys0 orphansPass
CardinalityMulti-tableChildren-per-parent distributionTail (p99) within 20% of realInvestigate: p99 lines/order 9 vs. 41
Utility (TSTR)TaskSame model, real holdoutAUC-PR gap under 0.03 overall and per segmentFail: minority-region gap 0.11

Record the generator version, seed snapshot date, random seeds and evaluation code hash with each scorecard, so a regenerated dataset can be re-scored identically. Provenance records for synthetic training data covers what to log.

Seed tables bound what any generator can produce

A tabular generator cannot reliably produce segments, edge cases or dependencies that are absent or barely present in its real seed tables. Generators also differ in how much structure they learn from the same seed [4], so seed coverage and generator choice should be evaluated together. If the seed has 40 chargeback disputes, oversampling them synthetically yields variations of 40 cases, not a population.

When specifying real seed tables, ask for:

  • Segment coverage. Minimum counts for the rare classes, regions, product lines and time periods the model must handle, stated per table.
  • Full relational context. Parent and child tables with keys intact and a schema or data dictionary, so cardinality and integrity can be learned and checked. Table grain, keys and history lists what to specify.
  • Time span. Enough history to include seasonality, process changes and code-list revisions; see how much tabular data you need.
  • A separate holdout. Real rows withheld from the generator for TSTR and memorization tests, licensed for evaluation.
  • Explicit rights. Permission to train a generator on the tables and to use or distribute its output must be written into the license. Those terms are covered in combining licensed and synthetic data rather than restated here.

For background on trade-offs, see synthetic data vs. real business data and the synthetic data glossary entry.

Where SourceX fits in synthetic tabular projects

SourceX sources real operational datasets from US companies, which is the seed and holdout material synthetic evaluation depends on. Typical sources include support and sales histories, engineering records, and finance and legal workflow data. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery; personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. Data is sourced on request rather than held in stock, so describe the tables, segments and holdout you need on the buyer request page.

Get real seed and holdout tables for synthetic evaluation

If your synthetic tabular program needs licensed real tables to train a generator or to hold out for TSTR, describe the data rather than the businesses. SourceX looks for US companies that hold it, assesses data and licensing permissions, and nothing is contracted until a supplier agrees. Start a buyer request.

Sources

  1. DataCebo (Synthetic Data Vault), "SDMetrics". https://docs.sdv.dev/sdmetrics
  2. DataCebo (Synthetic Data Vault), "Quality Report". https://docs.sdv.dev/sdmetrics/reports/quality-report
  3. DataCebo (Synthetic Data Vault), "Data quality metrics". https://docs.sdv.dev/sdmetrics/data-metrics/quality
  4. arXiv, "How Well Does Your Tabular Generator Learn the Structure of Tabular Data?" (2025). https://arxiv.org/pdf/2503.09453
  5. GovTech Singapore, "Synthetic data: Quality evaluation". https://public.practice.gto.tech.gov.sg/data/content/publications/synthetic-data/chapters/quality-evaluation/
  6. Amazon Web Services, "How to evaluate the quality of the synthetic data: measuring from the perspective of fidelity, utility, and privacy". https://aws.amazon.com/blogs/machine-learning/how-to-evaluate-the-quality-of-the-synthetic-data-measuring-from-the-perspective-of-fidelity-utility-and-privacy
  7. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data