Skip to content

Evaluation and benchmarking datasets

Synthetic evaluation data: when generated test sets mislead

Quick answer

Synthetic evaluation data is acceptable when a model generates the inputs (paraphrases, perturbations, adversarial variants, coverage fillers) and a person or deterministic check verifies the expected output. It misleads when the same class of model writes both the question and the answer, when generated items stand in for the real traffic distribution, or when the generator shares training data with the system under test. For release gates, regression baselines and vendor bake-offs, anchor the eval set in real, verified records and treat generated items as a labeled supplement.

By SourceX Editorial · Updated

Why a generated answer is not ground truth

A model-generated reference answer is a second model output, not ground truth, until something independent of the generator verifies it [1]. When an LLM writes a question from a chunk and then writes the "correct" answer, any misreading, hallucinated number or wrong inference becomes the gold label. Your candidate system is then scored on agreement with the generator, not on correctness.

The failure is quiet. Ragas-style records hold question, contexts, answer and ground_truth, and its tooling can generate test sets, including reference answers, from your documents [2]. Nothing in the schema says whether a human or a check confirmed that field. Vendors now market "fact-checked" generation [3], which can reduce errors, but an automated fact-check is another model judgment unless it resolves to a source span, a database value or a reviewer sign-off.

Even human-labeled test sets carry errors. An audit of 10 widely used benchmarks estimated an average test-label error rate of at least 3.3% [6]. A generated set with no audit has no measured error rate at all, so treat a small reported gain with caution until the label error rate is known.

Where synthetic generation earns its place

Synthetic items are most useful when they widen coverage of a distribution you already understand, rather than define that distribution. Four uses hold up well in practice:

  • Paraphrase and robustness probes. Rewrite a verified real question five ways (typos, different register, negation, a second language) and keep the verified answer. The truth is inherited, not generated.
  • Coverage fillers for known gaps. If real tickets have only 12 examples of a refund edge case, generate variants from those examples to stress retrieval, then report them as a separate slice. Ragas documents this "test set from your documents" pattern [2].
  • Adversarial and safety inputs. Prompt injections in retrieved documents, out-of-scope requests and jailbreak attempts are often rare in logs; generated inputs with rule-checked expected behavior ("refuse", "cite policy section", "escalate") are appropriate.
  • Deterministic tasks. Arithmetic over a ledger, SQL against a fixture database, or a field extraction from a templated form, where a program computes the expected output.

The common thread is that the expected output comes from a verified real item, a program, or a reviewer, never from the generator alone.

Five ways synthetic test sets mislead

Synthetic eval sets mislead in predictable ways, and each has a measurable symptom. Check for all five before you let generated items drive a decision.

  1. Self-preference and generator-judge overlap. When the model family that generated items also scores them, or is the candidate, results can tilt toward that family. Research on self-bias in automated evaluation documents this effect in LLM-built benchmarks [4]. Symptom: one vendor's model wins on generated items and loses on real ones.
  2. Distribution mismatch with real traffic. Generated questions tend to be well formed, single-intent and answerable from one chunk. Real support tickets, sales emails and engineering issues are multi-intent, reference earlier threads, and carry missing context. This is a working hypothesis to test, not a constant: compare length, intent count and "unanswerable" rate between your generated set and a sample of logs.
  3. Answerability bias in RAG. Generators draw questions from chunks that exist, so nearly every item is answerable from the corpus. Production RAG must also handle questions whose answer is absent, stale or split across documents. A set with no unanswerable items cannot measure abstention.
  4. Contamination through shared sources. If the generator or the candidate was trained on the public documents you generate from, scores reflect recall. Benchmarks degrade as test data leaks into newer models' training sets, which is why LiveBench refreshes questions [7]. See contamination through synthetic data for the training-side mirror of this problem.
  5. False precision from small sets. A 150-item generated set feels large. CLT-based confidence intervals are too narrow below a few hundred datapoints [5], so differences of a few points between models are often not real.

A verification workflow for generated items

A generated item should enter a decision-grade eval set only after it passes source-grounding, answer verification and a reviewer sample. The workflow below is what eval leads typically run; adjust the thresholds to your risk.

  1. Tag provenance on every row. Add item_origin (real, real_paraphrase, generated), generator_model, generator_prompt_hash and source_doc_id. Data Cards-style documentation of upstream sources and annotation methods applies to eval sets as much as training sets [10].
  2. Ground the answer to a span. Require evidence_span with character offsets into the source document. Reject items whose answer cannot be located verbatim or derived by a stated rule.
  3. Run an independent check. Use a different model family, or a deterministic check, to answer the question from the span alone. Disagreement routes the item to a human.
  4. Human-review a sample. Have domain reviewers grade a random sample of passed items. Track agreement with Cohen's kappa, the same statistic used to calibrate LLM judges against human labels [8].
  5. Report slices separately. Publish scores for real items and generated items side by side. If the ranking of candidate models flips between slices, the generated slice is not measuring what you need.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "item_id": "ev-00417",
  "item_origin": "real_paraphrase",
  "parent_item_id": "ev-00121",
  "generator_model": "model-b-2026-06",
  "generator_prompt_hash": "sha256:9f2c...",
  "source_doc_id": "kb/refund-policy-v7.pdf",
  "evidence_span": {"start": 4182, "end": 4310},
  "question": "if i returned it after 31 days do i still get store credit",
  "expected_behavior": "answer",
  "ground_truth": "Store credit is available for returns between 31 and 60 days.",
  "verified_by": "reviewer_03",
  "independent_check": "pass",
  "slice": "refunds/late-return",
  "answerable": true
}

Synthetic vs real evaluation data by decision type

The higher the cost of a wrong decision, the larger the real, verified share of the eval set must be. Use the table to set the mix before you generate anything.

Illustrative example: invented to show structure; it does not describe an available dataset.

Decision the eval drivesGenerated items acceptable forReal, verified data required forKey check
Prompt iteration in developmentMost items, with spot reviewA small real anchor sliceRanking stable on anchor slice
RAG retriever and chunking changesParaphrases, coverage fillersUnanswerable and multi-document questionsAbstention rate vs logs
Release gate or regression baselineRobustness probes onlyCore scored setLabel audit, confidence intervals
Model vendor bake-offAdversarial and safety inputsAll headline metricsGenerator differs from every candidate
Agent task successSimulated user turns grounded in real casesTask outcomes and end statesOutcome matches recorded result
Regulated or high-risk domainsRarelyNearly all itemsExpert sign-off per item

For a bake-off, the generator-judge overlap problem argues for a real held-out core; see running a model vendor bake-off with a private eval set. Agent evaluation has a parallel issue with simulated users; see user simulator scenarios grounded in real conversations.

What real data adds that generation cannot

Real operational records carry the distribution, the ambiguity and the outcome labels that a generator cannot know. A resolved support ticket has the customer's actual phrasing, the agent's reply, the macro used, and whether the case reopened. An invoice exception has the approver's decision. These outcomes are ground truth produced by the business, not by a model, which is the case made in outcome-labeled evaluation data.

Production failures are an especially strong seed. A golden set built from logged failures targets the cases your system actually gets wrong [9]. Building one from internal records is covered in golden evaluation datasets from business records, and the cluster map at LLM evaluation datasets shows where licensed test data fits next to public benchmarks.

Real data has costs of its own. Personal details must be removed without breaking what the test measures, and a third party's records need a license that covers evaluation use. If your logs are thin in a domain, licensed operational records from companies that hold them can fill the core; SourceX sources such datasets from US companies on request, and you can describe the eval data you need.

Requesting a mixed eval set from a supplier

When you commission or license eval data, specify the real and synthetic shares and the verification evidence in the request itself. Generic "eval set, 1,000 items" requests invite generated filler.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample entry
Task and decisionRAG answer quality for a billing support assistant; drives release gate
Real coreResolved support tickets with final agent reply and reopen flag, US English
Unanswerable shareAt least 15% of items with no answer in the provided corpus
Generated supplementParaphrases of verified real items only, tagged item_origin
Verification evidenceEvidence spans, reviewer IDs, kappa on a double-labeled sample
ProvenanceSource system (for example Zendesk or Salesforce Service Cloud export), date range, preparation method
De-identificationMethod recorded, with a checked sample
Allowed useEvaluation of named models; term and delivery stated in the license

Ask how the supplier would separate real from generated items in delivery, and run an acceptance gold-label audit before you rely on scores. For the training-side question of blending sources, see combining licensed and synthetic data and licensed vs synthetic vs scraped data. Ready-made use cases are described under AI evaluation datasets built from real business work and RAG evaluation datasets from real company documents.

Get real, verified evaluation data for your test sets

SourceX sources operational datasets such as support histories, engineering records and documents from US companies on request, and manages the commercial process, including the licensing agreement. Each dataset is rights-reviewed, personal details are removed or replaced before delivery with the method recorded, and a request does not guarantee a match. Describe the evaluation data you need.

Frequently asked questions

Can I use Ragas or a similar generator for my RAG test set?

Yes, for development and coverage, provided you verify groundtruth against evidence spans and add unanswerable items yourself [2]. Do not use an unreviewed generated set as the release gate.

Is LLM-as-a-judge scoring the same problem?

It is related but separate. A judge scores outputs; a generator creates the gold labels. Both need calibration against human labels [8], and the judge should not share a model family with the generator or the candidates [4].

How many real items do I need?

Enough that confidence intervals separate the models you are comparing. Below a few hundred items, standard CLT intervals understate uncertainty [5], so use alternatives such as Bayesian or Wilson score intervals and report them.

Sources

  1. OneUptime, "How to Build Ground Truth for RAG Evaluation When No Reference Answers Exist" (2026). https://oneuptime.com/blog/post/2026-08-31-build-rag-ground-truth-without-reference-answers/markdown
  2. Ragas, "Prepare data (Ragas v0.1.21 documentation)". https://docs.ragas.io/en/v0.1.21/getstarted/prepare_data.html
  3. Exa, "Generate fact-checked QA datasets automatically for RAG eval". https://tpc-websets.exa.ai/tool-generate-fact-checked-qa-datasets-automatically-rag-eval
  4. arXiv, "When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation" (2025). https://arxiv.org/pdf/2509.26600
  5. ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  6. Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  7. ICLR 2025, "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2025). https://iclr.cc/virtual/2025/poster/28134
  8. OneUptime, "How to Calibrate an LLM-as-a-Judge Against Human Labels with Cohen's Kappa" (2026). https://oneuptime.com/blog/post/2026-08-31-calibrate-llm-judge-cohens-kappa/markdown
  9. OneUptime, "How to Build a Golden Evaluation Dataset from Real LLM Production Failures" (2026). https://oneuptime.com/blog/post/2026-08-31-build-golden-evaluation-dataset-production-failures/markdown
  10. Pushkarna, Zaldivar, Kjartansson, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data