Evaluation and benchmarking datasets
Synthetic evaluation data: when generated test sets mislead
Quick answer
Synthetic evaluation data is acceptable when a model generates the inputs (paraphrases, perturbations, adversarial variants, coverage fillers) and a person or deterministic check verifies the expected output. It misleads when the same class of model writes both the question and the answer, when generated items stand in for the real traffic distribution, or when the generator shares training data with the system under test. For release gates, regression baselines and vendor bake-offs, anchor the eval set in real, verified records and treat generated items as a labeled supplement.
By SourceX Editorial · Updated
Why a generated answer is not ground truth
A model-generated reference answer is a second model output, not ground truth, until something independent of the generator verifies it [1]. When an LLM writes a question from a chunk and then writes the "correct" answer, any misreading, hallucinated number or wrong inference becomes the gold label. Your candidate system is then scored on agreement with the generator, not on correctness.
The failure is quiet. Ragas-style records hold question, contexts, answer and ground_truth, and its tooling can generate test sets, including reference answers, from your documents [2]. Nothing in the schema says whether a human or a check confirmed that field. Vendors now market "fact-checked" generation [3], which can reduce errors, but an automated fact-check is another model judgment unless it resolves to a source span, a database value or a reviewer sign-off.
Even human-labeled test sets carry errors. An audit of 10 widely used benchmarks estimated an average test-label error rate of at least 3.3% [6]. A generated set with no audit has no measured error rate at all, so treat a small reported gain with caution until the label error rate is known.
Where synthetic generation earns its place
Synthetic items are most useful when they widen coverage of a distribution you already understand, rather than define that distribution. Four uses hold up well in practice:
- Paraphrase and robustness probes. Rewrite a verified real question five ways (typos, different register, negation, a second language) and keep the verified answer. The truth is inherited, not generated.
- Coverage fillers for known gaps. If real tickets have only 12 examples of a refund edge case, generate variants from those examples to stress retrieval, then report them as a separate slice. Ragas documents this "test set from your documents" pattern [2].
- Adversarial and safety inputs. Prompt injections in retrieved documents, out-of-scope requests and jailbreak attempts are often rare in logs; generated inputs with rule-checked expected behavior ("refuse", "cite policy section", "escalate") are appropriate.
- Deterministic tasks. Arithmetic over a ledger, SQL against a fixture database, or a field extraction from a templated form, where a program computes the expected output.
The common thread is that the expected output comes from a verified real item, a program, or a reviewer, never from the generator alone.
Five ways synthetic test sets mislead
Synthetic eval sets mislead in predictable ways, and each has a measurable symptom. Check for all five before you let generated items drive a decision.
- Self-preference and generator-judge overlap. When the model family that generated items also scores them, or is the candidate, results can tilt toward that family. Research on self-bias in automated evaluation documents this effect in LLM-built benchmarks [4]. Symptom: one vendor's model wins on generated items and loses on real ones.
- Distribution mismatch with real traffic. Generated questions tend to be well formed, single-intent and answerable from one chunk. Real support tickets, sales emails and engineering issues are multi-intent, reference earlier threads, and carry missing context. This is a working hypothesis to test, not a constant: compare length, intent count and "unanswerable" rate between your generated set and a sample of logs.
- Answerability bias in RAG. Generators draw questions from chunks that exist, so nearly every item is answerable from the corpus. Production RAG must also handle questions whose answer is absent, stale or split across documents. A set with no unanswerable items cannot measure abstention.
- Contamination through shared sources. If the generator or the candidate was trained on the public documents you generate from, scores reflect recall. Benchmarks degrade as test data leaks into newer models' training sets, which is why LiveBench refreshes questions [7]. See contamination through synthetic data for the training-side mirror of this problem.
- False precision from small sets. A 150-item generated set feels large. CLT-based confidence intervals are too narrow below a few hundred datapoints [5], so differences of a few points between models are often not real.
A verification workflow for generated items
A generated item should enter a decision-grade eval set only after it passes source-grounding, answer verification and a reviewer sample. The workflow below is what eval leads typically run; adjust the thresholds to your risk.
- Tag provenance on every row. Add
item_origin(real, real_paraphrase, generated),generator_model,generator_prompt_hashandsource_doc_id. Data Cards-style documentation of upstream sources and annotation methods applies to eval sets as much as training sets [10]. - Ground the answer to a span. Require
evidence_spanwith character offsets into the source document. Reject items whose answer cannot be located verbatim or derived by a stated rule. - Run an independent check. Use a different model family, or a deterministic check, to answer the question from the span alone. Disagreement routes the item to a human.
- Human-review a sample. Have domain reviewers grade a random sample of passed items. Track agreement with Cohen's kappa, the same statistic used to calibrate LLM judges against human labels [8].
- Report slices separately. Publish scores for real items and generated items side by side. If the ranking of candidate models flips between slices, the generated slice is not measuring what you need.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"item_id": "ev-00417",
"item_origin": "real_paraphrase",
"parent_item_id": "ev-00121",
"generator_model": "model-b-2026-06",
"generator_prompt_hash": "sha256:9f2c...",
"source_doc_id": "kb/refund-policy-v7.pdf",
"evidence_span": {"start": 4182, "end": 4310},
"question": "if i returned it after 31 days do i still get store credit",
"expected_behavior": "answer",
"ground_truth": "Store credit is available for returns between 31 and 60 days.",
"verified_by": "reviewer_03",
"independent_check": "pass",
"slice": "refunds/late-return",
"answerable": true
}
Synthetic vs real evaluation data by decision type
The higher the cost of a wrong decision, the larger the real, verified share of the eval set must be. Use the table to set the mix before you generate anything.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Decision the eval drives | Generated items acceptable for | Real, verified data required for | Key check |
|---|---|---|---|
| Prompt iteration in development | Most items, with spot review | A small real anchor slice | Ranking stable on anchor slice |
| RAG retriever and chunking changes | Paraphrases, coverage fillers | Unanswerable and multi-document questions | Abstention rate vs logs |
| Release gate or regression baseline | Robustness probes only | Core scored set | Label audit, confidence intervals |
| Model vendor bake-off | Adversarial and safety inputs | All headline metrics | Generator differs from every candidate |
| Agent task success | Simulated user turns grounded in real cases | Task outcomes and end states | Outcome matches recorded result |
| Regulated or high-risk domains | Rarely | Nearly all items | Expert sign-off per item |
For a bake-off, the generator-judge overlap problem argues for a real held-out core; see running a model vendor bake-off with a private eval set. Agent evaluation has a parallel issue with simulated users; see user simulator scenarios grounded in real conversations.
What real data adds that generation cannot
Real operational records carry the distribution, the ambiguity and the outcome labels that a generator cannot know. A resolved support ticket has the customer's actual phrasing, the agent's reply, the macro used, and whether the case reopened. An invoice exception has the approver's decision. These outcomes are ground truth produced by the business, not by a model, which is the case made in outcome-labeled evaluation data.
Production failures are an especially strong seed. A golden set built from logged failures targets the cases your system actually gets wrong [9]. Building one from internal records is covered in golden evaluation datasets from business records, and the cluster map at LLM evaluation datasets shows where licensed test data fits next to public benchmarks.
Real data has costs of its own. Personal details must be removed without breaking what the test measures, and a third party's records need a license that covers evaluation use. If your logs are thin in a domain, licensed operational records from companies that hold them can fill the core; SourceX sources such datasets from US companies on request, and you can describe the eval data you need.
Requesting a mixed eval set from a supplier
When you commission or license eval data, specify the real and synthetic shares and the verification evidence in the request itself. Generic "eval set, 1,000 items" requests invite generated filler.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example entry |
|---|---|
| Task and decision | RAG answer quality for a billing support assistant; drives release gate |
| Real core | Resolved support tickets with final agent reply and reopen flag, US English |
| Unanswerable share | At least 15% of items with no answer in the provided corpus |
| Generated supplement | Paraphrases of verified real items only, tagged item_origin |
| Verification evidence | Evidence spans, reviewer IDs, kappa on a double-labeled sample |
| Provenance | Source system (for example Zendesk or Salesforce Service Cloud export), date range, preparation method |
| De-identification | Method recorded, with a checked sample |
| Allowed use | Evaluation of named models; term and delivery stated in the license |
Ask how the supplier would separate real from generated items in delivery, and run an acceptance gold-label audit before you rely on scores. For the training-side question of blending sources, see combining licensed and synthetic data and licensed vs synthetic vs scraped data. Ready-made use cases are described under AI evaluation datasets built from real business work and RAG evaluation datasets from real company documents.
Get real, verified evaluation data for your test sets
SourceX sources operational datasets such as support histories, engineering records and documents from US companies on request, and manages the commercial process, including the licensing agreement. Each dataset is rights-reviewed, personal details are removed or replaced before delivery with the method recorded, and a request does not guarantee a match. Describe the evaluation data you need.
Frequently asked questions
Can I use Ragas or a similar generator for my RAG test set?
Yes, for development and coverage, provided you verify groundtruth against evidence spans and add unanswerable items yourself [2]. Do not use an unreviewed generated set as the release gate.
Is LLM-as-a-judge scoring the same problem?
It is related but separate. A judge scores outputs; a generator creates the gold labels. Both need calibration against human labels [8], and the judge should not share a model family with the generator or the candidates [4].
How many real items do I need?
Enough that confidence intervals separate the models you are comparing. Below a few hundred items, standard CLT intervals understate uncertainty [5], so use alternatives such as Bayesian or Wilson score intervals and report them.
Sources
- OneUptime, "How to Build Ground Truth for RAG Evaluation When No Reference Answers Exist" (2026). https://oneuptime.com/blog/post/2026-08-31-build-rag-ground-truth-without-reference-answers/markdown
- Ragas, "Prepare data (Ragas v0.1.21 documentation)". https://docs.ragas.io/en/v0.1.21/getstarted/prepare_data.html
- Exa, "Generate fact-checked QA datasets automatically for RAG eval". https://tpc-websets.exa.ai/tool-generate-fact-checked-qa-datasets-automatically-rag-eval
- arXiv, "When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation" (2025). https://arxiv.org/pdf/2509.26600
- ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
- Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- ICLR 2025, "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2025). https://iclr.cc/virtual/2025/poster/28134
- OneUptime, "How to Calibrate an LLM-as-a-Judge Against Human Labels with Cohen's Kappa" (2026). https://oneuptime.com/blog/post/2026-08-31-calibrate-llm-judge-cohens-kappa/markdown
- OneUptime, "How to Build a Golden Evaluation Dataset from Real LLM Production Failures" (2026). https://oneuptime.com/blog/post/2026-08-31-build-golden-evaluation-dataset-production-failures/markdown
- Pushkarna, Zaldivar, Kjartansson, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.