Skip to content

Evaluation and benchmarking datasets

Writing an evaluation dataset specification: fields, slices, labels and acceptance

Quick answer

An evaluation dataset specification is the document a supplier prices and builds against. It should fix six things: the deployed task the set measures, a record schema that captures the full context the model saw, the gold format used to score each record, named slices with target counts, provenance and de-identification rules per record, and acceptance tests with rejection and replacement terms. If any one is vague, you will pay for records you cannot use.

By SourceX Editorial · Updated

This page covers the eval-specific parts of a request. For the general structure of a supplier request, see how to write a data request for suppliers. For where custom eval sets fit next to public benchmarks, start at the evaluation datasets hub.

Scope: tie the set to one deployed task and wall it off from training

The scope statement names the task the deployed system performs and states that no record may appear in any training, fine-tuning or retrieval-index corpus. Guidance on eval construction stresses both points: scope the dataset to the tasks the model will actually perform, and keep it distinct from training data [1]. A supplier cannot infer either from a list of topics.

Write the scope as a short paragraph covering four things:

  • Task. For example, "answer policy questions for tier-1 support agents using the attached knowledge base snapshot" or "extract remittance fields from vendor invoices into the AP schema."
  • System boundary. Say whether you are scoring the model alone, the retriever plus the model, or a multi-step agent with tools.
  • Exclusions. List adjacent tasks the set should not measure, such as escalation routing.
  • Held-out handling. Records live only in the eval store. The supplier does not reuse them in other deliveries, and you record a hash of each input so later contamination checks have something to match against.

Contamination design deserves its own treatment; see contamination-resistant evaluation design and private eval sets vs public benchmarks.

Record schema: capture every input that shaped the output

An eval record must hold everything the system saw at decision time, not only the user's message. Practitioner guidance on golden datasets recommends storing retrieved documents, tool results and prior turns with each case, because a failure often cannot be reproduced from the final prompt alone [2]. For RAG, the Ragas v0.1 documentation lists question, contexts, answer and ground_truth as the fields of a test record [4].

Specify fields by type and source system, not just by name. A "context" field should say whether it holds document IDs pointing into a frozen corpus snapshot or the full chunk text, and which version of the knowledge base it came from. For agents, the agent evaluation task suites page covers environment state; τ-bench is a useful public reference because it pairs simulated users with programmatic APIs, a database and a domain policy document [5].

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "eval-ap-000417",
  "task": "invoice_field_extraction",
  "input": {
    "document_ref": "corpus_v3/doc_88213.pdf",
    "page_count": 2,
    "instructions_version": "ap-prompt-2026-09"
  },
  "context": {
    "prior_turns": [],
    "tool_results": [{"tool": "vendor_lookup", "output_ref": "tools/88213_vendor.json"}],
    "corpus_snapshot": "kb-2026-08-31"
  },
  "gold": {
    "format": "field_exact",
    "fields": {"invoice_number": "INV-20931", "total_amount": "4812.50", "currency": "USD", "due_date": "2026-07-15"},
    "evidence_spans": [{"field": "total_amount", "page": 2, "bbox": [412, 690, 520, 708]}]
  },
  "slices": ["layout:multi_page", "vendor_tier:long_tail", "risk:credit_note"],
  "provenance": {"source_system": "ERP AP module", "record_created": "2025-11-03", "label_source": "posted_ledger_entry"},
  "deid": {"method": "replace_names_accounts", "applied_at": "2026-09-12"},
  "split": "eval_only"
}

Agree a container format up front. JSONL with one record per line is easy to diff; Parquet suits large tabular context. A Croissant JSON-LD descriptor can document the fields, files and record structure in a machine-readable way [8].

Gold format: choose how each record will be scored

The gold format defines what "correct" means for a record, and it decides whether your grader can be automated. Four formats cover most enterprise tasks:

Gold formatBest forWhat the supplier deliversMain failure mode
Exact answer or field valuesExtraction, classification, routingNormalized target values plus normalization rulesFormatting disputes (dates, currency, units)
Required and forbidden claimsRAG answers, summaries, policy repliesClaims that must appear, claims that must not, evidence spansClaims written too broadly to fail anything
Rubric scoresOpen generation, tone, multi-criteria qualityRubric, per-criterion scores, rater IDsLow agreement between raters
End-state checksTool-using agentsExpected database or system state after the episodeSeveral valid paths lead to different states

Required and forbidden claims let you score answers that are phrased differently but say the same thing, and evidence spans make each claim checkable against the source [3]. Rubric design is covered in evaluation rubrics with domain experts. If an LLM judge will grade the rubric, budget for a judge calibration set of human labels.

Where labels come from business outcomes, such as a posted ledger entry or a final claim adjudication, say which field is the label and how reversals are handled. Outcome-labeled evaluation data explains when those fields are reliable.

Slices: define them as tags with counts, not as themes

Slices are named, mutually understood tags with a target count each, so you can report results per segment rather than one blended score. Define each slice with a rule a supplier can apply to a source record, such as "document layout has more than one page" or "ticket was reopened after first resolution." A slice like "hard cases" is not a rule.

Plan three kinds of slices:

  • Coverage slices reflect the production mix: channel, product line, document type, language.
  • Risk slices oversample rare, high-cost cases: credit notes, regulated disclosures, refunds over a threshold. See stratified evaluation sets.
  • Regression slices hold known past failures, which practitioner guides recommend promoting into the golden set once they are triaged [2].

Allow records to carry several tags, but name one primary slice for allocation so counts add up.

Size: work back from the smallest difference you need to detect

Set slice sizes from the smallest change in a metric you need to detect, then add a margin for records rejected at acceptance. A set sized only by budget often produces confidence intervals wider than the gap between the two systems you are comparing. That makes a model-selection decision a coin toss.

A practical method is to state the decision the set must support ("ship the new retriever if correct-answer rate improves by at least five points on the policy slice"), estimate the baseline rate from a pilot, and compute the per-slice count needed for that difference. Paired comparisons on the same records need fewer examples than unpaired ones, which favors fixed eval sets over fresh samples each run. Write the per-slice minimum into the spec, along with a cap on how much any slice may exceed its target.

Provenance, timestamps and de-identification per record

Each record should state its source system, the date the underlying event occurred, the date it was labeled, and the de-identification method applied. Timestamps let you separate pre- and post-cutoff records when you test for contamination and let you drop records tied to a retired policy version. Provenance fields also tell you whether a label came from a business outcome or a post-hoc annotator.

Specify which identifiers must be removed or replaced, and how replacements stay consistent within a record so that a multi-turn conversation still reads coherently. For health data, the HIPAA rule at 45 CFR 164.514 sets the de-identification standard and its two methods, Safe Harbor and Expert Determination [9]. For broader context on source and rights fields, see the data provenance guide.

Acceptance: audit labels, set thresholds and agree replacement terms

Acceptance rules turn the spec into something both sides can verify. Write them before the build starts, because a supplier prices the rework risk they imply. Label errors in test sets are common even in widely used public benchmarks and can be enough to change which model ranks first [6], so an audit is not optional.

A workable acceptance section names:

  1. Schema validation. Every record parses, required fields are non-null, slice tags come from the controlled list, and referenced files exist.
  2. Audit sample. A random sample per slice, re-labeled independently by your reviewers.
  3. Agreement threshold. A minimum agreement rate or coefficient between supplier and auditor labels. Pick the agreement metric to fit the label type, since different metrics suit categorical, ordinal and free-text judgments [7].
  4. Rejection rule. What happens when a slice fails: whole-slice rework, or replacement of failed records only.
  5. Replacement terms. Who supplies replacements, how they are re-audited, and how disputed labels go to adjudication.
  6. Contamination check. Hashes of inputs compared against your training corpora and any public sources you can scan.

The detailed audit workflow, including adjudication, is on accepting a delivered eval set.

Spec checklist a supplier can price

A supplier can only quote a spec that answers each of these questions. Use this checklist before you send a request.

Illustrative example: invented to show structure; it does not describe an available dataset.

SectionQuestion the spec must answerExample entry
ScopeWhich deployed task and system boundary?Tier-1 policy Q&A, retriever plus model
Held-outWhere may records never appear?Training, fine-tuning, retrieval index
SchemaWhich fields, types and source systems?Input, context refs, tool results, gold, slices
ContextFrozen corpus version?kb-2026-08-31 snapshot
GoldWhich scoring format per task?Required and forbidden claims with evidence spans
SlicesRule and target count per slice?Refund over threshold: 150 records
SizeWhat difference must be detectable?Five points on the policy slice
ProvenanceEvent date, label date, label source?Ticket closed date, auditor ID
De-identificationWhich identifiers, which method?Names and account numbers replaced consistently
AcceptanceAudit sample, threshold, rejection, replacement?Per-slice sample, agreed metric, record-level replacement
FormatContainer and metadata?JSONL plus Croissant descriptor
LicensePermitted uses and publication?Internal eval only; see license page

License terms for eval data, including whether you may publish scores, are covered in evaluation-only data license terms and publishing results on licensed eval data. Cost drivers are on what a custom evaluation dataset costs.

How SourceX handles evaluation dataset specifications

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents and finance and legal workflows; nothing is held in stock, and a request does not guarantee a match. You describe the data your spec needs, not the businesses, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. See how the process works for AI teams buying evaluation data, or read about evaluation sets built from real business work.

Send your evaluation dataset specification

If your spec calls for real business records as eval inputs and labels, SourceX can look for US companies that hold the described data and manage the license from assessment through agreement and ongoing purchases. Nothing is contracted until a supplier agrees. Describe the evaluation data you need.

Sources

  1. arXiv, "A Practical Guide for Evaluating LLMs and LLM-Reliant Systems (arXiv:2506.13023)" (2025). https://arxiv.org/html/2506.13023v1
  2. OneUptime, "How to Build a Golden Evaluation Dataset from Real LLM Production Failures" (2026). https://oneuptime.com/blog/post/2026-08-31-build-golden-evaluation-dataset-production-failures/markdown
  3. OneUptime, "How to Build Ground Truth for RAG Evaluation When No Reference Answers Exist" (2026). https://oneuptime.com/blog/post/2026-08-31-build-rag-ground-truth-without-reference-answers/markdown
  4. Ragas, "Prepare your test dataset (Ragas v0.1.21 documentation)". https://docs.ragas.io/en/v0.1.21/getstarted/prepare_data.html
  5. arXiv (Sierra Research), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (arXiv:2406.12045)" (2024). https://export.arxiv.org/pdf/2406.12045
  6. arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks (arXiv:2103.14749)" (2021). https://arxiv.org/abs/2103.14749
  7. arXiv, "Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation (arXiv:2603.06865)" (2026). https://arxiv.org/pdf/2603.06865
  8. arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets (arXiv:2403.19546)" (2024). https://arxiv.org/pdf/2403.19546
  9. eCFR (Office of the Federal Register / HHS), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data