Skip to content

Evaluation and benchmarking datasets

Choosing an LLM provider with your own eval set: running a model bake-off

Quick answer

To choose an LLM provider on evidence, run every shortlisted model on the same private, held-out eval set built from your own workflows. Score paired results with confidence intervals, broken out by use-case slice. Weigh quality against cost per resolved task, latency, refusal behavior and run-to-run consistency, and lock the protocol before any vendor sees an item. Vendor leaderboard numbers are a starting filter, not a decision. Your bake-off result should be reproducible enough to defend in a contract negotiation.

By SourceX Editorial · Updated

Why vendor-published scores do not settle the decision

Published benchmark scores answer a different question from yours, and some are inflated. Research on contamination shows benchmark items can leak into training data and inflate public scores relative to fresh held-out tests [1]. Public suites also lack the task mix, document types and policies of a specific enterprise workflow [2]. OpenAI stopped reporting SWE-bench Verified in 2026 because it judged the benchmark increasingly contaminated [3].

Even eval tooling vendors frame public benchmarks as useful for rough comparison while products need custom evaluations [4]. For procurement, that means a model that leads MMLU-style tables can still lose on your claims-adjudication notes, your SQL dialect or your support macros. The comparison of private eval sets and public benchmarks covers when held-out data becomes necessary; this page assumes you have reached that point and need to pick a vendor.

What the bake-off eval set should contain

The eval set should be a stratified sample of real tasks from the workflow the contract will serve, with gold answers or rubrics written before any candidate runs. Pull items from production sources such as ticket exports, CRM case histories, contract repositories or GL journals, then remove direct identifiers. Weight rare but costly cases deliberately rather than at their natural frequency; see stratified eval sets for rare and high-risk cases.

Define slices that map to business decisions, because an aggregate score hides the trade-off you actually face [2]. Typical slices are task type, input length band, language, document format (scanned PDF, HTML email, DOCX), and policy-sensitive cases where the right answer is to decline or escalate. Each slice needs enough items to separate the candidates, which is a power question covered in eval set sizing.

Audit the labels before you trust the ranking. Northcutt and colleagues estimated an average label error rate of at least 3.3% across ten widely used test sets, enough to flip model rankings [8]. A two-annotator pass with adjudication on a random subset, as described in gold-label audits, is cheap insurance before a multi-year spend.

Paired design and error bars

Run every candidate on exactly the same items under the same prompt template, then compare models item by item rather than comparing two averages. Paired tests remove per-item difficulty from the comparison and make differences between models detectable with fewer items [5]. Plan the item count with a power calculation, as covered in eval set sizing, rather than a round number.

In practice, report each pairwise difference with a 95% confidence interval, and cluster by source document or conversation when several items come from one record. Use a fixed temperature and a fixed number of samples per item, and record the model version string, date and region for every call. Hosted models change behind stable names, so a result without a version and date is hard to defend at renewal.

For agentic workflows, score consistency as well as single-run success. The tau-bench work runs agents against simulated users, APIs and domain policies and shows that success on one attempt overstates reliability across repeated attempts [6]. Run each agent task several times and report the share solved on every attempt next to the share solved at least once.

Scoring cost, latency and refusals alongside quality

A vendor decision is a cost-quality-risk trade, so the scorecard needs operational columns next to accuracy. Measure cost per resolved task, not cost per million tokens, since verbose models and retries change the effective price. Capture time to first token and end-to-end latency at p50 and p95 under realistic concurrency, plus rate-limit errors and timeouts.

Track refusals in two directions. Over-refusal on legitimate items (a benefits question declined as medical advice) costs automation rate; under-refusal on items that should be declined creates compliance exposure. Label the expected behavior per item so both are counted, and keep refusal rate as a separate column rather than folding it into accuracy.

If you use an LLM judge for open-ended answers, do not use a candidate's own model family as the only judge. Research on automated evaluation finds that judges can favor outputs from their own model [7]. Use a judge from a vendor outside the shortlist, a panel of judges, or human rubric scoring on a calibration subset, and report judge agreement with humans.

Illustrative example: invented to show structure; it does not describe an available dataset.

ColumnExample value (Model B vs Model A)Why it matters
SliceInvoice line-item extraction, scanned PDFDecision is made per slice, not on the average
Items (n) / clusters320 items / 140 source documentsClustered errors when items share a document
Paired accuracy difference+4.1 pts, 95% CI [+1.2, +7.0]Interval excludes zero, so the gap is likely real
Consistent success over 4 runs71% vs 63%Reliability, not one lucky sample
Over-refusal / under-refusal1.2% / 0.0% vs 3.8% / 0.6%Automation loss versus compliance exposure
Cost per resolved task$0.031 vs $0.024Effective price after retries and verbosity
p95 latency6.8 s vs 4.1 sSLA fit for the user-facing step
Judge-human agreementCohen's kappa 0.78 on 60-item subsetWhether automated scores can be trusted
Model version and run dateVersion string, region, 2026-09-30Reproducibility at renewal

Protecting the eval set when items go to hosted APIs

Every item you send to a hosted model is exposure, so treat the bake-off as a controlled release of a valuable asset. Before the first call, confirm in writing each provider's API terms on whether inputs are used for training, how long prompts and outputs are retained, and whether abuse-monitoring logs are kept and who can read them. Where available, use zero-retention or enterprise API tiers and keep that configuration documented.

Split the set. Use a development split to tune prompts with each vendor and keep a sealed final split that each model sees only once, at the end, under the frozen protocol. Embed canary strings in a few items so later leakage can be detected, and rotate items after the bake-off, because a set exposed to several providers loses value as a held-out test. The guide to keeping a private eval set private covers access controls and API exposure in more depth, and SourceX's note on contamination checks for licensed eval data covers how to test for prior exposure.

Do not let a candidate vendor build or hold your eval set. Research on private data curators highlights conflicts of interest when the party curating an evaluation also has a stake in model outcomes [9]. The same logic applies to third-party eval suppliers; see eval vendor independence.

A bake-off protocol you can hand to procurement

The protocol should be written and signed off before any vendor runs, so results cannot be reinterpreted after the fact. Use the checklist below as a starting point and attach it to the RFP.

Illustrative example: invented to show structure; it does not describe an available dataset.

  1. Decision statement. The workflow, the volume per month and the slice weights that will drive the choice.
  2. Eval set manifest. Item count per slice, source system, de-identification method, label audit result, and the sealed final split's hash.
  3. Candidates and settings. Model names, version strings, region, temperature, max tokens, tool definitions and system prompt, identical across vendors except where an API forces a difference.
  4. Prompt-tuning budget. Equal hours or iterations per vendor on the development split only.
  5. Metrics. Paired quality difference with 95% CI per slice, consistency over k runs, over- and under-refusal, cost per resolved task, p50 and p95 latency, error and timeout rate.
  6. Judging. Rubric, judge model outside the shortlist or human panel, and the agreement threshold for accepting automated scores.
  7. Data handling. Confirmed retention and training-use terms per provider, canary list, access log, and the date the set will be retired.
  8. Decision rule. For example, choose the lowest cost per resolved task among models whose quality interval on the two highest-weighted slices is not worse than the leader's by more than an agreed margin.
  9. Re-run trigger. Model version change, price change, or a scheduled re-test before renewal.

The protocol also gives you leverage in negotiation. A documented gap on a named slice, with intervals, is a concrete basis for asking a vendor for price concessions, a model-version pin or remediation commitments.

Where the eval data comes from

Most teams lack enough labeled, held-out examples of their target workflow, especially for new use cases or rare cases. Options are internal logs (subject to your own privacy and retention rules), commissioned expert labeling, or licensed operational data from companies that run the same workflow. See the evaluation datasets hub for the full map and SourceX's overview of evaluation datasets built from real business work; the eval set glossary entry defines the term.

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing process. Nothing is held in stock and a request does not guarantee a match. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, with personal details removed or replaced before delivery. Buyers can describe the eval data they need at SourceX for buyers.

Get licensed eval data for your model bake-off

SourceX finds US companies that hold the operational data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Every release is approved by the supplying company, and delivery runs through private, access-controlled workflows after an executed agreement. Describe the workflow your bake-off must test at https://sourcex.si/buyers.

Frequently asked questions

How many items does a vendor bake-off need?

Enough per decision slice to make the paired confidence interval narrower than the difference you care about. Pairing reduces the count needed compared with unpaired averages [5]. Plan the size with a power calculation rather than a round number.

Can we use a vendor's own evaluation harness?

Use it only for cross-checking. Run the decision protocol in your own harness, with your prompts, judges and logs, so every candidate is measured the same way and no vendor sees the sealed split's gold answers.

Should fine-tuned variants be in the same bake-off?

Only if every vendor gets the same fine-tuning data and budget, and the fine-tuning data is disjoint from the eval set. Otherwise compare base models first and run a separate round for fine-tuned finalists.

Sources

  1. arXiv, "Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts" (2024). https://arxiv.org/pdf/2410.09247.pdf
  2. arXiv, "A Practical Guide for Evaluating LLMs and LLM-Reliant Systems" (2025). https://arxiv.org/html/2506.13023v1
  3. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  4. Evidently AI, "LLM evaluation benchmarks and datasets". https://www.evidentlyai.com/llm-evaluation-benchmarks-datasets
  5. PyPI, "abeval". https://pypi.org/project/abeval/
  6. arXiv (Sierra Research), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  7. arXiv, "When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation" (2025). https://arxiv.org/pdf/2509.26600
  8. arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  9. arXiv, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data