Skip to content

Evaluation and benchmarking datasets

How many examples does an LLM eval set need? Sizing for statistical power

Quick answer

There is no single number: an LLM eval set needs enough items to detect the smallest score difference that would change your decision, at the confidence and power you choose. For pass/fail scoring at 5% significance and 80% power, one practitioner guide calculates about 2,400 items per model to separate 82% from 85% accuracy, while 100 items detect only gaps larger than about 10 points [1]. Scoring both models on the same items, clustering and rare slices then move that count.

By SourceX Editorial · Updated

Size the set from the smallest difference that would change a decision

Fix the minimum detectable effect (MDE) first: the smallest true difference between two models that you would act on. Then choose the significance level α (the false-alarm rate you accept, usually 0.05) and the power (the chance of detecting a real difference of that size, usually 0.80); the item count follows from those choices and the variance of the metric [1].

For two independent samples scored pass/fail, the standard two-proportion approximation is:

n per model ≈ (z₁₋α/₂ + z₁₋β)² × [p₁(1 − p₁) + p₂(1 − p₂)] / (p₁ − p₂)²

At α = 0.05 and 80% power the first term is (1.96 + 0.84)² ≈ 7.85. For 82% versus 85% that gives 7.85 × (0.1476 + 0.1275) / 0.0009 ≈ 2,400, the figure the guide reports [1]. Because n grows with 1/Δ², halving the MDE roughly quadruples the items [1].

The MDE comes from the decision, not from statistics. A regression gate on a prompt change may need to catch a 2-point drop; a model bake-off on your own eval set only needs to resolve a gap that justifies switching costs; a research claim of a 1-point gain needs tens of thousands of independent items per model, or several thousand with a paired analysis. Fix the MDE before anyone labels an item; what goes into the eval set itself is covered in the buyer's map of LLM evaluation datasets.

What 100 to 2,000 items can resolve

A single accuracy score on n independent items carries a 95% margin of about 1.96 × √(p(1 − p)/n), so 100 items give roughly ±10 points at 50% accuracy and 1,000 items roughly ±3 points; precision improves with the square root of n, not with n.

Figures use the normal approximation. "Discordant" is the share of items on which exactly one of two models is correct; the paired columns use the McNemar sample-size approximation at α = 0.05 and 80% power.

Items (n)95% margin, one score at 50%95% margin, one score at 85%Smallest paired gap detectable, 10% discordantSmallest paired gap detectable, 20% discordant
100±9.8 points±7.08.8 points12.4 points
200±6.9±4.96.28.8
300±5.7±4.05.17.2
500±4.4±3.14.05.6
1,000±3.1±2.22.84.0
2,000±2.2±1.62.02.8

Margins for sets below a few hundred items are optimistic [2]; the section on small-n intervals covers alternatives.

The same arithmetic applies to public leaderboards. SWE-bench Verified has 500 human-validated tasks, and GPT-4o resolved 33.2% of them at launch [5]; a single score at that level carries a 95% margin of about ±4 points. The original OSWorld has 369 tasks [7], about ±5 points near 50%, and τ-bench's airline domain has 50 tasks [8], about ±14 points. Two models a few points apart on sets of this size are often statistically indistinguishable unless a paired analysis says otherwise.

Paired tests on the same items cut the count

When both models answer the same items, test the per-item differences rather than two separate scores: variation in item difficulty, shared by both models, cancels out, so a paired design detects smaller gaps with the same set [3].

For pass/fail scoring, only discordant items carry information about the difference, and McNemar's test is the standard test on them. The approximate paired sample size is n ≈ (z₁₋α/₂√p_d + z₁₋β√(p_d − Δ²))² / Δ², where p_d is the discordant share and Δ the gap. For a 3-point gap at α = 0.05 and 80% power:

Discordant share (p_d)Items needed, pairedCompare: independent design
6%about 520about 2,400 per model
10%about 870about 2,400 per model
15%about 1,310about 2,400 per model
20%about 1,740about 2,400 per model

You cannot know p_d in advance, so score a pilot of 100 to 200 items with both models and count disagreements. Closely related systems, such as two checkpoints of one model or one model with two prompts, are likely to disagree on fewer items than unrelated models, so regression gates gain most from pairing.

Pairing does not make small sets precise. The abeval documentation works through 200 paired items on which a 13-point gap had a 95% interval of roughly 4.5 to 21.5 points [3]: the direction was clear, the size was not. For graded scores such as a 1-to-5 rubric, F1 or a judge score, use a paired t-test or a paired bootstrap that resamples items and recomputes the difference in the metric each time.

Intervals and tests that hold up at small n

Below a few hundred items, do not rely on the textbook ±1.96 standard-error interval: an ICML 2025 position paper argues that central-limit-theorem (CLT) intervals come out too narrow on LLM evals of that size and recommends better-suited methods [2].

The clearest symptom is an observed score of 0% or 100%, where the normal interval collapses to zero width. Report the interval on the difference between models, not two per-model intervals: overlapping intervals do not show two models are equivalent. As a stability check, the WordPolo authors compared estimates from a 50-item subset with a 1,500-item set and found most fell within the larger set's intervals, though not all [10].

DesignMethodData you must keepWatch for
One pass rate, several hundred items or moreNormal (Wald) intervaln and pass countToo narrow below a few hundred items [2]
One pass rate, small nWilson score, Clopper–Pearson, Agresti–Coull [4] or a Bayesian Beta posteriorn and pass countClopper–Pearson is conservative
Two models, same items, pass/failMcNemar test and interval on the paired differencePer-item outcome for each modelReport the number of discordant items
Two models, same items, graded scorePaired t-test or paired bootstrapPer-item scoresResample items, not individual responses
Several items per source recordCluster bootstrap or cluster-robust standard errorsA source record ID on every itemEffective n falls toward the number of records
Several samples per item or trials per taskAverage within item first, then test across itemsRun ID, seed and temperatureExtra runs reduce only within-item noise
Many slices or metrics tested togetherHolm or Bonferroni adjustmentList of every test runChance "regressions" in some slice

Shared sources, repeated runs and agent trials shrink effective n

Items that share a source record are not independent, so the effective sample size is smaller than the item count: ten questions written from one contract, or five turns from one support conversation, tend to pass or fail together.

The usual survey-sampling correction is the design effect, 1 + (m − 1) × ICC, where m is the average number of items per source record and ICC is the intra-cluster correlation of outcomes. With 1,000 items drawn from 200 documents (m = 5) and an ICC of 0.3, the design effect is 2.2 and the effective n is about 450. For buyers this is a purchasing decision: capping items per record and licensing more distinct records often buys more precision than writing more questions per record. Ask for a source record ID on every item so a cluster bootstrap is possible.

At temperature above zero the same item can pass on one run and fail on the next; several samples per item measure that run-to-run noise but leave the between-item variance unchanged, so beyond a few samples per item the budget is better spent on new items. Agent benchmarks make the distinction explicit: τ-bench introduced pass^k, the probability that an agent succeeds in all k independent trials of a task, and its original release has 165 tasks, 115 retail and 50 airline [8]. Trials measure reliability on those tasks; only more tasks measure how far results generalize, as discussed in agent evaluation task suites.

Errors that more items cannot average away

Sample size controls random error only; wrong gold labels, a noisy grader and contamination bias the score by the same amount at any n.

  • Gold-label errors. Northcutt et al. estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets and showed that label errors can change model rankings [9]. If your MDE is 2 points, a gold-label audit and adjudication of a delivered set is worth more than another thousand items.
  • Grader noise. An LLM judge that disagrees with human raters adds noise to every score and can shift it in one direction. Measure it on a calibration set for the LLM judge, and use chance-corrected inter-annotator agreement statistics for human raters.
  • Contamination. In a post dated February 2026, OpenAI said it had stopped reporting SWE-bench Verified because the benchmark is increasingly contaminated, so score gains increasingly reflect training-time exposure [6]. A large contaminated set gives a precise estimate of the wrong quantity; see contamination-resistant evaluation design and SourceX's guide to contamination checks for licensed eval data.
  • Reuse. Each decision made on the same test split leaks information into prompts and model choices. Keep a locked held-out test split scored once per decision, iterate on a separate development split, and plan eval set refreshes.

Rare slices need their own minimums

If a decision depends on performance in a rare slice, size that slice directly, because a random sample gives it only its share of the total. In a 2,000-item random sample, a slice that is 4% of traffic gets about 80 items, and at 80% accuracy its 95% margin is about ±9 points.

To estimate a slice's pass rate to ±5 points at 80% accuracy you need about 250 items in that slice (1.96² × 0.80 × 0.20 / 0.05² ≈ 246). Oversample the slice, keep the sampling weights, and reweight to production frequencies when you report an overall score. Testing ten slices at α = 0.05 gives roughly a 40% chance of at least one false alarm when nothing changed, so apply Holm or Bonferroni corrections. Allocation across slices is covered in stratified evaluation sets for rare and high-risk cases.

Regulated deployments add a documentation reason to size slices. For high-risk AI systems under the EU AI Act, Article 10(2) includes assessing the quantity and suitability of data sets, and Article 10(3) requires testing data sets to be relevant, sufficiently representative and, as far as possible, free of errors and complete, with appropriate statistical properties, including, where applicable, for the persons or groups on which the system is intended to be used [11]. As of October 2026, secondary sources report that Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026, moved the start of high-risk obligations to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems [12]. See the AI training data compliance hub for the wider picture.

Worked example: sizing a model-switch eval for support replies

Working one invented decision through each step shows how a 3-point target becomes item, ticket and labeling counts.

Illustrative example: invented to show structure; it does not describe an available dataset.

A team plans to replace the model behind a support-reply assistant. The current model passes a policy-compliance rubric on about 82% of replies. The team will switch only if the candidate is at least 3 points better overall, and it wants each model's pass rate on refund-exception requests, about 4% of tickets, to within ±5 points.

StepInput or assumptionResult
1. Unpaired baseline82% vs 85%, α = 0.05, power 0.80About 2,400 items per model
2. Paired pilotBoth models score 200 items; 20 discordant (10%)About 870 paired items for a 3-point gap
3. Clustering3 items per ticket; pilot ICC of paired differences 0.2; design effect 1 + 2 × 0.2 = 1.4About 1,220 items from about 410 tickets
4. Rare sliceA random 1,220 items hold about 49 refund exceptions; ±5 points at an 80% pass rate needs about 250Add about 200 refund-exception items, one per ticket
5. Audit bufferPlan for 10% of items to be dropped as ambiguous at the gold-label auditOrder about 1,580 items from about 680 tickets
6. LabelingTwo independent raters per item, adjudication on disagreementsAbout 3,160 rater judgments plus adjudication

One alternative falls out of the same arithmetic: writing one item per ticket removes the design effect, so about 870 tickets would cover step 3, which is cheaper if tickets cost less to license than items cost to label. Either way the test split stays locked, and a separate development set of a few hundred items absorbs prompt iteration.

From power target to an item and labeling request

A supplier or labeling vendor needs the power calculation translated into counts it can deliver: distinct source records, items per record, slice minimums, rater passes and a rejection buffer.

  • Decision, metric, baseline estimate, MDE, α, power, and whether the test is one- or two-sided
  • Design: paired (every model scores the same items), with the pilot's discordant share
  • Unit of analysis: maximum items per source record, conversation or document, and a source record ID on each item
  • Locked test split size, plus a separate development split
  • Each slice with its minimum count and the source fields that identify it
  • Sampling frame: systems, date window, and exclusion of records that may already be in training data
  • Labeling: raters per item, qualifications, adjudication rule, minimum agreement, and audit sample size
  • Buffer for items rejected at audit
  • Runs per item or trials per task for stochastic models and agents
  • Reporting: interval method, paired difference with its interval, multiple-comparison correction

A written evaluation dataset specification turns this list into acceptance criteria. Eval sizing differs from sizing a training purchase, which is covered in how much training data to buy.

When the items must come from real operational records rather than synthetic prompts, SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing agreement. Datasets are sourced on request rather than held in stock, and a request does not guarantee a matching dataset, so state the distinct-record count and slice minimums up front. You can send SourceX the record counts and slices your power calculation requires.

Need real business records for a powered eval set?

Describe the task, the slices and the number of distinct source records your calculation calls for. SourceX looks for US businesses that hold that data, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Describe your evaluation data needs.

Sources

  1. Tian Pan (tianpan.co practitioner blog), "Statistical power in LLM evals" (2026). https://tianpan.co/blog/2026/04/15/statistical-power-llm-evals
  2. ICML 2025 (position paper poster page), "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  3. Python Package Index (PyPI), abeval project page, "abeval". https://pypi.org/project/abeval/
  4. The Moonlight (literature review site; German-language page), "How to Correctly Report LLM-as-a-Judge Evaluations (paper review)". https://www.themoonlight.io/de/review/how-to-correctly-report-llm-as-a-judge-evaluations
  5. OpenAI (with the SWE-bench authors), "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
  6. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  7. Xie et al., arXiv 2404.07972; NeurIPS 2024, "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
  8. Yao et al. (Sierra Research), arXiv 2406.12045, "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  9. Northcutt, Athalye, Mueller (arXiv 2103.14749; NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  10. arXiv (2609.19006), "WordPolo: Evaluating Language Models Through Iterative Semantic Feedback" (2026). https://arxiv.org/pdf/2609.19006
  11. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  12. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2026/1744 (Digital Omnibus on AI) amending Regulation (EU) 2024/1689" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data