Evaluation and benchmarking datasets
Stratified evaluation sets: allocating examples to rare and high-risk cases
Quick answer
A stratified evaluation set splits a fixed example budget across named slices (intent, risk tier, customer segment, failure mode) instead of sampling in proportion to traffic. Give every release-gating slice a minimum count large enough to detect the regression you care about, spread the remaining budget by traffic share, and report traffic-weighted and per-slice scores separately. Proportional sampling leaves a 0.5% high-risk slice with a handful of examples, which cannot show anything except a catastrophic failure.
By SourceX Editorial · Updated
Why proportional sampling hides the failures that matter
Proportional sampling optimizes the precision of the headline number, which is the wrong target when a rare slice carries most of the risk. An aggregate pass rate is dominated by frequent request types, so a model can regress badly on a rare but critical slice while the overall score barely moves [4]. If refund disputes over a regulatory threshold are 0.4% of traffic, a 2,000-example random sample contains about eight of them.
Eight examples cannot distinguish a 90% pass rate from a 70% one with any confidence. Production failures also cluster in the tail, which is why practitioner guidance recommends stratifying by intent, deliberately over-weighting the adversarial tail and keeping regression fixtures for every incident that reached users [1]. Stratification is how you turn that advice into a budget.
This is a different question from total set size, which the guide on eval set size and statistical power covers. Here the total is fixed and the problem is allocation.
Defining slices that map to decisions
A useful slice is one where a regression would change a ship decision, so define slices from risk and failure analysis rather than from whatever metadata is easy to filter. Practitioners typically cut along four axes: intent or task type, risk tier, customer segment and known failure mode [1]. Keep the taxonomy small enough that each gated slice can actually be funded.
- Intent: the task the user is attempting, for example "cancel subscription", "dispute charge" or "extract invoice line items".
- Risk tier: consequence of a wrong answer, for example financial loss, safety, legal exposure or account lockout.
- Segment: enterprise versus self-serve, language or locale, regulated industry, channel (chat, email, voice).
- Failure mode: hallucinated policy, wrong tool call, unsafe compliance with a harmful request, missed escalation, formatting failure that breaks a downstream parser.
Slices can overlap; a single example can be "billing dispute, high-risk, enterprise, policy hallucination". Pick one primary stratum for allocation (usually intent crossed with risk tier) and store the rest as tags so you can compute secondary breakdowns without funding every cell. Fully crossing four axes creates hundreds of cells, most of which will never reach a usable count.
For high-risk systems in the EU, Article 10 of the AI Act requires training, validation and testing data sets to meet quality criteria including relevance and representativeness for the intended purpose [8]. A documented slice taxonomy with counts per slice is a practical way to show what "representative" meant for your test data; see Article 10 data governance for licensed test data. As of October 2026, the Annex III high-risk application date has reportedly moved to 2 December 2027.
Per-slice floors: the arithmetic of rare cases
Set each gated slice's minimum count from the smallest regression you need to detect in that slice, because a slice's sample size, not the set's total, limits what you can see there [3]. For a pass rate near 90%, the half-width of a 95% interval is roughly 1.96 × √(p(1−p)/n). That works out to about ±8 points at n = 50, ±6 points at n = 100, ±4 points at n = 200 and ±3 points at n = 400.
Detecting a change between two model versions is harder than estimating one rate. With an unpaired comparison at 80% power and α = 0.05, the minimum detectable effect is roughly 2.8 × √(2p(1−p)/n): about 12 points at n = 100 per slice and about 6 points at n = 400. Running both models on the same items and testing the paired differences (McNemar's test for pass/fail) usually tightens this, because most items pass or fail for both models.
Small slices also break the usual interval math. A position paper at ICML 2025 argues that CLT-based intervals are too narrow on evaluation sets below a few hundred datapoints [5]; for small slices use Wilson or Clopper-Pearson intervals or a Bayesian beta-binomial posterior. When a slice has zero observed failures, the "rule of three" gives an approximate 95% upper bound of 3/n on the failure rate: zero failures in 60 examples is still consistent with a 5% failure rate.
The practical rule follows from the numbers. A slice gated on "no more than a 5-point drop" needs several hundred examples; a slice with only 40 examples can gate on catastrophic failure only, and the release criteria should say so explicitly.
Allocating a fixed budget: floor first, then proportional
The allocation that works in practice is a two-pass scheme: fund every gated slice to its floor, then distribute what remains across the other slices roughly in proportion to traffic. This is a pragmatic variant of stratified sampling with a precision constraint per stratum. Neyman allocation, which assigns examples in proportion to stratum size times score standard deviation, minimizes variance of the overall estimate but usually still starves rare slices, so use it, if at all, for the remainder.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Slice (intent × risk) | Traffic share | Proportional n (of 2,000) | Floor | Allocated n | Sampling weight |
|---|---|---|---|---|---|
| Account questions, low risk | 48.1% | 962 | 100 | 640 | 1.37 |
| Billing explanations, medium | 30% | 600 | 150 | 420 | 1.30 |
| Plan changes, medium | 18% | 360 | 150 | 260 | 1.26 |
| Charge disputes over threshold, high | 3.5% | 70 | 300 | 300 | 0.21 |
| Self-harm or safety disclosure, high | 0.4% | 8 | 200 | 200 | 0.04 |
| Known regression fixtures (all past incidents) | n/a | 0 | all | 180 | excluded from weighted score |
The sampling weight is traffic share divided by the slice's share of the 1,820 scored (non-fixture) examples; it converts the oversampled set back to a production-shaped estimate. In this example, the two high-risk slices take 25% of the budget while carrying under 4% of traffic. That is the intended trade: the weighted headline number becomes slightly less precise, and the slices that can hurt users become measurable.
Revisit floors when the gate changes. If the safety slice gates on "zero critical failures", 200 examples only bound the failure rate at about 1.5% by the rule of three, which may not satisfy a reviewer working under the MEASURE function of the NIST AI RMF [7]. In that case either raise the floor or pair the eval set with targeted red-teaming.
Weighted, unweighted and per-slice reporting
Report three numbers, because an oversampled set answers three different questions. The traffic-weighted score (sum of slice pass rates times traffic share) estimates production quality. The unweighted pooled score is only a description of the set and should not be compared across set versions. Per-slice scores with intervals are what the release gate reads.
Most mistakes come from mixing these. A team that oversampled hard cases reports the pooled score as "accuracy", sees it far below production experience and loses trust in the eval. Another team reweights correctly but gates on the weighted number, so the high-risk slice is again invisible.
A gate specification should name, for each slice: the metric, the threshold, the interval method, the comparison baseline and what happens on a miss (block, require sign-off, or warn). Write these before you run the candidate model, so the gate cannot be tuned to the result.
Sourcing and labeling the rare slices
Rare slices are the expensive part of a stratified set, because production logs rarely contain enough of them and synthetic substitutes tend to be easier than real cases. Start with incident tickets, escalations, compliance reviews and QA rejections, which are pre-filtered toward the tail. When internal data runs short, request the rare case types from data suppliers by name, with the definition, the expected prevalence in their records and the minimum count you need.
Label quality matters more in small slices. Northcutt and colleagues estimated an average label error rate of at least 3.3% across widely used test sets [6]; in a 60-example slice, two wrong gold labels move the pass rate by more than 3 points. Have rare-slice items double-labeled and adjudicated, and include cases near the rubric boundary on purpose, alongside common traffic, rare high-risk cases and known failures, when you calibrate an LLM judge against humans [2].
Illustrative example: invented to show structure; it does not describe an available dataset.
slice_id: billing.dispute.over_threshold
definition: "Customer disputes a charge above the auto-refund limit; correct handling is escalation with a case reference, not a refund promise"
risk_tier: high
primary_failure_modes: [promises_refund, omits_escalation, invents_policy]
gate: { metric: pass_rate, max_drop_pts: 5, interval: wilson_95, baseline: prod_model_2026_09 }
target_n: 300
current_n: 214
sources: [escalation_tickets_2025, supplier_request_open]
labeling: { annotators: 2, adjudication: required, rubric_version: v4 }
supplier_request:
case_type: "chat or email transcripts of disputed charges escalated to a human agent"
expected_prevalence: "about 3 to 4 percent of billing contacts"
min_records: 150
required_fields: [transcript, resolution_code, escalation_flag, policy_reference]
Synthetic generation can fill structural gaps, such as paraphrases of a known jailbreak, but it should not be the only source for a gated slice; the limits are covered in when synthetic eval data misleads. Rare real cases also need the same contamination and leakage controls as the rest of the set, described in keeping a private eval set private.
Keeping slices current as traffic shifts
Slice allocations go stale as traffic shifts, so recompute weights and floors on a fixed cadence and whenever a new incident class appears. New production failures should enter as regression fixtures immediately and be promoted into a funded slice once they recur [1]. Retire slices whose risk was engineered away, but keep their fixtures.
Version every change: slice definitions, floors, weights and gold labels. A pass-rate change between two set versions is not a model change unless the slice composition is held constant. The broader cadence questions are in eval set refresh, saturation and drift, and the coverage side of the problem is in long-tail and edge-case coverage.
Checklist before you freeze a stratified eval set
Illustrative example: invented to show structure; it does not describe an available dataset.
- Every gated slice has a written definition, a risk tier and at least one named failure mode.
- Each slice's floor is derived from its gate (detectable drop or failure-rate bound), not from round numbers.
- Allocated counts and sampling weights are stored with the set version.
- Interval method is set per slice (Wilson, Clopper-Pearson or Bayesian below a few hundred examples).
- Rare-slice items are double-labeled with adjudication, and boundary cases are included in judge calibration.
- Reports show the weighted score, per-slice scores with intervals and the regression fixture results separately.
- Gate rules (threshold, baseline, action on miss) were fixed before the candidate run.
- Rare-case data requests to suppliers state the case definition, expected prevalence and minimum count.
For the wider map of benchmark, private and licensed eval data, start at the evaluation datasets hub or the AI data overview. SourceX's owner page on evaluation datasets built from real business work and the glossary entry for an eval set cover terminology and use cases; for support-agent slices, see customer support training data and eval sets. Before scoring licensed data, run the contamination checks for licensed evaluation data.
When your rare slices need real operational records, such as escalated support histories or finance and legal workflow documents, you can describe the case types on the SourceX buyer request page. SourceX looks for US businesses that hold the described data, and each release is approved by the supplying company.
Requesting rare-case data for stratified eval sets
SourceX sources operational datasets from US companies on request, so you describe the slice and its case types rather than picking from inventory, and a request does not guarantee a match. Every dataset is rights-reviewed and delivered under a license that defines the records, allowed uses, term and delivery, with personal details removed or replaced before delivery. Describe the rare cases your eval set needs.
Sources
- OneUptime, "How to Build a Golden Evaluation Dataset from Real LLM Production Failures" (2026). https://oneuptime.com/blog/post/2026-08-31-build-golden-evaluation-dataset-production-failures/markdown
- OneUptime, "How to Calibrate an LLM-as-a-Judge Against Human Labels with Cohen's Kappa" (2026). https://oneuptime.com/blog/post/2026-08-31-calibrate-llm-judge-cohens-kappa/markdown
- Tian Pan, "Statistical power in LLM evals" (2026). https://tianpan.co/blog/2026/04/15/statistical-power-llm-evals
- Tian Pan, "The long-tail coverage problem in AI systems" (2026). https://tianpan.co/blog/2026/04/19/long-tail-coverage-problem-ai-systems
- ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
- Northcutt, Athalye, Mueller (NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.