Skip to content

Evaluation and benchmarking datasets

Eval set refresh cadence: handling saturation, drift and leakage

Quick answer

Refresh an eval set when it stops telling you something: when top models cluster near the ceiling, when the set no longer resembles production traffic, or when you suspect items have leaked into training data. A workable cadence for most standing programs is a frozen anchor subset for trend continuity plus scheduled additions and retirements, sized so each new slice has enough items to detect the differences you care about. Fresh items reduce contamination risk; they do not remove it.

By SourceX Editorial · Updated

This guide is for evaluation leads who own a recurring scorecard rather than a one-off bake-off. For the broader map of benchmarks, private sets and licensed test data, start at the evaluation datasets hub; for item structure that resists leakage by design, see contamination-resistant evaluation design.

Three signals that an eval set needs refreshing

An eval set needs attention when it saturates, drifts or leaks, and each failure shows up in different telemetry. Treat them as separate triggers with separate thresholds, because the fix differs.

Saturation. When the frontier models you compare all score in a narrow band near the top, the set has lost discriminative power. Needle-in-a-haystack retrieval is a familiar case: models reaching near-perfect scores can overstate real long-context ability [4]. Label noise makes the ceiling lower than 100%; an audit of 10 widely used test sets estimated an average label error rate of at least 3.3% [5], so on a noisy set the last few points of "improvement" can be models agreeing with wrong gold labels.

Drift. Production traffic moves as products, prompts, customer mixes and tool integrations change. A practitioner write-up on building golden sets from production failures notes that keeping a set representative for months while prompts, models and users change is the step most teams get wrong [9]. Drift shows up as a widening gap between eval scores and production quality metrics such as escalation rate, human override rate or task completion.

Leakage. Once items are public, shared with vendors through an API, or reused in training pipelines, scores start to reflect exposure rather than capability. The LiveBench authors note that test data can end up in a newer model's training set and make a benchmark obsolete quickly [1]. As of October 2026, OpenAI has stopped reporting SWE-bench Verified, stating that it is increasingly contaminated so gains reflect training-time exposure [3].

How to tell saturation from real progress

A score near the ceiling is saturation only if the remaining headroom is smaller than your measurement error. Do the arithmetic before you retire anything.

On a 500-item binary-scored set, a model at 95% accuracy has a standard error of roughly sqrt(0.95 x 0.05 / 500), about 1 point, so a 95% confidence interval spans roughly plus or minus 2 points. If your top three candidates sit at 94, 95 and 96, the set cannot rank them. Small refresh slices make this worse: a position paper at ICML 2025 argues that CLT-based intervals become too narrow below a few hundred items [6], so use bootstrap or Bayesian intervals on per-slice results. For sizing each slice, see how many examples an eval set needs.

Check three things before declaring saturation:

  • Error audit on the residual failures. Re-adjudicate the items every top model fails. If many are gold-label errors or ambiguous prompts, fix labels first; the gold-label audit and adjudication guide covers the process.
  • Per-stratum ceilings. An aggregate of 95% can hide a rare-case stratum at 60%. Saturation in the easy strata argues for reallocating items, not replacing the set; see stratified eval sets for rare and high-risk cases.
  • Contamination probe. Compare scores on items created before and after each candidate model's training cutoff. A large gap points to leakage rather than saturation.

The core trade-off is comparability over time against freshness, and most programs resolve it with a frozen anchor plus a rotating pool. Practitioner guidance to freeze and version a golden set [8] and benchmark designs built on frequently updated questions [1] are both right; they protect different things.

Frozen anchor subset. Keep a fixed, access-restricted slice, sized by your own power analysis, that never changes between versions. It carries trend lines across model generations and tells you whether a score change came from the model or the set.

Rolling window. Add items from a recent time window on a fixed cadence and retire the oldest rotating items. LiveBench draws questions from recent information sources with objective ground truth to limit exposure [1], and continuously evolving designs have since been applied to multimodal models [7]. The ICLR 2025 listing calls LiveBench "contamination-limited," a fair reminder that rotation lowers leakage risk without eliminating it [2].

Event-driven additions. Add items when production telemetry shows a new failure cluster, a new tool, a new product line or a regulatory change in the domain. These typically feed a "regression" slice that is reported separately.

Retirement rules. Retire an item when every tracked model passes it for N consecutive versions, when its gold label is disputed and cannot be adjudicated, or when it has been exposed (shared outside the eval boundary, posted publicly, or sent to a third-party API without a no-retention term). Log the retirement reason, not just the deletion.

A refresh policy you can adapt

A written refresh policy turns these patterns into decisions your team and suppliers can follow. The table below shows the structure; thresholds must come from your own power analysis and risk appetite.

Illustrative example: invented to show structure; it does not describe an available dataset.

TriggerMeasurementThreshold (example)ActionVersion bump
SaturationSpread of top-3 model scores vs. 95% CI widthSpread < CI width for 2 releasesAudit residual failures; add harder items to saturated strataMinor (v3.2 to v3.3)
DriftJensen-Shannon divergence between eval and 30-day production intent mixAbove team-set limitAdd rolling-window items from new intents; rebalance strataMinor
LeakagePre/post-cutoff score gap; canary string detection; exposure logGap > 2x CI or any confirmed exposureRetire exposed items; replace from reserve poolMajor if anchor touched
ScheduledCalendarQuarterlyAdd rolling slice; retire oldest rotating itemsMinor
Label method changeRubric or adjudication protocol revisedAny changeRe-score anchor under both methods; report bridgeMajor

Pair the table with a version manifest for each release. A minimal record looks like this:

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "eval_set": "support-agent-eval",
  "version": "3.3.0",
  "released": "2026-10-01",
  "anchor_items": 300,
  "rotating_items": 900,
  "added": {"count": 150, "source_window": "2026-07-01/2026-09-30", "trigger": "scheduled"},
  "retired": {"count": 120, "reasons": {"all_models_pass": 80, "exposed": 25, "label_dispute": 15}},
  "label_protocol": "rubric-v4, two-annotator + adjudicator",
  "schema_version": "2.1",
  "reserve_pool_remaining": 410
}

Document each version the way you would any dataset: sources, collection window, annotation method and intended use, in the spirit of Data Cards [10]. Keep a reserve pool of unreleased items so a leakage event does not force an emergency sourcing cycle.

Choosing a cadence by domain

Cadence should follow how fast the underlying distribution and the model landscape move, not the calendar alone. Fast-moving domains need short windows; stable domains can rely more heavily on the anchor.

  • Customer support and sales agents: product catalogs, policies and macros change often, so quarterly rolling additions drawn from recent tickets are a reasonable starting point, with event-driven slices after launches.
  • Coding agents: public repositories and issue trackers are heavily represented in pretraining corpora, so leakage dominates. Favor held-out tasks from private codebases; see held-out coding agent evaluation sets.
  • Document extraction and finance workflows: form layouts and field definitions change slowly, so a larger anchor with annual refreshes and event-driven additions for new templates often suffices.
  • Long-context and RAG: saturation on synthetic retrieval tests arrives early [4], so refresh toward real document sets and permission-aware cases rather than longer haystacks.

Whatever the domain, track exposure. Each time items go to an external model API or a third-party grader, record it; the guide on keeping a private eval set from leaking covers canaries and access controls.

Supply terms for recurring eval data

A refresh cadence only works if new items arrive on schedule in a stable shape, so recurring eval data needs supply terms that internal teams often leave implicit. Before you sign, ask suppliers to commit in writing to:

  • Delivery cadence and window: which time window each delivery draws from, and how far in arrears.
  • Schema stability: field names, types and enumerations held constant across deliveries, with versioned change notices.
  • Label method stability: the same rubric, annotator qualification and adjudication protocol, or a bridge study when any of them changes.
  • Non-overlap: no item, or near-duplicate, appearing in an earlier delivery or in data the supplier licensed to others for training.
  • Exposure controls: how the data is held before delivery and who can see it.
  • Recurring license scope: whether each delivery falls under one license or a renewal; the question can I license data every year covers renewal structures, and data supplier SLAs covers service levels.

Run incoming slices through your own contamination checks for licensed eval data and acceptance audit before they join the scored pool.

Where SourceX fits in a refresh program

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. The kinds of data include support and sales histories, engineering records, documents, and finance and legal workflows, which can supply rotating-pool items drawn from real work; see AI evaluation datasets built from real business work. Data is not held in stock, categories are not inventory, and a request does not guarantee a match.

Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows after an executed agreement and supplier approval. Buyers describe the data they need on the SourceX buyer page.

Sourcing fresh items for your eval refresh cycle

If your anchor is holding but your rotating pool is running dry, describe the data window, schema and label method you need, and SourceX will look for US businesses that hold it. Each release is approved by the supplying company, and pricing and allowed uses are agreed per deal in a license. Start a buyer request at sourcex.si/buyers.

Sources

  1. arXiv (White et al.), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
  2. ICLR 2025, "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2025). https://iclr.cc/virtual/2025/poster/28134
  3. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  4. arXiv, "Make Your LLM Fully Utilize the Context" (2024). https://arxiv.org/pdf/2404.16811
  5. arXiv / NeurIPS 2021 Datasets and Benchmarks (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  6. ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  7. arXiv, "MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models" (2026). https://arxiv.org/pdf/2607.01813
  8. Alpesh Nakrani (blog), "Golden eval set". https://alpeshnakrani.com/blog/golden-eval-set/
  9. OneUptime (blog), "Build a golden evaluation dataset from production failures" (2026). https://oneuptime.com/blog/post/2026-08-31-build-golden-evaluation-dataset-production-failures/markdown
  10. Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data