Evaluation and benchmarking datasets
Post-cutoff evaluation data: time-stamped records as temporal holdouts
Quick answer
Post-cutoff evaluation data is a test set built only from records that came into existence after the training cutoff of every model you compare, so no model could have seen them. The method depends on trustworthy creation timestamps, one shared cutoff boundary, and controls for the topic and format drift that newer records carry. Done well, it separates generalization from memorization. Done carelessly, it measures calendar drift and calls it capability.
By SourceX Editorial · Updated
This guide sits under the evaluation datasets hub and focuses narrowly on time as the holdout mechanism. For the broader toolkit (canaries, paraphrase checks, private splits), see designing contamination-resistant evaluation sets.
Why temporal holdouts beat static benchmarks for contamination control
A temporal holdout works because content that did not exist at training time cannot be in the training corpus, whatever the crawl or licensing pipeline did. Static public benchmarks lose this property over time: the LiveBench authors note that test data routinely ends up in newer models' training sets and cite evidence that LLM performance on Codeforces problems drops after a model's training cutoff date [1]. Their answer was to build questions from recent information sources and refresh them on a schedule, a design accepted at ICLR 2025 [2].
The same pressure shows up in agentic coding. OpenAI stated in February 2026 that SWE-bench Verified had become increasingly contaminated, so score gains increasingly reflected training-time exposure, and it stopped reporting the benchmark [3]. SWE-Bench Pro responded by adding held-out and commercial subsets drawn from codebases outside public training data [5]. Post-cutoff business records combine both ideas: they are new and they were never public.
Temporal splitting is not an LLM invention. Forecasting and process-mining researchers have long shown that random splits leak future information into training, and that time-ordered, case-level splits fix it [6][7]. If a field with far smaller models found leakage in random splits, a pretraining corpus of trillions of tokens deserves at least the same discipline.
Which timestamps to trust in operational records
Trust the timestamp the source system writes at creation and does not let users edit; treat every other date field as a hint. Operational systems carry several dates per record, and each answers a different question.
- Created (Zendesk
created_at, Salesforce CaseCreatedDate, Jiracreated): normally set by the server on insert, though import paths such as the Zendesk ticket import API or Salesforce's "Set Audit Fields upon Record Creation" permission can write historical values. This is your primary inclusion field. - Modified (
updated_at,LastModifiedDate,SystemModstamp, Jiraupdated): changes on any edit, including bulk admin updates and integrations. Never use it to prove novelty. - Closed or resolved (
ClosedDate, Jiraresolutiondate, asolvedstatus change in ticket audits): marks when the outcome became known. Use it for outcome-labeled items, because the label must also post-date the cutoff. - Business dates typed by users (invoice date, contract effective date, "date of incident"): editable and often backdated. Useful for slicing, not for the cutoff test.
Git history needs extra care. A commit carries an author date and a committer date; both can be set by the client, and rebases or cherry-picks rewrite the committer date. Prefer server-side evidence such as pull request creation events or push logs from the hosting platform.
Two failure modes recur. Migrated records may inherit the migration date as created_at, so a 2019 ticket imported into a new helpdesk in 2026 looks fresh; ask whether imports preserved original dates. Templated or recycled content (a macro reply, a reused contract clause, a ticket copied from an older one) can be newly created yet textually old, so a creation timestamp alone does not prove the text is novel.
Verifying timestamps in licensed records
Ask the supplier for evidence that the dates are system-generated, then test them yourself on a sample. Concretely, request the field's provenance (system of record, whether users or APIs can write it), an export of the audit or event log for a sample of records, and the dates of any platform migrations or bulk imports.
Then run checks on delivery:
- Monotonicity: record IDs that increase with creation time (common for auto-increment keys) should not show created dates jumping backward.
- Audit agreement: the first event in the audit trail should match
created_atwithin seconds. - Migration spikes: a histogram of creation dates by day should not show one day holding thousands of records.
- Text novelty: run n-gram overlap and nearest-neighbor checks against public corpora and older records [4], so recycled text is caught even when the date is new.
If the supplier keeps data in a lakehouse, snapshot tags (for example, Apache Iceberg tags on a specific snapshot) give a durable reference to the table state at extraction time, which makes later disputes about what was delivered easier to settle [8]. Our point-in-time correct data guide covers as-of extraction in more depth.
Matching cutoff windows across the models you compare
Set the holdout boundary at the latest training cutoff among all models under comparison, plus a buffer, and apply that single boundary to every model. If model A's cutoff is earlier than model B's, records from the gap are unseen by A but may be seen by B, and the comparison favors B on exactly that slice.
Published cutoffs need skepticism. Vendors may state a "knowledge cutoff" that differs from when the last training data was collected, and later post-training, retrieval tools or browsing can expose a model to newer material. Add a buffer of weeks to months and, where you can, measure it: score each model on monthly buckets around the stated cutoff and look for the kind of step change LiveBench cites [1]. A step that lands later than the stated date suggests the effective cutoff is later.
For models with tool access, disable web search and retrieval during the run, or log every retrieved URL and drop items whose answers were retrievable. Otherwise a post-cutoff item can be answered by lookup rather than reasoning.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Situation | Boundary rule | Risk if ignored |
|---|---|---|
| Comparing two models with stated cutoffs in different months | Use the later cutoff plus a buffer for both | Earlier-cutoff model penalized on items the other may have seen |
| Model receives continued post-training after release | Use the latest known data date, not the original cutoff | Fresh items may already be in post-training data |
| Model can browse or call retrieval tools | Disable tools or filter items with retrievable answers | Lookup inflates scores |
| Tracking one model across versions | Fix the boundary at the newest version's cutoff and re-score older versions | Version gains confounded with item recency |
Controlling for distribution shift in newer records
Post-cutoff records differ from older ones in more than novelty, so compare like with like before attributing a score drop to contamination. Product launches, policy changes, new error codes, seasonal volume and staff turnover all shift topic mix, length and difficulty; see temporal coverage, gaps and seasonality for how these show up in licensed data.
Three controls help. First, stratify both pre-cutoff and post-cutoff pools by the same slices (category, product line, channel, length band) and compare within slices, as in stratified evaluation sets. Second, keep a matched pre-cutoff control set from the same source and processing pipeline, so the only systematic difference is time. Third, report the share of post-cutoff items that reference entities, products or rules that did not exist before the cutoff; those items test knowledge the model could not have, not generalization, and usually belong in a separate slice.
The gap between control and holdout scores, within matched slices, is your contamination or memorization estimate. A gap that persists across slices is stronger evidence than a single headline number.
A record schema for date-stamped eval items
Each eval item should carry its timing evidence with it, so anyone can re-run the cutoff filter for a new model without going back to the supplier. Write these fields into your evaluation dataset specification.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"item_id": "ev-000412",
"source_system": "helpdesk",
"source_record_created_at": "2026-08-14T09:22:51Z",
"created_at_provenance": "server-generated; verified against audit log first event",
"label_available_at": "2026-08-19T16:05:10Z",
"migration_flag": false,
"text_novelty": {"max_ngram_overlap_public": 0.03, "near_duplicate_of_older_record": false},
"slices": {"category": "billing_dispute", "channel": "email", "length_band": "medium"},
"references_post_cutoff_entity": true,
"deidentification_method": "names, emails, phone numbers and account numbers replaced",
"eligible_for_models_with_cutoff_before": "2026-06-01"
}
The last field is computed, not stored by hand: it is the creation or label date, whichever is later, minus your buffer.
Sourcing fresh records on an ongoing basis
A temporal holdout decays as soon as the next model generation trains past it, so plan for repeated pulls of newly created records rather than one delivery. Any single set becomes contaminated once it leaks or once newer models train on overlapping public discussion of it. Keep the set private, as covered in private evaluation sets vs public benchmarks, and check our owner guide on contamination checks for licensed eval data.
When you license operational records for this purpose, write the time logic into the request: the minimum creation date, which timestamp defines it, how labels are dated, and whether later pulls may include records edited since the last one. Evaluation-specific license terms are covered in evaluation-only data license terms, and the definition of a held-out data set applies here as usual.
SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records, and finance and legal workflows, and manages licensing and ongoing purchases; each release is approved by the supplying company, and a request does not guarantee a match. You can describe the records and date window you need on the buyer request page.
Request post-cutoff evaluation records
SourceX looks for US businesses that hold the time-stamped records you describe, assesses the data and licensing permissions, and delivers each rights-reviewed dataset under a license defining records, uses, term and delivery. Personal details such as names, emails and account numbers are removed or replaced before delivery. Describe the record type, date window and timestamp requirements at sourcex.si/buyers.
Frequently asked questions
How long after a cutoff should the holdout window start?
There is no universal number. Start at the latest stated cutoff among compared models plus a buffer, then use monthly score buckets to see where performance changes and move the boundary if the step appears later [1].
Can I use modified dates if created dates are missing?
Only as an upper bound. A record modified after the cutoff may have been created years earlier, so without a creation date or audit log you cannot prove it is unseen.
Does a post-cutoff record guarantee the text is new?
No. Templates, copied tickets and reused clauses can carry old text in a new record, so pair the timestamp filter with overlap and near-duplicate checks [4].
Sources
- arXiv (White et al.), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
- ICLR, "LiveBench: A Challenging, Contamination-Limited LLM Benchmark (ICLR 2025 poster)" (2025). https://iclr.cc/virtual/2025/poster/28134
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- arXiv, "Investigating Data Contamination in Modern Benchmarks for Large Language Models" (2023). https://arxiv.org/html/2311.09783v2
- arXiv, "SWE-Bench Pro" (2025). https://arxiv.org/html/2509.16941v1
- arXiv (Weytjens and De Weerdt), "Creating Unbiased Public Benchmark Datasets with Data Leakage Prevention for Predictive Process Monitoring" (2021). https://export.arxiv.org/abs/2107.01905
- temporalcv documentation, "Why Time Series Is Different". https://temporalcv.readthedocs.io/en/latest/guide/why_time_series_is_different.html
- Apache Iceberg, "Branching and Tagging". https://iceberg.apache.org/docs/1.7.0/branching
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.