Data quality, coverage and contamination
Training Data Freshness: Measuring Staleness and Setting Recency Requirements
Quick answer
Training data freshness is how closely a dataset reflects the conditions your model will face in production. Measure it from record timestamps, not from the delivery date: median record age at delivery, the share of records older than a cutoff, and the lag from a business event to its export. Then set recency requirements by content type. Policy, pricing and product facts need short windows. Procedural know-how and stable workflows tolerate older records. Research shows that temporal mismatch between training and evaluation data degrades performance and is hard to fix afterward [1].
By SourceX Editorial · Updated
Why stale training data hurts models in production
Stale data teaches a model a version of the world that no longer exists, and the damage shows up most on inputs drawn from later periods. Longpre et al. measured the effect of pretraining data age and found that a temporal gap between pretraining data and evaluation data degrades downstream performance, and that fine-tuning on newer data does not fully close the gap [1]. Fabre and colleagues report that sequentially trained models match shuffled baselines and exhibit more up-to-date and temporally precise knowledge [2].
For a production team, staleness produces recognizable failure modes:
- Superseded facts. A support model trained on 2023 tickets recommends a plan tier, return window or API endpoint that was retired.
- Vocabulary drift. New product names, error codes and SKUs appear as out-of-vocabulary noise, so intent classifiers route them to "other".
- Process drift. A claims or onboarding workflow added a step, and trajectories in the training set skip it. Agents trained on those records learn the old sequence.
- Label drift. An outcome field changed meaning after a code-set revision, so old and new rows disagree on what "resolved" means.
Fine-tuning is a weak patch for missing recent knowledge. Gekhman et al. found that models learn new factual knowledge slowly through fine-tuning, and that examples introducing new knowledge, once learned, increase hallucination [5]. If a model must know current facts, retrieval over a maintained corpus is usually the better route, and the SFT set should teach behavior rather than carry volatile facts. The knowledge-corpus quality guide for RAG covers stale document versions on the retrieval side.
Staleness metrics to compute on every delivery
You measure staleness by profiling the event timestamps inside the records against a reference date, usually the delivery date or the date your model ships. The delivery date alone says nothing; a dataset exported last week can be mostly five years old. Compute these metrics per delivery and per slice (product line, region, channel), because a fresh aggregate can hide a stale segment. ISO/IEC 5259-4 frames data quality as a managed process across training and evaluation, which is the right home for these checks [7].
| Metric | Definition | Typical field used | What it catches |
|---|---|---|---|
| Median record age | Reference date minus the median event timestamp | created_at, closed_at, event_time | Overall skew toward old history |
| P90 record age | Age of the 90th-percentile oldest record | Same | Long tail of very old rows |
| Share older than N months | Records with age > N divided by total | Same | Direct test against a recency requirement |
| Recent-window density | Records in the last 90 days divided by the 12-month monthly average | Same | Thin or missing recent coverage |
| Event-to-export lag | Export timestamp minus event timestamp, median and max | exported_at vs event_time | Pipeline delay and backlog |
| Max-date gap | Reference date minus the newest record | Same | Truncated extracts, frozen sources |
| Superseded share | Records referencing a retired product, policy version or code | policy_version, sku, code fields | Content that is valid in form but obsolete |
Pick the timestamp deliberately. In ticket data, created_at and closed_at can differ by weeks, and a reopened ticket may carry an updated_at that makes old content look new. In document corpora, file-system modified dates change on migration, so prefer an effective date or version field. Ask the supplier which field drives each metric and record it in the dataset's documentation; the Data Cards framework explicitly covers upstream sources and decisions that affect model performance [8].
Illustrative example: invented to show structure; it does not describe an available dataset.
freshness_profile:
dataset: support_tickets_extract_v3
reference_date: 2026-10-01
age_field: closed_at
records: 412000
median_age_days: 410
p90_age_days: 1180
share_older_than_24_months: 0.31
recent_90d_density_ratio: 0.42 # last 90 days vs 12-month monthly average
event_to_export_lag_days: {median: 9, max: 63}
newest_record: 2026-08-14
superseded_share:
check: "references to plan tiers retired before 2025-06-01"
value: 0.18
slices_flagged: ["region=EMEA (median age 790d)", "channel=chat (no records after 2026-05)"]
A recent-window density ratio well below 1.0 with a max-date gap of seven weeks usually means a truncated extract, not a quiet quarter. Confirm before you accept the delivery; the acceptance sampling guide shows how to treat freshness failures as defects in an acceptance plan.
Setting recency requirements by use and content type
A recency requirement should follow how fast the underlying content changes, not a single company-wide number. Short windows suit volatile facts; long windows suit stable procedures, and they also buy you the long-tail coverage that only years of history provide. The table below is a starting point to adapt to your own measured drift.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Content type | Example records | Suggested max median age | Suggested cap on records older than | Main risk if stale |
|---|---|---|---|---|
| Pricing, policy and product facts | Plan terms, return rules, release notes | 3–6 months | 12 months | Confident wrong answers |
| Customer conversations | Support tickets, chat transcripts | 6–12 months | 36 months | Outdated resolutions, missing new intents |
| Sales and CRM activity | Call notes, opportunity histories | 12 months | 36 months | Old objections and competitors |
| Engineering records | Incidents, code review threads | 12–18 months | 48 months | Deprecated APIs and tooling |
| Finance and legal workflows | Invoices, contract redlines | 18–24 months | 60 months, with version tags | Old clause standards, tax rules |
| Procedural know-how | Recorded hands-on work, SOP walkthroughs | 36+ months | Case by case | Low, unless equipment or steps changed |
Two refinements make these numbers defensible. First, set a recency requirement per slice that matters to your deployment, which ties freshness to coverage gap analysis against your deployment distribution. Second, allow older records when they carry an explicit version tag, so you can filter or reweight rather than discard them. Multi-year sets need that tagging anyway because definitions shift; see code-set revisions in multi-year operational datasets.
Measuring staleness against current conditions
The most direct staleness test is behavioral: train or fine-tune on the candidate data, then evaluate on a held-out slice from the most recent period. Split by time, not at random. Random splits of event-log and case data leak information across cases, and temporal, case-level splits are the common fix in process-mining benchmarks [6]. The same applies to tickets, claims and transactions: hold out the last one to three months by event date and keep whole cases on one side of the cut.
Run three checks on that split:
- Temporal decay curve. Train on windows ending at different dates and plot accuracy on the newest holdout. A steep slope means your use case is time-sensitive and the recency cap should be tight.
- Recent-only ablation. Compare a model trained on the full history with one trained on the last 12 months. If the full history wins on long-tail slices but loses on common recent intents, keep history but reweight toward recent records.
- Fact-currency audit. Sample 200 records, extract the factual claims (prices, versions, policies), and check them against current sources. The share of obsolete claims estimates the superseded share. The factual accuracy checks for licensed text guide covers the sampling design.
Record the metrics alongside your other dimensions, such as completeness and consistency, as described in training data quality metrics. Timeliness belongs in the same scorecard, not in a separate spreadsheet, alongside the other checks in the training data quality assessment hub.
Evaluation sets go stale too
Your benchmark ages on the same clock as your training data, so a stale eval can hide a stale model or punish a fresh one. Research on question answering under temporal conflict shows that facts keep changing after a benchmark is built, so a fixed gold answer can become outdated and penalize a model that gives the current one [3]. Contamination compounds the problem: test items leak into newer models' training data and make benchmarks obsolete quickly, which is why LiveBench refreshes questions from recent sources on a schedule [4].
For production teams, the practical rules are:
- Version every eval set with a
valid_as_ofdate and re-verify answers that depend on volatile facts before each major model release. - Keep a rolling "recent" eval slice built from the last quarter's operational records, separate from the fixed regression suite.
- Never source the recent eval slice from the same export as training data without a time cut and a case-level overlap check.
Refresh cadence for licensed data
Refresh cadence should equal the time it takes your measured decay curve to cost more than a retrain is worth, and that number comes from the checks above. If accuracy on the newest holdout drops a few points per quarter, a quarterly or semiannual top-up is a reasonable hypothesis to test. If the curve is flat for two years, buy once and revisit annually. Ask suppliers early whether a source can produce recurring extracts with stable schemas and consistent timestamp fields, because a refresh with renamed columns costs as much as a new integration.
Contract terms for recurring deliveries, including freshness and latency commitments, belong in procurement rather than in the quality spec. The SLA guide for recurring data deliveries and the broader AI training data procurement guide cover those terms. For buyer patterns on cadence, see how often AI buyers want fresh data and whether data can be licensed every year.
Writing freshness into a data request
A data request should state recency as measurable acceptance criteria so a supplier can check fit before anything is extracted. SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents and finance and legal workflows; nothing is held in stock, and a request does not guarantee a match. Buyers describe the data rather than the businesses, so a precise freshness spec is part of describing the data you need.
Illustrative example: invented to show structure; it does not describe an available dataset.
Freshness requirements (support conversation data, SFT + eval refresh)
- Age field: ticket closed_at (UTC), plus product_version on each ticket
- Median record age at delivery: <= 12 months
- Records older than 36 months: <= 15%, all carrying product_version
- Newest record: within 45 days of export
- Recent-window density (last 90 days vs 12-month average): >= 0.8
- Event-to-export lag: report median and max
- Holdout: last 60 days delivered as a separate file, whole tickets only
- Recurring extract wanted: yes, same schema; cadence to be agreed
Find time-sensitive training data for your models
If your models depend on recent operational records, start with the freshness profile you need and the slices that matter. SourceX looks for US businesses that hold the described data, reviews ownership and consents, and delivers only under a license that defines records, uses, term and delivery, after the supplying company approves the release. Describe your recency requirements at SourceX for AI data buyers.
Sources
- arXiv (Longpre et al.), "A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity" (2023). https://arxiv.org/pdf/2305.13169
- arXiv (Fabre et al.), "Understanding Data Temporality Impact on Large Language Models Pre-training" (2026). https://arxiv.org/pdf/2605.22769
- arXiv, "Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs" (2025). https://arxiv.org/html/2506.07270v1
- arXiv (White et al.; ICLR 2025), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
- arXiv (Gekhman et al.), "Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?" (2024). https://arxiv.org/pdf/2405.05904
- arXiv (Weytjens and De Weerdt), "Creating Unbiased Public Benchmark Datasets with Data Leakage Prevention for Predictive Process Monitoring" (2021). https://export.arxiv.org/abs/2107.01905
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Artificial intelligence - Data quality for analytics and machine learning - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
- Google Research (Pushkarna, Zaldivar, Kjartansson), FAccT 2022, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.