Data quality, coverage and contamination
Temporal Coverage in Licensed Datasets: Date Ranges, Gaps and Seasonality
Quick answer
Temporal coverage is whether a dataset's records are spread evenly and plausibly across the period a supplier says it covers. To verify it, bin records per month by source system and by timestamp field, then look for empty months, step changes, migration seams and spikes. Confirm what each timestamp means and which time zone it uses. Finally, decide whether your task needs complete annual cycles, because a "three-year" export that thins out in year one behaves like a much shorter one.
By SourceX Editorial · Updated
Why the date range on a data sheet is not enough
A stated date range tells you the minimum and maximum timestamps, not how records are distributed between them. A ticket export described as "January 2021 to June 2024" can hold 2% of its rows before 2023 because an older help desk was only partly migrated. That dataset is effectively 18 months of history with a long, thin tail.
Time also matters to model quality directly. Longpre et al. measured that the age of pretraining data relative to evaluation data changes downstream performance, and that the effect does not disappear with fine-tuning [1]. The practical consequence is that a model trained on one period can underperform on records from a later one, and random train/test splits that mix periods tend to hide that. Representativeness is defined against a target population, and that population moves over time [2], so a skewed time distribution is a coverage defect, not a cosmetic one.
This page covers span, gaps and cycles across the history you license. Recency and refresh cadence are a separate question, covered in how often AI buyers want fresh data; shifts in label meaning across years are covered in code-set revisions in multi-year operational datasets.
The records-per-month histogram, done properly
The core artifact is a monthly count table split by source system and by each timestamp field, not one global histogram. A single chart hides the most common failure, where one system's volume falls to zero while another's rises to replace it.
Build it in four passes:
- Parse every timestamp column to UTC with an explicit source offset, and count unparseable or null values per column. A null
closed_aton 30% of rows is a completeness finding; a nullcreated_atis a delivery defect. - Group by
source_systemxmonth(created_at)and, separately, bymonth(closed_at)ormonth(updated_at). Divergence between the two shapes tells you about backlog and resolution lag. - Normalize against a known driver where one exists: active customers, open accounts, invoices issued or headcount. Raw volume growth can reflect business growth rather than better coverage.
- Flag rule-based anomalies: zero-count months, any month below 25% or above 300% of its trailing 6-month median, first and last months per system, and months where the share of a key category moves by more than 10 points.
The thresholds above are starting points to tune, not standards. Pair the histogram with the schema, range and null checks in validation checks for structured dataset deliveries, because many temporal anomalies are really parsing failures.
Illustrative example: invented to show structure; it does not describe an available dataset.
| month | source_system | records (created_at) | records (closed_at) | trailing 6-mo median | flag |
|---|---|---|---|---|---|
| 2022-09 | legacy_helpdesk | 4,210 | 3,980 | 4,050 | none |
| 2022-10 | legacy_helpdesk | 1,140 | 2,960 | 4,100 | about 28% of median, just above the 25% threshold; investigate migration |
| 2022-10 | new_crm | 2,870 | 310 | n/a | first month for system |
| 2022-11 | legacy_helpdesk | 0 | 1,050 | 3,600 | zero created; closures continue |
| 2022-11 | new_crm | 4,390 | 3,720 | n/a | none |
| 2023-02 | new_crm | 13,800 | 4,100 | 4,300 | spike above 300%: check bulk import |
Read across rows, this sample shows a clean cut-over in October 2022, a tail of legacy closures that must be joined to their creation records, and a February spike that is probably a backfill rather than real demand.
Timestamp semantics and time zones
Every timestamp field encodes a different event, and choosing the wrong one silently reshapes your coverage. Ask the supplier, field by field, what writes the value and when.
Common fields in operational records and what they actually mean:
created_at: when the record was inserted, which after a migration may be the import date rather than the original event date.opened_at/submitted_at: when the business event started; often the field you want for modeling.closed_at/resolved_at: when the case ended; reopened cases may overwrite it.updated_at/last_modified: changes on any edit, including bulk updates, so it is a poor coverage axis.exported_at/ extract date: a property of the delivery, not the data.
Time zones produce subtler defects. A system that stores local time without an offset will create a one-hour duplicate or gap at each daylight saving transition, and records near midnight can shift across day and month boundaries. Store UTC alongside the original local value and the source zone. Process-mining formats such as OCEL 2.0 make this explicit by modeling each event with its own timestamp and linked objects, and exchanging logs in SQLite, XML or JSON [3]; asking for an event-log shape rather than a flattened case table makes semantics auditable.
If any of these questions cannot be answered, write them into your request up front. The guide to writing a data request for suppliers shows how to specify the timestamp fields you need.
System migrations, backfills and boundary duplicates
Migrations between systems are the most common cause of gaps and duplicates in multi-year operational data. When a company moves from one help desk, CRM or ERP to another, history is usually handled in one of three ways: fully migrated, partially migrated (open items only, or the last N months), or left behind in an archive.
Failure modes to test at each boundary month:
- Duplicate records that exist in both systems with different IDs. Match on external reference fields, customer plus subject plus opened date, or text similarity; near-duplicate detection with MinHash and LSH handles free-text cases.
- Import-date collapse, where thousands of migrated rows share one
created_atvalue. A single day holding a month's volume is the signature. - Truncated threads, where only the final message or status was carried over. See completeness checks for case and ticket records.
- Field mapping drift, where a status or category picklist was remapped, so the same outcome has different codes on either side of the seam.
Ask the supplier for a migration log: cut-over date, what was moved, and any IDs that map old to new. It turns a mystery dip into a documented boundary.
Seasonality and how many years of history you need
Seasonality is a repeating calendar pattern in volume or content, and whether you need complete annual cycles depends on whether your task's inputs change with the calendar. Support volume rises after product launches and holidays, finance workflows cluster at month-end and fiscal year-end, and tax, enrollment or insurance renewal records concentrate in fixed windows.
To see it, decompose the monthly series into trend, seasonal and remainder components. STL (seasonal-trend decomposition using Loess) is a standard method [7]; implementations expose a seasonal window, a trend window and a robust option that limits the influence of outlier months such as a migration spike. A seasonal component that is large relative to the remainder means a model trained on part of a year will see a biased mix of request types.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Use case | Calendar sensitivity | Practical history target | Why |
|---|---|---|---|
| SFT on support replies | Medium | At least 2 complete cycles | Holiday, launch and billing-cycle topics recur yearly |
| Finance close or AP workflow agents | High | 2 to 3 fiscal years | Month-end, quarter-end and year-end steps differ |
| Classification of document types | Low | 12 months often adequate | Form types change slowly; check for template revisions |
| Temporal holdout evaluation | High | Training years plus a later held-out window | The test window must post-date training data |
| Forecasting-adjacent tasks | Very high | 3+ cycles | Fewer cycles cannot separate trend from season |
One cycle shows a pattern once; two let you tell seasonality from a one-off event. More years are not automatically better, because older records may reflect retired products, policies or code sets.
Splitting by time for training and evaluation
Time-aware splits are the main defense against overstated results on time-stamped data. Random splits let near-identical records from the same week sit in both train and test, which inflates scores. They also cannot show how a model ages, and Longpre et al. found that a gap between training and evaluation periods degrades performance [1].
Use a cutoff date for evaluation, keep a buffer of a few weeks between train and test to limit leakage through long-running cases, and confirm each split covers the seasonal phases you care about. For feature-based work, as-of joins prevent future values from leaking into training rows; see point-in-time correct training data. For evaluation built from later records, see post-cutoff evaluation data as temporal holdouts.
If records are health data de-identified under HIPAA Safe Harbor, all date elements other than the year that relate to an individual must be removed, which rules out monthly analysis on those dates; the Safe Harbor dates guide explains the trade-off.
Temporal acceptance checklist for a multi-year delivery
A delivery passes temporal review when its span, density and semantics match what you agreed and what your task needs. Use this before signing off.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Min and max of each timestamp field match the agreed range, per source system.
- No zero-count months inside the range, or each one is explained in writing.
- Month-over-month changes above your threshold are tied to a known cause (migration, launch, backfill, policy change).
- Each timestamp field has a documented meaning and source time zone; UTC values are provided or derivable.
- Migration boundaries are identified, with a duplicate check across the seam.
- No single day holds an implausible share of a month's records.
- Category and outcome mix per quarter is stable or explained.
- The number of complete annual or fiscal cycles meets the target for your use case.
- A held-out later window exists for evaluation, if needed.
- Findings are recorded in the dataset card or acceptance memo.
Record the outcome. Dataset cards are meant to tell later users about a dataset's limitations and potential biases [6], and a monthly coverage chart with explained anomalies is one of the most useful things to put there. The NIST AI RMF's MAP and MEASURE functions give a structure for documenting such data characteristics in a risk program [4]. For high-risk systems in the EU, Article 10 of the AI Act requires training, validation and testing data to meet stated quality criteria [5]; as of October 2026 the high-risk application dates have reportedly been moved by Regulation (EU) 2026/1744, so check the current timeline with counsel.
For the wider set of quality dimensions, start at the training data quality assessment hub, and use coverage gap analysis to map the time dimension against your deployment distribution.
Where temporal coverage fits in a SourceX request
When you describe a dataset to SourceX, the period, the timestamp fields and the cycles you need belong in the description. SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents and finance and legal workflows; it does not hold stock, and a request does not guarantee a match. You describe the data, not the businesses, and every release is approved by the supplying company.
Each dataset is rights-reviewed and delivered under a license that defines the records, uses, term and delivery, with diligence materials on source and preparation prepared per dataset. That is the point to ask for the migration log and field definitions above. You can describe the history and date coverage you need on the buyers page.
Request time-stamped operational records
If your project depends on multi-year history with known date coverage, describe the period, systems and timestamp fields you need. SourceX looks for US businesses that hold that data and manages the process from assessment through license and delivery; nothing is contracted until a supplier agrees. Start a buyer request.
Sources
- arXiv (Longpre et al.), "A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity" (2023). https://arxiv.org/pdf/2305.13169
- arXiv, "Representativeness in Statistics, Politics, and Machine Learning" (2021). https://arxiv.org/pdf/2101.03827
- OCEL standard authors (arXiv:2403.01975), "OCEL (Object-Centric Event Log) 2.0 Specification" (2023). https://arxiv.org/pdf/2403.01975
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- Hugging Face, "Create a dataset card (Datasets library documentation)". https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card
- Cleveland et al., "STL: A Seasonal-Trend Decomposition Procedure Based on Loess" (1990). https://www.jstor.org/stable/1390761
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.