Tables, time series and transactional data
Held-Out Business Time Series for Evaluating Forecasting Foundation Models
Quick answer
Public time series forecasting benchmark datasets are the right starting point, but they cannot prove zero-shot skill on their own: many of the same public archives feed pretraining corpora, and each model vendor trained on a different mix. Credible evaluation of forecasting foundation models needs a held-out set of business series that were never published, cut with rolling origins and horizon-sized gaps, scored with scale-free point and probabilistic metrics, and shipped with provenance documentation showing the series stayed private.
By SourceX Editorial · Updated
What public forecasting benchmarks measure well, and where they stop
Public suites are strong for reproducible, like-for-like comparison, but they say little about series a model has never seen. Widely used suites aggregate public archives from energy, traffic, weather, web traffic and retail, with fixed horizons and standardized scoring. The M5 competition's hierarchical retail sales data remains one of the most reused public forecasting datasets [5].
These suites are public by design, which is exactly the limitation. Once a dataset is downloadable, it can end up inside a pretraining corpus, directly or through a derived archive. Google describes TimesFM as pretrained on a large corpus that mixes real-world and synthetic series [4], and other model families draw on their own large public archives. Use public benchmarks for regression tracking and for comparability with published numbers. Use private held-out series to answer the question your stakeholders actually ask: how will this model do on our kind of data?
For the broader decision of when private data is worth the cost, see private evaluation sets versus public benchmarks.
Why pretraining overlap breaks zero-shot comparisons
Overlap means a "zero-shot" score may partly measure memorization, and it does so unevenly across models. Foundation-model papers present zero-shot forecasting as their headline capability [1], and that claim only holds if the test series are absent from training. Public arenas now compare many time-series foundation models trained on different corpora [2], so two models can face the same benchmark with very different exposure to it.
Some benchmark maintainers now publish pretraining splits designed to exclude their test sets, such as the non-leaking pretraining data provided with the GIFT-Eval benchmark [7]. That helps models trained on those splits, but a vendor model trained on an undisclosed corpus can still have seen part of a public benchmark. Overlap also hides in less obvious places:
- Derived copies. Resampled, aggregated or renamed versions of the same public series (hourly electricity rolled to daily, store-item sales rolled to store level).
- Overlapping time windows. A model trained on a series through 2023 and tested on the same series' 2024 values has seen the history, so it is not zero-shot even if the test window is new.
- Shared generating process. Different public uploads drawn from the same upstream source, such as the same utility, exchange or retail panel.
- Synthetic augmentations. Synthetic series generated from public seeds that keep their seasonality and level.
The practical fix is a set of series no model could have seen, plus a written check for each model's disclosed corpus. The same logic, applied to code, appears in held-out coding agent evaluation sets. The general method for screening licensed data against training corpora is covered in the guide to contamination checks for licensed evaluation data.
Protocol: rolling-origin evaluation with horizon-sized gaps
A rolling-origin protocol, with a gap of at least one horizon between any training or tuning data and the first test window, is a sound default for time-series evaluation. Random train-test splits leak future information, so validation needs chronological splits and horizon-sized gaps between training and test windows [3]. For a zero-shot model there is no training split on your data and each forecast starts at its origin, but the same discipline still applies to any fine-tuning arm and any hyperparameter tuning.
Specify the protocol before you request data, because it drives the history length you need:
- Horizons by frequency. For example, 14 and 28 steps for daily data, 13 and 26 for weekly, 24 and 168 for hourly.
- Number of origins. Several forecast origins per series, spaced across different seasons, so one holiday or outage does not dominate the score.
- Context length. Fixed per model where the model allows it, and recorded for each run.
- Fine-tuned arm. If you also compare fine-tuned models, freeze a chronological training cutoff and keep a gap before the first test origin.
Choices such as the number of origins, the gap length and the aggregation method can reorder a leaderboard. Write the protocol down, version it, and score every model with the same code.
Metrics: scale-free point accuracy plus probabilistic calibration
Report at least one scale-free point metric and one probabilistic metric, because business series differ in scale by orders of magnitude and most decisions use quantiles. MASE scales errors by an in-sample naive forecast, so a spare part selling three units a month and a fast-moving SKU selling thousands can share one table. CRPS scores the whole predictive distribution, which matters when the forecast feeds safety stock or staffing levels, so report both side by side.
Add metrics that fit the domain. For intermittent demand, report scaled errors on non-zero periods separately and check calibration of the probability of zero. For hierarchical data, check coherence between SKU, store and region levels. Aggregate with geometric means of relative scores against a seasonal-naive baseline, so a few volatile series do not dominate.
What held-out business series should cover
The most useful held-out sets cover the operational patterns that public suites under-represent. Public archives lean toward energy, traffic, weather, web traffic and retail panels. The patterns forecasting models break on in enterprise use are often elsewhere:
- B2B order books. Lumpy orders from a small number of accounts, contract renewals, and price changes that shift volume between periods.
- Service operations. Ticket arrivals, call volumes and field-service jobs with intraday seasonality, staffing feedback and outage spikes.
- Intermittent spare-parts demand. Long runs of zeros, occasional large pulls, and part supersessions that end one series and start another.
- Finance workflows. Invoice counts, receivables aging and payment timing with month-end and quarter-end effects.
- Structural breaks. Catalog changes, new sites, system migrations and promotions, recorded as dated events rather than left as unexplained shifts.
Ask for at least two to three full seasonal cycles of history per series where the business has it, plus the event log. If your evaluation includes covariates, specify them separately; see multivariate time series with covariates. Pretraining data is a different purchase with different volume needs, covered in time-series foundation model pretraining data.
Provenance documentation that makes "unseen" credible
An evaluation set is only "held out" if you can show where it came from and that it was never released. Ask the supplier to document the source system, extraction date, any transformation, and a statement that the series were not published, shared on a data marketplace or submitted to a competition. Structured documentation in the style of Data Cards covers upstream sources, collection methods and intended use [6], and is easy to attach to an evaluation report.
Keep the set held out after delivery too. Restrict access to the evaluation team, never send series to a hosted model API that may log or train on inputs unless the terms rule that out, and rotate in fresh origins over time as old test windows age into possible training data.
Specification artifact: forecasting evaluation data request
The request below shows the fields an evaluation lead should define before talking to any supplier.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example specification |
|---|---|
| Purpose | Zero-shot and fine-tuned comparison of 5 forecasting foundation models against seasonal-naive and ETS baselines |
| Domains | B2B distributor order lines; field-service job counts; spare-parts withdrawals |
| Grain and frequency | SKU-by-warehouse daily; region weekly; part-by-depot weekly |
| History | 36 months minimum per series; series start and end dates recorded |
| Volume | 500 to 2,000 series per domain, including intermittent series |
| Record schema | series_id (pseudonymous), timestamp (ISO 8601), value, unit, hierarchy_path, event_flags |
| Event log | Promotions, price changes, stockouts, system migrations, part supersessions with dates |
| Protocol | Rolling origin; 4 origins per series; gap equal to the horizon; horizons 28 days and 13 weeks |
| Metrics | MASE, CRPS, weighted quantile loss at P10/P50/P90, coherence check |
| Provenance | Source system, extraction date, transformations, statement of non-publication |
| Privacy | Customer and account identifiers removed or replaced; method recorded |
| Format | Parquet, one file per domain, plus a data dictionary |
| Allowed use | Internal model evaluation; no training; no redistribution of series |
How SourceX supports held-out forecasting evaluation sets
SourceX sources operational datasets from US companies on request, including the sales, support, engineering, finance and legal records that business time series are built from. Data is sourced per request rather than held in stock, so a request does not guarantee a match. Buyers describe the data they need, not the businesses, and every release is approved by the supplying company.
Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, allowed uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which supports the provenance record an evaluation report needs. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. SourceX does not train models. You can describe your evaluation data needs to SourceX and see how the evaluation dataset use case works.
For related structured-data purchases, start from the tabular, time-series and transactional data hub.
Sourcing held-out time series for forecasting model evaluation
If public benchmarks no longer separate the models you are comparing, specify the domains, horizons and provenance you need and SourceX will look for US businesses that hold matching data. Nothing is contracted until a supplier agrees, and pricing and allowed uses are set in each license. Request held-out business time series.
Sources
- arXiv, "TimeRAF: Retrieval-Augmented Foundation model for Zero-shot Time Series Forecasting" (2024). https://arxiv.org/pdf/2412.20810
- Hugging Face Space (dag-upb), "TS Arena: models". https://dag-upb-ts-arena.hf.space/models
- temporalcv documentation, "Why time series is different". https://temporalcv.readthedocs.io/en/latest/guide/why_time_series_is_different.html
- Google Research, "A decoder-only foundation model for time-series forecasting" (2024). https://research.google/blog/a-decoder-only-foundation-model-for-time-series-forecasting/
- University of Nicosia, Institute For the Future, "M5 competition". https://www.unic.ac.cy/iff/research/forecasting/m5-competition/
- Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- Salesforce AI Research, "GIFT-Eval: A Benchmark for General Time Series Forecasting Model Evaluation" (2024). https://arxiv.org/html/2410.10393v2
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.