Tables, time series and transactional data
Time-Series Foundation Model Pretraining Data: What Real-World Series to License
Quick answer
To extend a time-series foundation model (TSFM) beyond public archives, license real business series that add what those archives lack: domains such as orders, ticket volumes, payments and plant telemetry; a frequency mix from seconds to months; long histories with structural breaks, promotions and intermittency; and clean timestamps and units. Buy for diversity per time point, not raw volume, deduplicate against every evaluation suite you report on, and make sure the license covers pretraining and the resulting weights.
By SourceX Editorial · Updated
Why data, not architecture, now separates time-series foundation models
Data is the main lever left because most TSFMs are trained on overlapping public corpora with similar transformer recipes. Google describes TimesFM as pretrained on roughly 100 billion real-world time points [1], and a scaling-law study for TSFMs reports that forecasting performance improves predictably as training data grows [2], consistent with the power-law behavior seen in other domains [9]. Recent work that deconstructs a strong generic transformer baseline suggests that a well-tuned generic architecture closes much of the gap to specialized designs [3].
The market is crowded: a public arena lists many competing models scored on shared tasks [5], and a 2026 overview of TSFMs treats fine-tuning on domain series as a standard step after pretraining [4]. When everyone draws on the same Monash-style archives, LOTSA (the Moirai pretraining archive), the Chronos training corpora and the GIFT-Eval pretraining split, a model's edge comes from series nobody else has. Size and composition figures for those archives differ between secondary summaries, so check each against its primary paper before you quote it in a model card.
What public archives under-represent
Public time-series archives skew toward a few well-worn domains and miss the messy commercial series that real forecasting users care about. Energy load, traffic sensors, weather, web traffic and the M-competition retail sets dominate; enterprise operational series are scarce because companies rarely publish them.
Gaps worth licensing to fill:
- Business calendars. Fiscal weeks (4-4-5), holiday shifts, month-end closes and payroll cycles that show up in finance and order data.
- Interventions. Promotions, price changes, stockouts and outages, ideally with the event flagged (see time series with covariates).
- Intermittency and counts. Sparse SKU-store demand, spare-parts consumption, claim counts and support escalations, where zeros dominate.
- Structural breaks. Product launches, migrations to a new ERP, COVID-era regime changes and reclassified accounts.
- Reporting artifacts. Restatements, late-arriving records, backfilled values and changes of unit or currency that a real deployment must survive.
- Hierarchies. SKU to category to region totals, and ticket queues to product lines, which support coherent multivariate training.
Synthetic generators, including Gaussian-process kernel compositions used in some open models, produce clean shapes but rarely reproduce these artifacts. Search summaries describe Chronos as mixing public real series with Gaussian-process synthetic data [10], and Chronos-2 as relying on synthetic data for covariate structure.
Diversity axes to specify when you license series
Specify the corpus along explicit axes so suppliers can tell you what they actually hold and you can measure marginal value. Volume alone is a weak purchase criterion; one billion points from one meter fleet often adds less than one hundred million points spread across ten unlike domains.
| Axis | What to ask for | Why it matters for pretraining |
|---|---|---|
| Domain | Orders, invoices, payments, support tickets, web sessions, machine telemetry, inventory, staffing | Cross-domain transfer and zero-shot coverage |
| Frequency | Sub-minute, hourly, daily, weekly, monthly, with native rather than resampled timestamps | Frequency-conditioned models need real mixes, not only aggregates |
| Length | Median and p95 history length; share of series longer than three seasonal cycles | Long-context training and seasonality learning |
| Sparsity | Share of zeros, missing-interval rate, irregular sampling flags | Intermittent-demand behavior |
| Dimensionality | Univariate versus multivariate panels; covariates such as price, promo, calendar | Covariate-aware and multivariate heads |
| Regime | Known break dates, interventions, outages | Robustness to non-stationarity |
| Provenance | Source system (SAP, NetSuite, Zendesk, OSIsoft PI, a historian, a data warehouse), extraction query, timezone | Lets you debug artifacts and document lineage |
For sensor-heavy needs such as vibration or SCADA tags, the generic licensing route is covered on licensing sensor and IoT data, and what makes sensor and IoT data valuable for AI explains signal quality. This page focuses on the corpus-design decisions specific to TSFM pretraining.
Real versus synthetic series in the pretraining mix
Use synthetic series for coverage of shapes and real licensed series for the behavior that only operations produce. The two are complements: synthetic data is cheap, unlimited and carries few third-party rights constraints, while real series carry business calendars, interventions and measurement quirks that generators omit.
A tabular analog is instructive. TabPFN was pretrained only on synthetic tables, and Real-TabPFN reports gains from a continued-pretraining stage on curated real-world data [6]; the same staged pattern is a reasonable hypothesis for TSFMs. The tabular foundation model data-mix page covers that evidence, and the broader trade-offs are on licensed vs synthetic vs scraped data and combining licensed and synthetic data.
| Situation | Lean toward | Reason |
|---|---|---|
| Early pretraining, shape coverage | Synthetic plus public archives | Cheap breadth; no license overhead |
| Zero-shot gaps on business domains | Licensed real series | Generators miss calendars and interventions |
| Covariate-aware forecasting | Real panels with flagged covariates | Causal links between promo, price and demand are hard to simulate |
| Intermittent or count demand | Real sparse series | Synthetic zero patterns are usually too regular |
| Continued pretraining for a vertical | Real series from that vertical | Domain shift dominates |
Deduplication and contamination against evaluation suites
Deduplicate every licensed and public series against the benchmarks you report, or your zero-shot claims will not hold up. Benchmark sets such as the Monash archive, the M4 and M5 competitions, GIFT-Eval and the datasets in public arenas [5] are often re-hosted inside pretraining archives under different names, frequencies or truncations.
Practical checks:
- Hash normalized windows (z-scored, fixed length, rounded) across both corpora to catch exact and rescaled copies.
- Compare resampled versions: an hourly training series aggregated to daily can match a daily test series.
- Check for overlapping date ranges from the same source entity, even when values were transformed.
- Hold out whole suppliers or whole time spans for evaluation, not random windows, which leak seasonality.
The companion page on held-out business time series for evaluating forecasting models covers building private test sets. For as-of logic and late-arriving records, see point-in-time correct training data.
Rights, confidentiality and documentation for licensed series
License terms must cover pretraining, continued pretraining and the model weights you release, not only internal analytics. Many public datasets carry missing or wrong license metadata; one large audit found license omission above 70% and error rates above 50% on popular hosting sites [7], so re-verify the terms of every public archive component as well.
Business series raise a confidentiality issue that text corpora do not: a single company's daily revenue, order counts or churn can be competitively sensitive even without personal data. Ask suppliers about aggregation, value scaling, date shifting and removal of entity labels, and agree which transformations are acceptable before extraction. If series derive from customer-level records, personal identifiers should be removed before aggregation.
If you provide a general-purpose model in the EU, Article 53 obligations, which have applied since 2 August 2025, include technical documentation, a copyright policy and a public summary of training content; as of October 2026, the AI Office's enforcement powers have applied since 2 August 2026 [8]. Keep a per-series manifest so that summary and your model card can be produced from records rather than recollection. The rights grant itself is covered in pre-training data license rights.
Request template for a TSFM pretraining corpus
A precise request lets a supplier say quickly whether it holds matching data. Describe the series, not the companies you hope will supply them.
Illustrative example: invented to show structure; it does not describe an available dataset.
request: tsfm-continued-pretraining
intended_use: pretraining and continued pretraining; derived weights may be released
domains:
- b2b order lines aggregated to sku-warehouse-day
- support ticket arrivals by queue-hour
- card payment authorizations by merchant-category-15min
- compressor telemetry by tag-1s (historian export)
frequency_mix: {sub_minute: 15%, hourly: 30%, daily: 40%, weekly_monthly: 15%}
min_history: three seasonal cycles at native frequency
covariates_wanted: [price, promo_flag, holiday_flag, outage_flag]
known_breaks: list dates of migrations, reorganizations, reclassifications
transformations_allowed: [value_scaling, entity_label_removal, date_shift_by_constant]
exclusions: series published in Monash, M4, M5, GIFT-Eval or other public benchmarks
manifest_fields_per_series:
- series_id (opaque)
- domain
- native_frequency
- start_ts, end_ts, timezone
- unit, currency
- aggregation_level
- missing_rate, zero_share
- source_system
- transformations_applied
format: Parquet, long layout (series_id, ts, value, covariate columns)
How SourceX fits a TSFM data request
SourceX sources operational datasets from US companies on request, including sales histories, support records, engineering records and finance workflows, and manages licensing and ongoing purchases. Data is not held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery; personal details such as names, emails and account numbers are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. You can describe the series you need to SourceX and work through Find, Assess, Agree, Transact and Manage, with nothing contracted until a supplier agrees.
For broader context, the structured data buyer's guide and the page on pretraining data sourcing for foundation-model teams cover adjacent decisions.
Request real business time series for TSFM pretraining
SourceX sources operational time series from US companies on request, under licenses that define records, uses, term and delivery, with each release approved by the supplying company. Describe the domains, frequencies and history you need at sourcex.si/buyers.
Frequently asked questions
Is TimesFM's training data available to license?
No public listing makes TimesFM's full corpus licensable; Google describes it at a high level as about 100 billion real-world time points [1]. Teams extending their own models typically combine public archives with synthetic series and separately licensed business data.
How much licensed data moves the needle?
Scaling results suggest more data helps [2], but marginal value depends on novelty. Measure it with ablations: hold the public mix fixed, add each licensed domain, and track zero-shot error on held-out suppliers.
Can I use LOTSA or Chronos corpora commercially?
Check each component dataset's license; archive-level terms do not override the terms of individual sources, and license metadata is often incomplete [7].
Sources
- Google Research, "A decoder-only foundation model for time-series forecasting" (2024). https://research.google/blog/a-decoder-only-foundation-model-for-time-series-forecasting/
- arXiv, "Towards Neural Scaling Laws for Time Series Foundation Models" (2024). https://arxiv.org/pdf/2410.12360
- arXiv, "Revisiting the Generic Transformer: Deconstructing a Strong Baseline for Time Series Foundation Models" (2026). https://arxiv.org/pdf/2602.06909
- arXiv, "Foundation Models and Fine-Tuning: Toward a New Generation of Models for Time Series Forecasting" (2026). https://arxiv.org/pdf/2607.23146
- TS Arena (Hugging Face Space), "TS Arena: models". https://dag-upb-ts-arena.hf.space/models
- arXiv, "Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data" (2025). https://arxiv.org/pdf/2507.03971
- Longpre et al., arXiv, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- Hestness et al., Baidu Research (arXiv), "Deep Learning Scaling is Predictable, Empirically" (2017). https://arxiv.org/abs/1712.00409v1
- arXiv, "Chronos: Learning the Language of Time Series" (2024). https://arxiv.org/abs/2403.07815
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.