Skip to content

Tables, time series and transactional data

Time-Series Foundation Model Pretraining Data: What Real-World Series to License

Quick answer

To extend a time-series foundation model (TSFM) beyond public archives, license real business series that add what those archives lack: domains such as orders, ticket volumes, payments and plant telemetry; a frequency mix from seconds to months; long histories with structural breaks, promotions and intermittency; and clean timestamps and units. Buy for diversity per time point, not raw volume, deduplicate against every evaluation suite you report on, and make sure the license covers pretraining and the resulting weights.

By SourceX Editorial · Updated

Why data, not architecture, now separates time-series foundation models

Data is the main lever left because most TSFMs are trained on overlapping public corpora with similar transformer recipes. Google describes TimesFM as pretrained on roughly 100 billion real-world time points [1], and a scaling-law study for TSFMs reports that forecasting performance improves predictably as training data grows [2], consistent with the power-law behavior seen in other domains [9]. Recent work that deconstructs a strong generic transformer baseline suggests that a well-tuned generic architecture closes much of the gap to specialized designs [3].

The market is crowded: a public arena lists many competing models scored on shared tasks [5], and a 2026 overview of TSFMs treats fine-tuning on domain series as a standard step after pretraining [4]. When everyone draws on the same Monash-style archives, LOTSA (the Moirai pretraining archive), the Chronos training corpora and the GIFT-Eval pretraining split, a model's edge comes from series nobody else has. Size and composition figures for those archives differ between secondary summaries, so check each against its primary paper before you quote it in a model card.

What public archives under-represent

Public time-series archives skew toward a few well-worn domains and miss the messy commercial series that real forecasting users care about. Energy load, traffic sensors, weather, web traffic and the M-competition retail sets dominate; enterprise operational series are scarce because companies rarely publish them.

Gaps worth licensing to fill:

  • Business calendars. Fiscal weeks (4-4-5), holiday shifts, month-end closes and payroll cycles that show up in finance and order data.
  • Interventions. Promotions, price changes, stockouts and outages, ideally with the event flagged (see time series with covariates).
  • Intermittency and counts. Sparse SKU-store demand, spare-parts consumption, claim counts and support escalations, where zeros dominate.
  • Structural breaks. Product launches, migrations to a new ERP, COVID-era regime changes and reclassified accounts.
  • Reporting artifacts. Restatements, late-arriving records, backfilled values and changes of unit or currency that a real deployment must survive.
  • Hierarchies. SKU to category to region totals, and ticket queues to product lines, which support coherent multivariate training.

Synthetic generators, including Gaussian-process kernel compositions used in some open models, produce clean shapes but rarely reproduce these artifacts. Search summaries describe Chronos as mixing public real series with Gaussian-process synthetic data [10], and Chronos-2 as relying on synthetic data for covariate structure.

Diversity axes to specify when you license series

Specify the corpus along explicit axes so suppliers can tell you what they actually hold and you can measure marginal value. Volume alone is a weak purchase criterion; one billion points from one meter fleet often adds less than one hundred million points spread across ten unlike domains.

AxisWhat to ask forWhy it matters for pretraining
DomainOrders, invoices, payments, support tickets, web sessions, machine telemetry, inventory, staffingCross-domain transfer and zero-shot coverage
FrequencySub-minute, hourly, daily, weekly, monthly, with native rather than resampled timestampsFrequency-conditioned models need real mixes, not only aggregates
LengthMedian and p95 history length; share of series longer than three seasonal cyclesLong-context training and seasonality learning
SparsityShare of zeros, missing-interval rate, irregular sampling flagsIntermittent-demand behavior
DimensionalityUnivariate versus multivariate panels; covariates such as price, promo, calendarCovariate-aware and multivariate heads
RegimeKnown break dates, interventions, outagesRobustness to non-stationarity
ProvenanceSource system (SAP, NetSuite, Zendesk, OSIsoft PI, a historian, a data warehouse), extraction query, timezoneLets you debug artifacts and document lineage

For sensor-heavy needs such as vibration or SCADA tags, the generic licensing route is covered on licensing sensor and IoT data, and what makes sensor and IoT data valuable for AI explains signal quality. This page focuses on the corpus-design decisions specific to TSFM pretraining.

Real versus synthetic series in the pretraining mix

Use synthetic series for coverage of shapes and real licensed series for the behavior that only operations produce. The two are complements: synthetic data is cheap, unlimited and carries few third-party rights constraints, while real series carry business calendars, interventions and measurement quirks that generators omit.

A tabular analog is instructive. TabPFN was pretrained only on synthetic tables, and Real-TabPFN reports gains from a continued-pretraining stage on curated real-world data [6]; the same staged pattern is a reasonable hypothesis for TSFMs. The tabular foundation model data-mix page covers that evidence, and the broader trade-offs are on licensed vs synthetic vs scraped data and combining licensed and synthetic data.

SituationLean towardReason
Early pretraining, shape coverageSynthetic plus public archivesCheap breadth; no license overhead
Zero-shot gaps on business domainsLicensed real seriesGenerators miss calendars and interventions
Covariate-aware forecastingReal panels with flagged covariatesCausal links between promo, price and demand are hard to simulate
Intermittent or count demandReal sparse seriesSynthetic zero patterns are usually too regular
Continued pretraining for a verticalReal series from that verticalDomain shift dominates

Deduplication and contamination against evaluation suites

Deduplicate every licensed and public series against the benchmarks you report, or your zero-shot claims will not hold up. Benchmark sets such as the Monash archive, the M4 and M5 competitions, GIFT-Eval and the datasets in public arenas [5] are often re-hosted inside pretraining archives under different names, frequencies or truncations.

Practical checks:

  1. Hash normalized windows (z-scored, fixed length, rounded) across both corpora to catch exact and rescaled copies.
  2. Compare resampled versions: an hourly training series aggregated to daily can match a daily test series.
  3. Check for overlapping date ranges from the same source entity, even when values were transformed.
  4. Hold out whole suppliers or whole time spans for evaluation, not random windows, which leak seasonality.

The companion page on held-out business time series for evaluating forecasting models covers building private test sets. For as-of logic and late-arriving records, see point-in-time correct training data.

Rights, confidentiality and documentation for licensed series

License terms must cover pretraining, continued pretraining and the model weights you release, not only internal analytics. Many public datasets carry missing or wrong license metadata; one large audit found license omission above 70% and error rates above 50% on popular hosting sites [7], so re-verify the terms of every public archive component as well.

Business series raise a confidentiality issue that text corpora do not: a single company's daily revenue, order counts or churn can be competitively sensitive even without personal data. Ask suppliers about aggregation, value scaling, date shifting and removal of entity labels, and agree which transformations are acceptable before extraction. If series derive from customer-level records, personal identifiers should be removed before aggregation.

If you provide a general-purpose model in the EU, Article 53 obligations, which have applied since 2 August 2025, include technical documentation, a copyright policy and a public summary of training content; as of October 2026, the AI Office's enforcement powers have applied since 2 August 2026 [8]. Keep a per-series manifest so that summary and your model card can be produced from records rather than recollection. The rights grant itself is covered in pre-training data license rights.

Request template for a TSFM pretraining corpus

A precise request lets a supplier say quickly whether it holds matching data. Describe the series, not the companies you hope will supply them.

Illustrative example: invented to show structure; it does not describe an available dataset.

request: tsfm-continued-pretraining
intended_use: pretraining and continued pretraining; derived weights may be released
domains:
  - b2b order lines aggregated to sku-warehouse-day
  - support ticket arrivals by queue-hour
  - card payment authorizations by merchant-category-15min
  - compressor telemetry by tag-1s (historian export)
frequency_mix: {sub_minute: 15%, hourly: 30%, daily: 40%, weekly_monthly: 15%}
min_history: three seasonal cycles at native frequency
covariates_wanted: [price, promo_flag, holiday_flag, outage_flag]
known_breaks: list dates of migrations, reorganizations, reclassifications
transformations_allowed: [value_scaling, entity_label_removal, date_shift_by_constant]
exclusions: series published in Monash, M4, M5, GIFT-Eval or other public benchmarks
manifest_fields_per_series:
  - series_id (opaque)
  - domain
  - native_frequency
  - start_ts, end_ts, timezone
  - unit, currency
  - aggregation_level
  - missing_rate, zero_share
  - source_system
  - transformations_applied
format: Parquet, long layout (series_id, ts, value, covariate columns)

How SourceX fits a TSFM data request

SourceX sources operational datasets from US companies on request, including sales histories, support records, engineering records and finance workflows, and manages licensing and ongoing purchases. Data is not held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery; personal details such as names, emails and account numbers are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. You can describe the series you need to SourceX and work through Find, Assess, Agree, Transact and Manage, with nothing contracted until a supplier agrees.

For broader context, the structured data buyer's guide and the page on pretraining data sourcing for foundation-model teams cover adjacent decisions.

Request real business time series for TSFM pretraining

SourceX sources operational time series from US companies on request, under licenses that define records, uses, term and delivery, with each release approved by the supplying company. Describe the domains, frequencies and history you need at sourcex.si/buyers.

Frequently asked questions

Is TimesFM's training data available to license?

No public listing makes TimesFM's full corpus licensable; Google describes it at a high level as about 100 billion real-world time points [1]. Teams extending their own models typically combine public archives with synthetic series and separately licensed business data.

How much licensed data moves the needle?

Scaling results suggest more data helps [2], but marginal value depends on novelty. Measure it with ablations: hold the public mix fixed, add each licensed domain, and track zero-shot error on held-out suppliers.

Can I use LOTSA or Chronos corpora commercially?

Check each component dataset's license; archive-level terms do not override the terms of individual sources, and license metadata is often incomplete [7].

Sources

  1. Google Research, "A decoder-only foundation model for time-series forecasting" (2024). https://research.google/blog/a-decoder-only-foundation-model-for-time-series-forecasting/
  2. arXiv, "Towards Neural Scaling Laws for Time Series Foundation Models" (2024). https://arxiv.org/pdf/2410.12360
  3. arXiv, "Revisiting the Generic Transformer: Deconstructing a Strong Baseline for Time Series Foundation Models" (2026). https://arxiv.org/pdf/2602.06909
  4. arXiv, "Foundation Models and Fine-Tuning: Toward a New Generation of Models for Time Series Forecasting" (2026). https://arxiv.org/pdf/2607.23146
  5. TS Arena (Hugging Face Space), "TS Arena: models". https://dag-upb-ts-arena.hf.space/models
  6. arXiv, "Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data" (2025). https://arxiv.org/pdf/2507.03971
  7. Longpre et al., arXiv, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  8. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  9. Hestness et al., Baidu Research (arXiv), "Deep Learning Scaling is Predictable, Empirically" (2017). https://arxiv.org/abs/1712.00409v1
  10. arXiv, "Chronos: Learning the Language of Time Series" (2024). https://arxiv.org/abs/2403.07815

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data