Tables, time series and transactional data
Sales and Order Histories for Demand Forecasting Models
Quick answer
Demand forecasting training data is a multi-year history of units sold or ordered at SKU × location × day (or week) grain, joined to the covariates that move demand: price, promotions, holidays, inventory position and product lifecycle dates. Buyers should license it with product and location hierarchies, explicit stockout or availability flags so censored sales can be separated from true demand, and enough history to hold out a horizon-gapped backtest. Public benchmarks are useful starting points, but production-grade models usually need real operational histories from businesses that sell.
By SourceX Editorial · Updated
Why public retail benchmarks are not enough for production forecasting
Public benchmarks show the shape of the problem but typically cover one retailer, one period and a narrow covariate set. The best-known public retail competitions use hierarchical item-store unit sales aggregated upward to departments, categories, stores and regions, with prices, calendar events and many intermittent series that have frequent zero-sale days.
That design is a good template for what to ask suppliers for, but a single benchmark cannot represent B2B distributors with order-line demand, grocery with short shelf lives, or specialty retail with heavy assortment churn. Forecasting foundation models are already pretrained on very large real corpora; Google's TimesFM used roughly 100 billion real-world time points [1]. Recent work frames fine-tuning those models on domain data as the next step for practitioners [2], which is exactly where licensed retail and distribution histories earn their place. For pretraining-scale corpora, see time-series foundation model pretraining data.
Grain and hierarchy: SKU × location × day, with masters you can reconcile
The core table should be one row per SKU, location and period, built from point-of-sale lines, e-commerce orders or ERP sales-order lines, with a stated rule for returns and cancellations. Ask whether units are net of returns, whether the date is order date, ship date or invoice date, and how multi-unit packs and kits are counted. Mixing order date and ship date within one series is a common silent error that shifts demand by days.
Hierarchies matter as much as the facts. A product master (SKU → subcategory → category → department, plus brand and pack size) and a location master (store or warehouse → region → channel) let you train at multiple levels and reconcile forecasts so that item-level forecasts sum consistently to category, store and total levels. Request stable surrogate keys and a mapping table for SKU and store renumbering; without it, a relaunched item looks like a new product and history is lost. Our guide to table grain, keys and history covers the general specification.
Censored demand: stockouts and availability flags
Recorded sales understate demand whenever the item was out of stock, so a dataset without availability information trains models to forecast scarcity. Ask suppliers for daily on-hand inventory, an out-of-stock flag, or at minimum a "zero sales with positive on-hand" versus "zero sales with zero on-hand" distinction. For e-commerce, page-level availability or "add to cart blocked" events serve the same purpose.
Also ask how the supplier handles allocation and substitution. If a store received a reduced allocation, or customers switched to a neighboring pack size, sales for both SKUs are distorted. These signals usually live in replenishment and order-management systems, so cross-reference ERP transaction and master data and the supply-chain workflows described on our supply chain and logistics datasets page.
Covariates: price, promotions, calendars and what is known in advance
Covariates should be delivered as separate, time-stamped tables, each tagged as known-future (planned price, promotion calendar, holidays, store openings) or past-only (realized weather, competitor activity, traffic). Leaking a past-only covariate into the forecast horizon is one of the most common reasons backtests look better than production. Promotion tables should carry mechanic (percent off, BOGO, multibuy), display or feature flags, and start and end dates, not just a binary indicator.
Price should be the shelf or list price plus the effective selling price, because markdowns and loyalty discounts differ. The time series with covariates page goes deeper on calendar and price feeds.
Product lifecycle: launches, substitutions and discontinuations
Lifecycle events need a dated product master: first-sale date, planned and actual discontinuation date, and a predecessor-successor map for replacements. Cold-start models for new items learn from analogous launches, so the history of past launches with their attributes is itself training data. Without discontinuation dates, a model cannot tell a dead SKU from a sustained stockout.
Request specification for a demand forecasting dataset
Use the table below as a starting brief when you describe what you need to a supplier or an intermediary.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Element | What to specify | Why it matters |
|---|---|---|
| Fact table | sku_id, location_id, date, units_sold, units_returned, net_revenue, order_channel | Defines grain and net-versus-gross logic |
| History | Number of years, including at least two full seasonal cycles | Seasonality and year-over-year effects |
| Product master | category path, brand, pack_size, launch_date, discontinue_date, successor_sku | Hierarchy, reconciliation, cold start |
| Location master | location_type, region, channel, open_date, close_date | Hierarchy, new-store effects |
| Availability | on_hand_units or oos_flag per SKU-location-day | Separates censored sales from demand |
| Price | list_price, effective_price, markdown_flag | Price elasticity without leakage |
| Promotions | promo_id, mechanic, display_flag, start_date, end_date, known_in_advance | Lift modeling, known-future covariates |
| Calendar | holidays, events, fiscal week mapping | Aligns retail and Gregorian calendars |
| Data dictionary | units, time zone, date semantics, null rules | Avoids silent misinterpretation |
| Format | Parquet partitioned by date, or a share via an open protocol such as Delta Sharing [5] | Scales to billions of rows |
Pair the brief with quality measures you can test on arrival, such as completeness, consistency and timeliness; ISO/IEC 5259-2 defines data quality measures for analytics and ML datasets [4]. A data dictionary with metric definitions prevents the most frequent misreadings of revenue and unit fields.
Backtesting and holding out data you can trust
Evaluation must use chronological splits with a gap equal to the forecast horizon between training and test windows, because random splits leak future information into training [3]. Keep the most recent period, ideally including a peak season, completely out of training and out of any data shared with model vendors. Score every aggregation level, from total down to SKU-location: a multi-level scorecard catches models that do well on totals but fail on intermittent SKUs.
If the same histories will benchmark third-party or foundation models, treat the holdout as a separate asset; our page on held-out business time series for evaluating forecasting models explains why contamination matters.
Licensing and privacy checks for sales histories
Sales histories look impersonal, but order lines often carry customer IDs, loyalty numbers, shipping addresses and B2B account names. Ask suppliers to aggregate to SKU-location-day, or to remove and replace customer identifiers before release, and confirm which fields remain. Check that the license names forecasting model training and evaluation as permitted uses; a large audit of AI training datasets found license omission rates above 70% on popular dataset hosting sites [6], so do not assume terms carry over from a portal or a reseller.
Retailers and distributors also have contracts with brands and marketplaces that may restrict sharing of sell-through data, so ask the supplier whether vendor or marketplace agreements cover the fields you want. If customer contracts or data processing agreements apply, read our page on customer contracts and DPAs for AI training use.
How SourceX sources sales and order histories
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Nothing is held in stock, and a request does not guarantee a match. You describe the data, not the businesses; SourceX looks for US businesses that hold it, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded, and a sample is checked, though no method is perfect.
The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Delivery runs through private, access-controlled workflows after an executed agreement. Teams anywhere can submit a buyer request on SourceX. Related context sits on the structured data buyer's guide, the multi-brand retail buyers page and logistics operations AI.
License sales histories for your demand forecasting models
Describe the grain, history length, covariates and stockout signals you need, and SourceX will look for US businesses that hold that data; nothing is contracted until a supplier agrees. Every dataset is rights-reviewed and delivered under a license that defines the records, uses, term and delivery. Start a demand forecasting data request.
Sources
- Google Research, "A decoder-only foundation model for time-series forecasting" (2024). https://research.google/blog/a-decoder-only-foundation-model-for-time-series-forecasting/
- arXiv, "Foundation Models and Fine-Tuning: Toward a New Generation of Models for Time Series Forecasting" (2026). https://arxiv.org/pdf/2607.23146
- temporalcv documentation, "Why time series is different". https://temporalcv.readthedocs.io/en/latest/guide/why_time_series_is_different.html
- ISO, "ISO/IEC 5259-2:2024 Artificial intelligence: Data quality for analytics and machine learning, Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- Databricks, "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing" (2021). https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.