Tables, time series and transactional data
Real-World Tables for Tabular Foundation Models: Pretraining, Data Mix and Held-Out Evaluation
Quick answer
Tabular foundation model training data works best as a layered mix: synthetic priors teach the in-context learning mechanism, and a curated set of licensed real tables then adds, through continued pretraining, the semantics, dependencies and messiness that generators miss. Budget the purchase in independent tables and cells, not rows, and keep column headers and units. Fingerprint every delivered table against OpenML, Kaggle and UCI so it never overlaps your evaluation suites. Hold back a separate set of never-published business tables, under evaluation-only terms, to measure generalization.
By SourceX Editorial · Updated
Why real tables matter after a synthetic prior
Real tables add generalization that a synthetic prior alone does not supply. TabPFN, the reference in-context tabular model, was pretrained on more than 100 million synthetic datasets sampled from structural causal models, with no real-world semantics [3]. The TabDPT authors note that most tabular foundation models are trained this way. They report that adding real data during pretraining gave significantly faster training and better generalization to unseen data [1].
Synthetic generators also have known blind spots. Work on tabular generators finds they can fail to learn the structure of real tables, such as dependencies between columns [5]. Business tables carry constraints that a random causal graph rarely produces: invoice totals that equal line-item sums, status codes that only move in one direction, and foreign keys into customer and product masters. For the general trade-off, see licensed vs synthetic vs scraped data and synthetic data vs. real business data.
How to frame the data mix: prior first, real tables as the adaptation layer
Treat licensed real tables as a continued-pretraining and domain-adaptation stage, not as a replacement for the synthetic prior. Real-TabPFN does exactly this. It starts from a synthetically pretrained TabPFN and adds a real-data stage on top [2]. The synthetic stage keeps the mechanism general, and the real stage pulls the model toward the distributions it will meet in production.
The same paper's key finding is about curation. A small, curated collection of large real-world datasets outperformed broader, noisier corpora such as CommonCrawl- or GitTables-derived tables for continued pretraining [2]. Web-extracted tables are often tiny, lack a meaningful target and repeat the same templates. That makes the buying question less "how many tables can we get" and more "how many independent, large, well-documented tables with a defensible target can we get". The mechanics of blending sources are covered in combining licensed and synthetic data.
Three practical mix decisions follow:
- Stage order. Run synthetic pretraining to convergence, then continued pretraining on real tables at a lower learning rate. Keep a synthetic replay fraction so the model does not forget the prior.
- Weighting by source, not by row. One supplier's 40-million-row event log should not dominate the real stage. Sample tasks per source table, then subsample rows to the context length your architecture supports.
- Task construction. Each real table yields many training tasks: different target columns, feature subsets and row samples. A table with three defensible targets is worth more than three near-identical tables from one system.
How much real tabular data to buy: count tables and cells
Measure a tabular pretraining purchase in independent tables and total cells, not rows. TabDPT's scaling experiments span roughly 52 million to 2 billion cells and show performance following power laws in model and data size [1]. Cells (rows times columns) capture width, which matters for attention over features. Independent tables capture diversity, which matters for in-context generalization.
"Independent" is the term to define in the contract. Twelve monthly exports of the same Salesforce Opportunity object are one table, not twelve. A Snowflake schema with an orders fact and five dimensions is one relational source; it can produce several flat tables, but they share entities and leak into each other. Count by supplier and source system first, then by table. For general sizing guidance, see how much tabular data you need.
Specification checklist for TFM pretraining tables
A useful request specifies diversity, structure and target quality per table, not just volume. Use the checklist below as a starting request and adapt the thresholds to your architecture's context window and feature limits.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | What to specify | Why it matters for a TFM |
|---|---|---|
| Source diversity | Number of supplying companies and source systems (ERP, CRM, ticketing, billing, MES) | Prevents one schema style from dominating the real stage |
| Table count | Independent tables after de-duplication of exports and snapshots | The unit of in-context diversity |
| Shape | Rows and columns per table; share of wide tables (100+ columns) | Tests feature-attention and context limits |
| Column-type mix | Numeric, ordinal, low- and high-cardinality categorical, dates, free-text codes | Synthetic priors under-represent high-cardinality codes such as SKU or GL account |
| Missingness | Native nulls preserved, with a code for "not applicable" vs "unknown" | Real missingness is informative; imputed tables hide it |
| Targets | At least one defensible target per table, with label definition and timing | A table without a target is weak pretraining signal |
| Semantics | Original column headers, units, currency, code lists and data dictionary | LLM-based tabular models use column semantics [4] |
| Time | Event and as-of timestamps on every row | Enables temporal splits and leakage checks |
| Keys | Primary and foreign keys kept, pseudonymized consistently | Enables relational tasks and de-duplication |
| Provenance | Supplier, system, extraction date, never-published attestation | Required for contamination control |
| Format | Parquet with explicit schema, plus a manifest in JSON | Preserves types that CSV loses |
Keep headers and dictionaries even if your current architecture ignores them. Microsoft Research's tabular foundation model work fine-tunes an LLM over serialized tables, where column names and descriptions carry signal [4]. Stripping them to col_1 … col_n at delivery throws away information you cannot recover later. The tabular data license specification page covers grain, keys and history in more depth, and target leakage audits cover label timing.
Keeping pretraining tables out of your evaluation suites
Contamination control is a delivery acceptance step, not an afterthought. Public tabular TFMs that use real data have tended to draw it from public repositories, and Real-TabPFN reports its headline results on datasets from the OpenML AutoML Benchmark [2]. When pretraining and evaluation both come from the same public pools, the risk of overlap is real, and a model can look strong on tables it has effectively seen.
Before accepting a delivery, run three checks:
- Schema fingerprint. Hash the sorted, normalized column names and types of each delivered table and compare them against OpenML, Kaggle and UCI metadata dumps.
- Content fingerprint. Hash row-level MinHash signatures over a sample of rows and compare them against the same repositories and your own benchmark suites.
- Provenance attestation. Record the source system and confirm the table was never published, then log it in the dataset manifest.
The same logic applies in reverse: once a table enters pretraining, tag it so it can never be promoted into an evaluation set. For a full procedure, see contamination checks for licensed eval data.
Building a held-out evaluation suite from private business tables
A private evaluation suite should be a separate purchase from pretraining data, under evaluation-only terms, from tables that were never published. Public suites are useful for comparison with the literature, but they cannot tell you how the model behaves on a mid-market distributor's order lines or an insurer's claims table. The broader argument is on private evaluation sets vs public benchmarks and on AI evaluation datasets built from real business work.
Design the suite as a coverage matrix, so a single headline number does not hide a failure mode.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Slice | Example business table | Failure mode it exposes |
|---|---|---|
| Binary classification, imbalanced | Support tickets with an "escalated" flag at 3% prevalence | Calibration and minority-class recall |
| Multiclass, many classes | Expense lines with a GL account target | Large label spaces beyond the model's class limit |
| Regression, heavy tails | Invoice days-to-pay | Outlier sensitivity and loss choice |
| Small n | 300-row equipment maintenance log | The regime where in-context models should shine |
| Large n | Multi-million-row order history, subsampled into context | Context-selection and retrieval strategy |
| Wide tables | 400-column product master | Feature limits and attention cost |
| High-cardinality categoricals | Customer ID, SKU, ZIP code | Encoding and memorization |
| Temporal drift | Train on 2023 orders, test on 2025 | Distribution shift that random splits hide |
| Relational | Orders plus customers plus products | Whether flattening loses signal; see RelBench [6] |
Fix the splits before anyone trains. Use time-based splits where the table has a natural clock, freeze the test labels, and keep the suite in an access-controlled store that pretraining pipelines cannot read. Multi-table tasks are a different setting from single flat tables [6]; if your roadmap includes relational models, see multi-table relational data for relational deep learning.
License and rights terms specific to TFM pretraining
Pretraining licenses need to cover what happens to the weights, not just the data. Ask for explicit grants for pretraining and continued pretraining, for derivative model weights, and for keeping and deploying trained models after the data term ends. Evaluation tables need different terms: evaluation-only use, no training, and restrictions on publishing rows or per-table scores that would reveal the data.
Each table needs its own rights review, because tables from different suppliers carry different ownership and consent histories. Rows about people (customers, employees, patients) need de-identification before delivery; free-text columns and rare-value combinations are where tabular de-identification usually fails. See de-identifying tabular and transactional data for machine learning and how to license proprietary data for AI training.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
How SourceX fits a tabular model data request
SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, and finance and legal workflows, and it manages licensing and ongoing purchases. Nothing is held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, with personal details removed or replaced before delivery. SourceX does not train models and does not source scraped web content. You can describe the tables you need on the buyer page.
Sourcing real tables for your tabular foundation model
SourceX looks for US businesses that hold the tables you describe, runs an assess and agree step on data and licensing permissions, and delivers only after an executed agreement and supplier approval. Start with your specification checklist and evaluation coverage matrix, and see the structured data buyer's guide, the AI data hub and the time-series foundation model data guide for related requests. Describe your tabular data requirements to SourceX.
Frequently asked questions
Do tabular foundation models need real data at all?
Synthetic-only models such as the original TabPFN perform well [3], but TabDPT reports faster training and better generalization when real data is added [1]. Real tables are most useful as a later, curated stage rather than the whole corpus [2].
Are web-extracted tables a substitute for licensed business tables?
Not for continued pretraining, based on published evidence as of October 2026. Real-TabPFN found a curated set of large real datasets beat broader web-derived corpora [2]. Web tables also rarely carry a defensible target column.
Can the same supplier provide both pretraining and evaluation tables?
It can, but keep them separate at the table and entity level. Split by source system or by time, record the split in the manifest, and treat any shared customer or product keys as leakage risk.
Sources
- arXiv (Layer 6 AI authors), "TabDPT: Scaling Tabular Foundation Models on Real Data" (2024). https://arxiv.org/pdf/2410.18164
- arXiv, "Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data" (2025). https://arxiv.org/pdf/2507.03971
- Nature / PubMed Central, "Accurate predictions on small data with a tabular foundation model (TabPFN)" (2025). https://pmc.ncbi.nlm.nih.gov/articles/PMC11711098/
- Microsoft Research, "Towards Foundation Models for Learning on Tabular Data". https://microsoft.com/en-us/research/?p=1042506
- arXiv, "How Well Does Your Tabular Generator Learn the Structure of Tabular Data?" (2025). https://arxiv.org/pdf/2503.09453
- arXiv, "RelBench: A Benchmark for Deep Learning on Relational Databases" (2024). https://arxiv.org/pdf/2407.20060
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.