Tables, time series and transactional data
Tabular, Time-Series and Transactional Data for AI: A Buyer's Guide
Quick answer
Tabular data for AI spans several families of structured business records: single tables, multi-table relational databases, schemas with SQL, spreadsheets, transactions and ledgers, time series, logs and event streams, and master data or knowledge graphs. Each family trains different models, from tabular foundation models to text-to-SQL agents and forecasters. Real enterprise versions rarely reach the open market, so buyers license them and should fix grain, keys, history, label timing and join-safe de-identification before discussing volume.
By SourceX Editorial · Updated
This hub covers the structured data cluster of the SourceX guide to AI data. For definitions, read structured vs unstructured AI training data.
Eight families of structured data and the models they train
Scope structured data for AI training by family, because each family comes from different systems, trains different models and fails acceptance for different reasons. Settle the last column before discussing volume.
| Family | Typical source systems | Models and tasks | Pin down first |
|---|---|---|---|
| Single wide tables | Warehouse marts, feature stores | Gradient-boosted trees, tabular foundation models (TFMs) | Row meaning, target definition, units |
| Multi-table relational data | ERP, CRM, billing and order databases | Relational deep learning, field autocompletion | Primary and foreign keys, timestamps on every table |
| Schemas, SQL and metric definitions | Warehouse catalogs, query logs, dbt models | Text-to-SQL, schema linking, analytics agents | Column descriptions, SQL dialect, executable gold queries |
| Spreadsheets and workbooks | Finance models, trackers, exported reports | Spreadsheet agents, formula generation, table detection | Native .xlsx with formulas and links intact |
| Transactions and ledgers | Payment processors, general ledgers, accounts payable | Fraud and anomaly detection, reconciliation | Label source and date, dispute history, merchant codes |
| Time series and telemetry | Order histories, sensor historians, observability metrics | Forecasting, time-series foundation models, predictive maintenance | Frequency, gaps, time zones, price and promotion covariates |
| Logs and event streams | Application logs, clickstreams, business event tables | Log anomaly detection, root-cause analysis, recommenders | Incident labels aligned to timestamps, entity IDs |
| Master data and knowledge graphs | Customer, supplier and product masters | Entity resolution, product matching, LLM grounding | Match labels, merge history, licenses of external graph data |
Several record types have their own SourceX pages: spreadsheets and financial models, financial transaction data, sensor and IoT data, sales CRM histories and accounting reconciliations. Licensing operational data for AI covers the wider category.
Why real enterprise tables are hard to find
Real multi-table business data rarely reaches the open market because it sits inside ERP, billing and payment systems under privacy, confidentiality and commercial constraints.
- Scarcity. When SAP released SALT, anonymized sales tables from one customer's ERP system, it said text is plentiful but data with multiple linked tables is scarce, because privacy, confidentiality and commercial interests make company datasets hard to procure [1].
- Tabular foundation models. TabDPT's authors report that existing TFMs are predominantly trained on synthetic data, and that adding real data to pretraining can give significantly faster training and better generalization [2]. Real-TabPFN reports that continued pretraining on a small, curated set of large real datasets beat broader, potentially noisier corpora such as CommonCrawl or GitTables [3].
- Time series. Google Research describes TimesFM as pretrained on 100 billion real-world time points [4], so fair evaluation needs held-out business time series outside public corpora.
| Public resource | What it contains | Gap a licensed dataset fills |
|---|---|---|
| RelBench and v2 (relational) | Prediction tasks over multi-table databases with time-based splits [5]; v2 adds ERP, consumer-platform, clinical and scholarly databases [6] | Few industries, mostly public databases |
| Spider 2.0 (text-to-SQL) | 632 tasks on real-application databases, often over 1,000 columns, on systems such as BigQuery and Snowflake [7] | An o1-preview-based agent solved 17.0% in the arXiv version [8]; your schemas are absent |
| BIRD (text-to-SQL) | 12,751 question-SQL pairs over 95 databases in 37 domains, with dirty values; in the 2023 paper ChatGPT reached 40.08% execution accuracy against 92.96% for humans [9] | No private business logic |
| SpreadsheetBench | 912 forum-sourced instructions; 35.7% of files hold multiple tables and 42.7% non-standard relational tables [10] | Forum problems, not business workbooks |
| Loghub (logs) | 19 real-world log datasets, of which 6 are labeled [11] | Few labels; system, not business, logs |
| Open fraud sets | One study merged IEEE-CIS, Sparkov and a Kaggle e-commerce set and added synthetic attributes after finding open sources insufficient [12] | Partly synthetic; may not match your payment mix |
Specify grain, keys and history before row counts
A structured data request should define what one row means, how tables join, how history is kept and when each label became known; row counts mean little until those are fixed. The tabular license specification guide adds clause-level detail.
- Grain. "One row per invoice line" and "one row per invoice" train different models; mixed grains double-count amounts.
- Keys. Name primary and foreign keys and agree a maximum orphan-row rate. Subsets must keep whole entity graphs (sampling relational extracts without breaking keys).
- History. Ask for change history or dated snapshots, with both business event time and the time each row was written; current-state tables cannot support point-in-time correct training data.
- Labels. Record each label's source and when it became known: a dispute flag set after settlement did not exist when the payment was scored.
- Metadata. Spider 2.0 tasks require database metadata and dialect documentation [7]; ask for column descriptions, code lists, units and metric definitions (data dictionaries as grounding data).
- Coverage. State entities and time span: ten million rows from 200 customers is a narrow sample.
Acceptance thresholds can borrow the data quality model and measures of ISO/IEC 5259-2:2024 [13]; validation checks for structured deliveries turns them into tests. The manifest below applies these points to a four-table extract.
Illustrative example: invented to show structure; it does not describe an available dataset.
dataset: order_to_cash_extract_example
history: change_log # every update kept with valid_from / valid_to
time_zone: UTC
coverage: {from: 2021-01-01, to: 2025-12-31}
tables:
customers:
grain: one row per customer version
primary_key: [customer_token, valid_from]
customer_token: {method: keyed_hash, key_held_by: supplier, same_key_all_tables: true}
postal_code: {method: generalize, keep_digits: 3}
orders:
grain: one row per order
primary_key: [order_id]
foreign_keys: {customer_token: customers.customer_token} # as-of join on valid_from / valid_to
event_time: order_created_at
written_at: row_inserted_at
order_lines:
grain: one row per order line
primary_key: [order_id, line_no]
foreign_keys: {order_id: orders.order_id}
payments:
grain: one row per payment attempt
primary_key: [payment_attempt_id]
foreign_keys: {order_id: orders.order_id}
label: {column: disputed, source: processor_dispute_file, known_at: dispute_opened_at}
free_text_columns: {payment_memo: pii_scanned_and_masked}
acceptance:
duplicate_primary_keys: 0
orphan_foreign_keys_max_pct: 0.1
license:
license_id: LIC-EXAMPLE-01
permitted_uses: [model_training, internal_evaluation]
derived_artifacts: [features, embeddings] # name them explicitly
known_atlets you drop any label or field that did not exist at prediction time.- One keyed hash applied across all tables pseudonymizes customers without breaking joins.
event_timeandwritten_atseparate when something happened from when the system recorded it.
Leakage: the defect that makes a weak table look strong
Leakage means a training table holds information that would not exist at prediction time. In licensed business data it usually enters through post-outcome fields, random splits over time-ordered rows, or the same entity on both sides of a split.
- Post-outcome fields. A
write_off_reasonin a credit table orresolved_byin a ticket table is filled only after the outcome. Ask which columns are written after the event you predict. - Time. Split chronologically; for time-series validation, one library recommends a gap between training and test equal to the forecast horizon [14]. RelBench uses time-based splits so models cannot use future data to predict earlier events [5].
- Entities and preprocessing. The same customer on both sides inflates scores, and encoders fitted on the full dataset before splitting leak statistics; fit preprocessing on training folds only [15].
Write timestamps make these checks possible, so history belongs in the license (target leakage audit).
De-identifying tables without breaking joins
Removing names and account numbers does not de-identify a table, because combinations of ordinary columns single people out, and inconsistent key replacement destroys the joins that give relational data its value. Treat direct identifiers, quasi-identifiers (ZIP code, birth date, job title) and free-text columns as separate problems.
- Quasi-identifiers. Rocher and colleagues estimated with a generative model that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes; it is a model estimate, not a count of actual re-identifications [16]. Generalize dates and locations, then test residual risk, for example with k-anonymity.
- Keys. Use one keyed pseudonymization scheme across all tables, with the key held by the supplier, so joins survive and tokens cannot be recomputed from known identifiers.
- Free text. Memo and description columns carry identifiers that column rules miss.
- Alternatives. NIST SP 800-188 cautions about the limits of traditional de-identification compared with formal privacy methods, and describes synthetic data and protected enclaves as other sharing models [17].
As of October 2026, US rules attach to particular families:
- Health tables. HIPAA de-identification uses Safe Harbor, which removes 18 listed identifiers, or Expert Determination [18]. SourceX requires one of these methods for health records before anything is considered for a license.
- Consumer records. California Civil Code section 1798.140(m) treats information as deidentified only if the business takes reasonable measures, publicly commits not to re-identify it and contractually obligates recipients to comply [19], so expect no-re-identification terms.
- Financial transactions. Under Regulation P (12 CFR 1016.11), a recipient of nonpublic personal information from a nonaffiliated financial institution faces reuse and redisclosure limits even if it is not itself a financial institution [20].
For datasets sourced through SourceX, personal details such as names and account numbers are removed or replaced before delivery, the method is recorded per dataset, and a sample is checked after processing. No de-identification method is perfect. Techniques are compared in de-identifying tabular and transactional data and the de-identified data guide.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Public, synthetic or licensed: choosing a route per use case
Public benchmarks suit reproducing baselines, synthetic tables suit pipeline tests and privacy-constrained prototypes, and licensed real records are needed when a model must learn real distributions, dirty values and cross-table behavior.
| If you are building | Public data is enough to | License real records when | First check on a sample |
|---|---|---|---|
| A tabular foundation model | Reproduce published baselines | You need large, clean tables from domains absent in public corpora | Table-level overlap with public corpora |
| A text-to-SQL or analytics agent | Compare against Spider 2.0 or BIRD | Target schemas are wide, poorly documented or dialect-specific | Gold SQL executes on the delivered schema |
| A spreadsheet agent | Test forum-style tasks | You need business workbooks with formulas and cross-sheet links | Formulas preserved, not flattened to values |
| A fraud or anomaly model | Prototype features | Label rates and patterns must match your traffic | Label source and date on every row |
| A forecasting model | Benchmark on public series | You need series no public model has seen | Stockouts, promotions and gaps flagged |
| A synthetic data generator | Prototype on public seed tables | It must reproduce a business domain public tables lack, scored against a real holdout | Fidelity and copy checks on the holdout |
Synthetic tables are scored against real data: the SDMetrics library measures statistical differences, machine-learning efficacy and privacy [21]. See evaluating synthetic tabular data and combining licensed and synthetic data.
SourceX works the licensed route: it sources operational datasets from US companies, including sales histories and finance workflows, and manages the licensing agreement and later purchases. Datasets are sourced on request, not held in stock, so a request does not guarantee a match. You can send SourceX a structured data specification.
Delivery formats that keep types, keys and metadata intact
Deliver structured data as typed columnar files or a governed warehouse share with the data dictionary alongside, not as CSV dumps that lose types and keys.
- Columnar files. Apache Parquet writes file metadata in a footer that records where each column chunk starts, so readers fetch only the columns they need [22]. Open table formats add snapshots on top (Iceberg and Delta Lake delivery).
- Warehouse shares. In Snowflake Secure Data Sharing no data is copied between accounts and shared objects are read-only for the consumer [23] (zero-copy warehouse data sharing).
- Metadata. Croissant describes a dataset in JSON-LD built on schema.org, covering dataset metadata, file resources and record-level structure [24]. Transfer channels for every data type are in the dataset delivery guide.
Start here: structured data guides by family
Read the guide that matches your family and model goal.
- Tables and relational databases: real-world tables for tabular foundation models, multi-table data for relational deep learning, how much tabular data you need.
- SQL, schemas and spreadsheets: text-to-SQL training data, private text-to-SQL evaluation sets, spreadsheet benchmarks from real workbooks, messy multi-table spreadsheets.
- Transactions and ERP: ERP transaction and master data, fraud-labeled transactions, journal entries for anomaly detection.
- Time series, logs and events: time-series foundation model data, order histories for demand forecasting, production logs for anomaly detection, labeled equipment failure data for predictive maintenance.
- Master data and graphs: entity resolution data, product attribute data, enterprise knowledge graph data.
Mistakes that sink structured data purchases
Beyond leakage and broken keys, five errors routinely leave purchased structured data unusable or under-licensed.
- Joining systems without a crosswalk. CRM, billing and ERP rarely share a customer ID; ask for the mapping table and its match method.
- Ignoring currencies, units and time zones. Mixed currencies and local timestamps across daylight-saving changes corrupt amounts and sequences.
- Keeping test and demo records. Business systems hold test transactions and demo accounts; ask the supplier to flag or remove them.
- Evaluating on public benchmark data. It may already sit in a model's training mix; keep a private held-out set (LLM evaluation datasets).
- Leaving derived artifacts out of the license. Features, embeddings, fine-tuned weights and synthetic tables made from licensed data need explicit terms (AI training data licensing).
Describe the tables, series or transactions your model needs
At the SourceX buyer page you can submit a data request, talk to SourceX or browse the dataset catalog. Describe the data family, grain, history, labels and permitted uses you need; SourceX looks for US companies that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery. Nothing is contracted until a supplier agrees.
Guides in this section
- Data Dictionaries and Schema Docs for Text-to-SQL GroundingHow to source real data dictionaries, column descriptions, business glossaries and metric definitions to train and ground text-to-SQL and BI copilots.
- Demand Forecasting Training Data: Sales and Order HistoriesWhat to license for demand forecasting models: SKU-location sales and order histories with prices, promotions, stockout flags and product lifecycle dates.
- Entity Resolution Training Data from Real Master DataHow to source entity resolution training data: real duplicate-laden master records, gold match clusters, merge histories and entity-level splits.
- ERP Transaction and Master Data for AI TrainingHow to specify and license real ERP document flows and master data (orders, deliveries, invoices, vendors) for ERP copilots, agents and forecasting.
- Fraud Detection Training Data: Labels, Latency, Base RatesHow to specify and license real fraud-labeled card, ACH and e-commerce transactions: label maturity, base rates, account histories and PAN tokenization.
- Private Enterprise Text-to-SQL Evaluation Sets for BuyersHow to source held-out, execution-verified text-to-SQL and data-agent eval sets on real enterprise data when Spider and BIRD scores stop discriminating.
- Product Attribute Extraction Data from Verified PIM RecordsHow to source product attribute extraction data: verified PIM/ERP attribute values with raw titles, B2B spec fields, normalization targets and rights.
- Production Log Data for Anomaly Detection TrainingHow to source and license real, labeled production logs for log parsing, anomaly detection and LLM log analysis beyond HDFS, BGL and other public sets.
- Relational Database Datasets for Machine LearningHow to source multi-table relational data for relational deep learning and relational foundation models: schema, keys, timestamps and join-safe privacy.
- Spreadsheet Benchmarks for LLMs from Real WorkbooksHow to build or license a private spreadsheet benchmark for LLM agents: real workbooks, task instructions, multiple test cases and contamination controls.
- Tabular Foundation Model Training Data: What to LicenseWhich real-world tables to license for tabular foundation model pretraining, how they sit beside synthetic priors, and how to build held-out eval suites.
- Text-to-SQL Training Data from Real Enterprise SchemasHow to specify and license text-to-SQL training data on real enterprise schemas: gold SQL, dialects, query-log sourcing, value privacy and rights.
- Time-Series Foundation Model Training Data to LicenseWhich real business time series to license for TSFM pretraining: domains, frequencies, scale, dedup against benchmarks, and real vs synthetic data mix.
- Churn Prediction Training Data: Billing and Usage HistoriesHow to specify and license real subscription, invoice, usage and support histories with observed churn outcomes for churn and customer-health models.
- Enterprise Knowledge Graph Data for LLMs: A Buyer's GuideHow to source entity-relationship triples, ontologies and per-fact provenance from real business systems for KG-grounded LLMs and KG completion models.
- Event Sequence Data for ML: Schema, Splits and SourcingHow to structure, split and license entity-keyed business event streams (orders, payments, logins, status changes) for next-event and time-to-event models.
- Excel Formula Datasets: Context, Dependencies and RepairsHow to specify an Excel formula dataset from real workbooks: extraction schema, dependency graphs, error-to-fix repair pairs, dedup and de-identification.
- Held-Out Time Series for Forecasting Model EvaluationWhy public forecasting benchmarks can overlap pretraining corpora, and how to specify unseen business time series for credible zero-shot evaluation.
- Incident-Labeled Telemetry for Root-Cause Analysis ModelsHow to source logs, metrics and traces aligned to confirmed incident timelines and root-cause labels for training and evaluating AIOps RCA models.
- Journal Entry Data for Anomaly Detection and Audit AIHow to source real journal entries with posting metadata and reviewed exception labels for audit analytics, JE testing and close anomaly models.
- Messy Spreadsheet Data for Table Detection and NormalizationWhat messy multi-table spreadsheet data and labels you need to train table detection, header inference and spreadsheet-to-relational normalization models.
- Point-in-Time Correct Training Data from Licensed ExtractsHow to require snapshots, change logs and event vs system time in licensed tables so as-of joins use only what was known at prediction time.
- Product Matching Datasets: Offer Pairs, GTINs and VariantsWhat a product matching dataset needs: labeled offer pairs and clusters, GTIN and MPN coverage, variant hard negatives, and public benchmark gaps.
- Schema Linking Datasets for Wide Enterprise WarehousesWhat labeled data trains schema retrieval for text-to-SQL on warehouses with thousands of columns: labels, hard negatives, metrics and a buyer spec.
- Spend Classification Training Data: UNSPSC, GS1, CustomHow to source real PO, AP and catalog line items with verified UNSPSC, GS1 GPC or custom taxonomy codes for spend and product classification models.
- Synthetic Tabular Data Evaluation: Metrics and Seed TablesEvaluate synthetic tabular and multi-table data: SDMetrics scores, train-on-synthetic-test-on-real, copy checks, and the real seed tables behind them.
- Table Autocompletion Training Data for ERP and CRM FieldsThe historical records, time-respecting labels and correction history needed to train field autocompletion and attribute recommendation for business forms.
- Tabular Data for LLM Fine-Tuning: Table Reasoning TasksHow post-training teams turn licensed business tables into LLM fine-tuning tasks: serialization, table QA, prediction, verification and leakage controls.
- Tabular Data Request Spec: Grain, Keys, History, CodesHow to specify a tabular data request: table grain, primary and foreign keys, snapshot vs change history, units, nulls, code tables and dictionaries.
- Target Leakage in Tabular Data: A Pre-Purchase AuditHow buyers find columns that encode the outcome in a supplier's tabular sample: prediction-point mapping, timestamp lineage and single-feature AUC tests.
Sources
- Silicon Saxony, "SAP advancing enterprise AI research with first real ERP dataset". https://silicon-saxony.de/en/sap-advancing-enterprise-ai-research-with-first-real-erp-dataset/
- arXiv, "TabDPT: Scaling Tabular Foundation Models on Real Data" (2024). https://arxiv.org/pdf/2410.18164
- arXiv, "Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data" (2025). https://arxiv.org/pdf/2507.03971
- Google Research, "A decoder-only foundation model for time-series forecasting" (2024). https://research.google/blog/a-decoder-only-foundation-model-for-time-series-forecasting/
- arXiv, "RelBench: A Benchmark for Deep Learning on Relational Databases" (2024). https://arxiv.org/pdf/2407.20060
- arXiv, "RelBench v2: A Large-Scale Benchmark and Repository for Relational Data" (2026). https://arxiv.org/html/2602.12606v1
- ML Anthology (ICLR 2025), "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows" (2025). https://mlanthology.org/iclr/2025/lei2025iclr-spider
- arXiv, "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows (arXiv:2411.07763)" (2024). https://www.arxiv.org/pdf/2411.07763
- arXiv (Li et al.), "Can LLM Already Serve as a Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs (BIRD)" (2023). https://arxiv.org/abs/2305.03111v2
- arXiv, "SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation" (2024). https://arxiv.org/html/2406.14991v2
- arXiv, "Loghub: A Large Collection of System Log Datasets for AI-driven Log Analytics" (2020; revised 2023). https://arxiv.org/pdf/2008.06448
- ITMM journal (journals.nmetau.edu.ua), "Methodology of dataset preparation for training e-commerce fraud detection models" (2026). https://journals.nmetau.edu.ua/index.php/itmm/en/article/view/2468
- ISO/IEC, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- temporalcv documentation, "Why Time Series Is Different". https://temporalcv.readthedocs.io/en/latest/guide/why_time_series_is_different.html
- Machine Learning Mastery, "3 subtle ways data leakage can ruin your models (and how to prevent it)". https://machinelearningmastery.com/3-subtle-ways-data-leakage-can-ruin-your-models-and-how-to-prevent-it/
- Rocher, Hendrickx and de Montjoye (Nature Communications), "Estimating the success of re-identifications in incomplete datasets using generative models" (2019). https://pmc.ncbi.nlm.nih.gov/articles/PMC6650473
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- California Legislature, "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
- Synthetic Data Vault documentation, "SDMetrics". https://docs.sdv.dev/sdmetrics
- The Apache Software Foundation, "Apache Parquet: File Format". https://parquet.apache.org/docs/file-format/
- Snowflake Documentation, "About Secure Data Sharing". https://docs.snowflake.com/en/user-guide/data-sharing-intro.html
- Akhtar et al. (MLCommons), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.