Skip to content

Tables, time series and transactional data

Tabular, Time-Series and Transactional Data for AI: A Buyer's Guide

Quick answer

Tabular data for AI spans several families of structured business records: single tables, multi-table relational databases, schemas with SQL, spreadsheets, transactions and ledgers, time series, logs and event streams, and master data or knowledge graphs. Each family trains different models, from tabular foundation models to text-to-SQL agents and forecasters. Real enterprise versions rarely reach the open market, so buyers license them and should fix grain, keys, history, label timing and join-safe de-identification before discussing volume.

By SourceX Editorial · Updated

This hub covers the structured data cluster of the SourceX guide to AI data. For definitions, read structured vs unstructured AI training data.

Eight families of structured data and the models they train

Scope structured data for AI training by family, because each family comes from different systems, trains different models and fails acceptance for different reasons. Settle the last column before discussing volume.

FamilyTypical source systemsModels and tasksPin down first
Single wide tablesWarehouse marts, feature storesGradient-boosted trees, tabular foundation models (TFMs)Row meaning, target definition, units
Multi-table relational dataERP, CRM, billing and order databasesRelational deep learning, field autocompletionPrimary and foreign keys, timestamps on every table
Schemas, SQL and metric definitionsWarehouse catalogs, query logs, dbt modelsText-to-SQL, schema linking, analytics agentsColumn descriptions, SQL dialect, executable gold queries
Spreadsheets and workbooksFinance models, trackers, exported reportsSpreadsheet agents, formula generation, table detectionNative .xlsx with formulas and links intact
Transactions and ledgersPayment processors, general ledgers, accounts payableFraud and anomaly detection, reconciliationLabel source and date, dispute history, merchant codes
Time series and telemetryOrder histories, sensor historians, observability metricsForecasting, time-series foundation models, predictive maintenanceFrequency, gaps, time zones, price and promotion covariates
Logs and event streamsApplication logs, clickstreams, business event tablesLog anomaly detection, root-cause analysis, recommendersIncident labels aligned to timestamps, entity IDs
Master data and knowledge graphsCustomer, supplier and product mastersEntity resolution, product matching, LLM groundingMatch labels, merge history, licenses of external graph data

Several record types have their own SourceX pages: spreadsheets and financial models, financial transaction data, sensor and IoT data, sales CRM histories and accounting reconciliations. Licensing operational data for AI covers the wider category.

Why real enterprise tables are hard to find

Real multi-table business data rarely reaches the open market because it sits inside ERP, billing and payment systems under privacy, confidentiality and commercial constraints.

  • Scarcity. When SAP released SALT, anonymized sales tables from one customer's ERP system, it said text is plentiful but data with multiple linked tables is scarce, because privacy, confidentiality and commercial interests make company datasets hard to procure [1].
  • Tabular foundation models. TabDPT's authors report that existing TFMs are predominantly trained on synthetic data, and that adding real data to pretraining can give significantly faster training and better generalization [2]. Real-TabPFN reports that continued pretraining on a small, curated set of large real datasets beat broader, potentially noisier corpora such as CommonCrawl or GitTables [3].
  • Time series. Google Research describes TimesFM as pretrained on 100 billion real-world time points [4], so fair evaluation needs held-out business time series outside public corpora.
Public resourceWhat it containsGap a licensed dataset fills
RelBench and v2 (relational)Prediction tasks over multi-table databases with time-based splits [5]; v2 adds ERP, consumer-platform, clinical and scholarly databases [6]Few industries, mostly public databases
Spider 2.0 (text-to-SQL)632 tasks on real-application databases, often over 1,000 columns, on systems such as BigQuery and Snowflake [7]An o1-preview-based agent solved 17.0% in the arXiv version [8]; your schemas are absent
BIRD (text-to-SQL)12,751 question-SQL pairs over 95 databases in 37 domains, with dirty values; in the 2023 paper ChatGPT reached 40.08% execution accuracy against 92.96% for humans [9]No private business logic
SpreadsheetBench912 forum-sourced instructions; 35.7% of files hold multiple tables and 42.7% non-standard relational tables [10]Forum problems, not business workbooks
Loghub (logs)19 real-world log datasets, of which 6 are labeled [11]Few labels; system, not business, logs
Open fraud setsOne study merged IEEE-CIS, Sparkov and a Kaggle e-commerce set and added synthetic attributes after finding open sources insufficient [12]Partly synthetic; may not match your payment mix

Specify grain, keys and history before row counts

A structured data request should define what one row means, how tables join, how history is kept and when each label became known; row counts mean little until those are fixed. The tabular license specification guide adds clause-level detail.

  • Grain. "One row per invoice line" and "one row per invoice" train different models; mixed grains double-count amounts.
  • Keys. Name primary and foreign keys and agree a maximum orphan-row rate. Subsets must keep whole entity graphs (sampling relational extracts without breaking keys).
  • History. Ask for change history or dated snapshots, with both business event time and the time each row was written; current-state tables cannot support point-in-time correct training data.
  • Labels. Record each label's source and when it became known: a dispute flag set after settlement did not exist when the payment was scored.
  • Metadata. Spider 2.0 tasks require database metadata and dialect documentation [7]; ask for column descriptions, code lists, units and metric definitions (data dictionaries as grounding data).
  • Coverage. State entities and time span: ten million rows from 200 customers is a narrow sample.

Acceptance thresholds can borrow the data quality model and measures of ISO/IEC 5259-2:2024 [13]; validation checks for structured deliveries turns them into tests. The manifest below applies these points to a four-table extract.

Illustrative example: invented to show structure; it does not describe an available dataset.

dataset: order_to_cash_extract_example
history: change_log                 # every update kept with valid_from / valid_to
time_zone: UTC
coverage: {from: 2021-01-01, to: 2025-12-31}
tables:
  customers:
    grain: one row per customer version
    primary_key: [customer_token, valid_from]
    customer_token: {method: keyed_hash, key_held_by: supplier, same_key_all_tables: true}
    postal_code: {method: generalize, keep_digits: 3}
  orders:
    grain: one row per order
    primary_key: [order_id]
    foreign_keys: {customer_token: customers.customer_token}   # as-of join on valid_from / valid_to
    event_time: order_created_at
    written_at: row_inserted_at
  order_lines:
    grain: one row per order line
    primary_key: [order_id, line_no]
    foreign_keys: {order_id: orders.order_id}
  payments:
    grain: one row per payment attempt
    primary_key: [payment_attempt_id]
    foreign_keys: {order_id: orders.order_id}
    label: {column: disputed, source: processor_dispute_file, known_at: dispute_opened_at}
    free_text_columns: {payment_memo: pii_scanned_and_masked}
acceptance:
  duplicate_primary_keys: 0
  orphan_foreign_keys_max_pct: 0.1
license:
  license_id: LIC-EXAMPLE-01
  permitted_uses: [model_training, internal_evaluation]
  derived_artifacts: [features, embeddings]   # name them explicitly
  • known_at lets you drop any label or field that did not exist at prediction time.
  • One keyed hash applied across all tables pseudonymizes customers without breaking joins.
  • event_time and written_at separate when something happened from when the system recorded it.

Leakage: the defect that makes a weak table look strong

Leakage means a training table holds information that would not exist at prediction time. In licensed business data it usually enters through post-outcome fields, random splits over time-ordered rows, or the same entity on both sides of a split.

  • Post-outcome fields. A write_off_reason in a credit table or resolved_by in a ticket table is filled only after the outcome. Ask which columns are written after the event you predict.
  • Time. Split chronologically; for time-series validation, one library recommends a gap between training and test equal to the forecast horizon [14]. RelBench uses time-based splits so models cannot use future data to predict earlier events [5].
  • Entities and preprocessing. The same customer on both sides inflates scores, and encoders fitted on the full dataset before splitting leak statistics; fit preprocessing on training folds only [15].

Write timestamps make these checks possible, so history belongs in the license (target leakage audit).

De-identifying tables without breaking joins

Removing names and account numbers does not de-identify a table, because combinations of ordinary columns single people out, and inconsistent key replacement destroys the joins that give relational data its value. Treat direct identifiers, quasi-identifiers (ZIP code, birth date, job title) and free-text columns as separate problems.

  • Quasi-identifiers. Rocher and colleagues estimated with a generative model that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes; it is a model estimate, not a count of actual re-identifications [16]. Generalize dates and locations, then test residual risk, for example with k-anonymity.
  • Keys. Use one keyed pseudonymization scheme across all tables, with the key held by the supplier, so joins survive and tokens cannot be recomputed from known identifiers.
  • Free text. Memo and description columns carry identifiers that column rules miss.
  • Alternatives. NIST SP 800-188 cautions about the limits of traditional de-identification compared with formal privacy methods, and describes synthetic data and protected enclaves as other sharing models [17].

As of October 2026, US rules attach to particular families:

  • Health tables. HIPAA de-identification uses Safe Harbor, which removes 18 listed identifiers, or Expert Determination [18]. SourceX requires one of these methods for health records before anything is considered for a license.
  • Consumer records. California Civil Code section 1798.140(m) treats information as deidentified only if the business takes reasonable measures, publicly commits not to re-identify it and contractually obligates recipients to comply [19], so expect no-re-identification terms.
  • Financial transactions. Under Regulation P (12 CFR 1016.11), a recipient of nonpublic personal information from a nonaffiliated financial institution faces reuse and redisclosure limits even if it is not itself a financial institution [20].

For datasets sourced through SourceX, personal details such as names and account numbers are removed or replaced before delivery, the method is recorded per dataset, and a sample is checked after processing. No de-identification method is perfect. Techniques are compared in de-identifying tabular and transactional data and the de-identified data guide.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Public, synthetic or licensed: choosing a route per use case

Public benchmarks suit reproducing baselines, synthetic tables suit pipeline tests and privacy-constrained prototypes, and licensed real records are needed when a model must learn real distributions, dirty values and cross-table behavior.

If you are buildingPublic data is enough toLicense real records whenFirst check on a sample
A tabular foundation modelReproduce published baselinesYou need large, clean tables from domains absent in public corporaTable-level overlap with public corpora
A text-to-SQL or analytics agentCompare against Spider 2.0 or BIRDTarget schemas are wide, poorly documented or dialect-specificGold SQL executes on the delivered schema
A spreadsheet agentTest forum-style tasksYou need business workbooks with formulas and cross-sheet linksFormulas preserved, not flattened to values
A fraud or anomaly modelPrototype featuresLabel rates and patterns must match your trafficLabel source and date on every row
A forecasting modelBenchmark on public seriesYou need series no public model has seenStockouts, promotions and gaps flagged
A synthetic data generatorPrototype on public seed tablesIt must reproduce a business domain public tables lack, scored against a real holdoutFidelity and copy checks on the holdout

Synthetic tables are scored against real data: the SDMetrics library measures statistical differences, machine-learning efficacy and privacy [21]. See evaluating synthetic tabular data and combining licensed and synthetic data.

SourceX works the licensed route: it sources operational datasets from US companies, including sales histories and finance workflows, and manages the licensing agreement and later purchases. Datasets are sourced on request, not held in stock, so a request does not guarantee a match. You can send SourceX a structured data specification.

Delivery formats that keep types, keys and metadata intact

Deliver structured data as typed columnar files or a governed warehouse share with the data dictionary alongside, not as CSV dumps that lose types and keys.

  • Columnar files. Apache Parquet writes file metadata in a footer that records where each column chunk starts, so readers fetch only the columns they need [22]. Open table formats add snapshots on top (Iceberg and Delta Lake delivery).
  • Warehouse shares. In Snowflake Secure Data Sharing no data is copied between accounts and shared objects are read-only for the consumer [23] (zero-copy warehouse data sharing).
  • Metadata. Croissant describes a dataset in JSON-LD built on schema.org, covering dataset metadata, file resources and record-level structure [24]. Transfer channels for every data type are in the dataset delivery guide.

Start here: structured data guides by family

Read the guide that matches your family and model goal.

Mistakes that sink structured data purchases

Beyond leakage and broken keys, five errors routinely leave purchased structured data unusable or under-licensed.

  • Joining systems without a crosswalk. CRM, billing and ERP rarely share a customer ID; ask for the mapping table and its match method.
  • Ignoring currencies, units and time zones. Mixed currencies and local timestamps across daylight-saving changes corrupt amounts and sequences.
  • Keeping test and demo records. Business systems hold test transactions and demo accounts; ask the supplier to flag or remove them.
  • Evaluating on public benchmark data. It may already sit in a model's training mix; keep a private held-out set (LLM evaluation datasets).
  • Leaving derived artifacts out of the license. Features, embeddings, fine-tuned weights and synthetic tables made from licensed data need explicit terms (AI training data licensing).

Describe the tables, series or transactions your model needs

At the SourceX buyer page you can submit a data request, talk to SourceX or browse the dataset catalog. Describe the data family, grain, history, labels and permitted uses you need; SourceX looks for US companies that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery. Nothing is contracted until a supplier agrees.

Describe your structured data needs to SourceX

Guides in this section

Sources

  1. Silicon Saxony, "SAP advancing enterprise AI research with first real ERP dataset". https://silicon-saxony.de/en/sap-advancing-enterprise-ai-research-with-first-real-erp-dataset/
  2. arXiv, "TabDPT: Scaling Tabular Foundation Models on Real Data" (2024). https://arxiv.org/pdf/2410.18164
  3. arXiv, "Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data" (2025). https://arxiv.org/pdf/2507.03971
  4. Google Research, "A decoder-only foundation model for time-series forecasting" (2024). https://research.google/blog/a-decoder-only-foundation-model-for-time-series-forecasting/
  5. arXiv, "RelBench: A Benchmark for Deep Learning on Relational Databases" (2024). https://arxiv.org/pdf/2407.20060
  6. arXiv, "RelBench v2: A Large-Scale Benchmark and Repository for Relational Data" (2026). https://arxiv.org/html/2602.12606v1
  7. ML Anthology (ICLR 2025), "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows" (2025). https://mlanthology.org/iclr/2025/lei2025iclr-spider
  8. arXiv, "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows (arXiv:2411.07763)" (2024). https://www.arxiv.org/pdf/2411.07763
  9. arXiv (Li et al.), "Can LLM Already Serve as a Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs (BIRD)" (2023). https://arxiv.org/abs/2305.03111v2
  10. arXiv, "SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation" (2024). https://arxiv.org/html/2406.14991v2
  11. arXiv, "Loghub: A Large Collection of System Log Datasets for AI-driven Log Analytics" (2020; revised 2023). https://arxiv.org/pdf/2008.06448
  12. ITMM journal (journals.nmetau.edu.ua), "Methodology of dataset preparation for training e-commerce fraud detection models" (2026). https://journals.nmetau.edu.ua/index.php/itmm/en/article/view/2468
  13. ISO/IEC, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  14. temporalcv documentation, "Why Time Series Is Different". https://temporalcv.readthedocs.io/en/latest/guide/why_time_series_is_different.html
  15. Machine Learning Mastery, "3 subtle ways data leakage can ruin your models (and how to prevent it)". https://machinelearningmastery.com/3-subtle-ways-data-leakage-can-ruin-your-models-and-how-to-prevent-it/
  16. Rocher, Hendrickx and de Montjoye (Nature Communications), "Estimating the success of re-identifications in incomplete datasets using generative models" (2019). https://pmc.ncbi.nlm.nih.gov/articles/PMC6650473
  17. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  18. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  19. California Legislature, "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  20. Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
  21. Synthetic Data Vault documentation, "SDMetrics". https://docs.sdv.dev/sdmetrics
  22. The Apache Software Foundation, "Apache Parquet: File Format". https://parquet.apache.org/docs/file-format/
  23. Snowflake Documentation, "About Secure Data Sharing". https://docs.snowflake.com/en/user-guide/data-sharing-intro.html
  24. Akhtar et al. (MLCommons), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data