Tables, time series and transactional data
Multi-Table Relational Data for Relational Deep Learning and Relational Foundation Models
Quick answer
Relational deep learning and relational foundation models need the database itself, not a flattened feature table: entity tables, event tables, primary and foreign keys, and created/updated timestamps on every row that changes over time [1]. When you source relational database datasets for machine learning, specify the schema graph, history depth, foreign-key coverage after de-identification and temporal-split rules up front. Public benchmarks such as RelBench show the shape, but real linked business tables remain scarce [2][4].
By SourceX Editorial · Updated
What relational models consume that tabular models do not
Relational models consume the schema graph, so the unit you license is a set of linked tables rather than one wide table. RelBench frames a database as a heterogeneous temporal graph: rows become nodes, foreign keys become edges, and a GNN learns over that graph instead of a data scientist hand-writing aggregates such as "orders in the last 90 days" [1]. That changes procurement in three ways.
- Keys are the signal. A supplier that drops
customer_idfrom the order lines, or re-keys tables independently, destroys the edges the model trains on. - Event tables need time. Each event row needs a timestamp so the training graph at time t contains only rows that existed at t [1].
- Breadth matters for pretraining. A relational foundation model learns transferable patterns across many schemas, so diversity of databases matters as much as rows in any one of them [2].
If your team is still deciding between flat and linked data, the owner page on structured vs unstructured AI training data and the tabular foundation model pretraining data guide cover single-table diversity. This page is about multi-table extracts.
Where public relational benchmarks stop
Public relational benchmarks are good for method comparison but thin for enterprise pretraining. RelBench v2, which appears in the ICLR 2026 program, adds ERP, consumer-platform, scholarly and clinical databases, introduces autocomplete tasks that predict a missing cell from linked context, and points to the ReDeLEx collection of 70+ real-world databases for pretraining [2][3].
That is a small number of schemas relative to the variety inside real companies. SAP's SALT (Sales Autocompletion Linked Business Tables) release, built from anonymized ERP data, was presented as a first of its kind precisely because large, cleansed, multi-table company data is hard to procure for privacy, confidentiality and commercial reasons [4]. Expect the same friction when you license: the supplier is releasing its operational structure, not only its records.
Typical gaps buyers try to fill with licensed data:
- Schema depth. Real ERP or CRM extracts have dozens of tables, code tables, soft deletes and status histories that public benchmarks simplify away. See ERP transaction and master data for AI training.
- Held-out schemas. A pretrained relational model must be evaluated on databases it never saw, so a private, unpublished database is valuable as a test set.
- Long history. Churn, lifetime value and next-purchase tasks need multiple years of events per entity, which sampled public extracts rarely keep.
How to specify a relational extract before you license it
A usable specification names the tables, keys, grain, time columns and coverage numbers you will accept, so the supplier can say yes or no before anyone exports data. The general rules for grain and history are in what to specify when licensing tabular data; the relational additions are below.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field in your request | What to ask for | Why it matters for relational models |
|---|---|---|
| Schema diagram | Every table, PK, FK and cardinality (1:N, N:M via bridge table) | Defines the graph edges; missing bridge tables break N:M links |
| Row counts per table | Counts at extract time and per year | Shows graph size and event density per entity |
| Entity counts | Distinct customers, accounts, products, suppliers | Bounds the number of prediction targets |
| Timestamps | created_at and updated_at on every event table, with time zone | Required for temporal neighbor sampling and splits [1] |
| History depth | Earliest and latest event date; whether updates overwrite rows or append history | Overwritten rows leak the future into the past |
| Code tables | Lookups for status, reason and product category codes | Without them, categorical columns are opaque integers |
| FK coverage after de-identification | Share of FK values that resolve to a parent row, per FK | Detects joins broken by pseudonymization |
| Deletions and soft deletes | How deleted parents and is_deleted flags are represented | Orphans can be real business behavior or an extract bug |
| Format | Parquet per table plus a schema file (DDL or JSON) [8] | Preserves types and makes per-table loading cheap |
Ask for the DDL, not only a diagram image. Column types, nullability and declared constraints tell you which foreign keys the source system enforced and which were application-level conventions that may already contain orphans.
Temporal splits and leakage in multi-table data
Split relational data by time, not by random row, because a future row in any linked table can reveal the label. RelBench uses temporal splits for this reason: training uses events before a cutoff, validation and test use later windows, and the model may only see rows with timestamps earlier than the prediction time [1]. Random splits on linked tables leak through neighbors; a customer's future order lines, refunds or support tickets sit one hop away from the training node.
Leakage sources to check in a delivered sample:
- Overwritten status columns. An
orders.status = 'returned'value written months later, with no history table, tells the model the outcome. - Aggregates materialized by the source system. Columns such as
lifetime_valueorlast_order_dateon the customer table were computed at extract time. - Backfilled timestamps. A migration that set every
created_atto the load date collapses history into one instant. - Snapshot tables without validity ranges. Slowly changing dimensions need
valid_fromandvalid_to, or you cannot reconstruct the database as it looked at time t.
The same discipline applies to process data; the timestamped business event sequences guide covers single-sequence leakage, and object-centric logs in OCEL 2.0 already ship a relational SQLite format that maps events to multiple objects [9].
Keeping referential integrity after de-identification
De-identification must preserve joins, which means the same identifier gets the same replacement in every table and every delivery. If customer_id 48213 becomes c_9f2a in orders but c_71bd in tickets, the graph silently splits one customer into two disconnected nodes and the training signal disappears without an error.
Practical requirements to put in the request:
- Keyed, deterministic pseudonyms. Use one keyed hash or token vault per key domain, applied across all tables, and keep the key with the supplier so later refreshes map to the same values. Stable cross-delivery IDs are covered in sampling relational extracts without breaking keys.
- An FK coverage report. For each foreign key, report the orphan rate (child rows whose FK has no matching parent) before and after de-identification. A jump after processing signals a tooling error, not a business pattern.
- Free-text and quasi-identifiers. Names and emails inside
notesordescriptioncolumns, and rare combinations of ZIP code, dates and job title, survive key pseudonymization. Event-level business logs carry measurable uniqueness and re-identification risk even without names [5]. - Legal standard by data type. Consumer records may need to meet a deidentified standard such as the CCPA definition, which also attaches conditions to the business holding the data [6]. Health records need HIPAA Expert Determination or Safe Harbor [7]; Safe Harbor's date generalization can damage event ordering, so relational health projects often look at Expert Determination. State-by-state differences are compared in deidentified data under US state privacy laws.
Illustrative example: invented to show structure; it does not describe an available dataset.
FK coverage report (delivery sample)
fk child_rows orphans_before orphans_after
order_lines.order_id -> orders.id 1,204,330 0 0
orders.customer_id -> customers.id 310,912 41 41
tickets.customer_id -> customers.id 88,407 212 3,960 <- investigate
invoices.order_id -> orders.id 298,115 0 0
In that example, the jump on tickets.customer_id suggests the ticketing extract was pseudonymized with a different key than the CRM, which is a common relational de-identification failure.
License terms specific to relational models
The license should name the derivative artifacts a relational model produces, because embeddings and pretrained weights encode the supplier's schema and behavior. Ask counsel to address node and table embeddings, pretrained relational foundation model checkpoints, and synthetic databases generated from the licensed data, alongside the usual allowed-use, term and deletion language. If the database will also serve as a held-out evaluation set, state whether it may be shared with benchmark collaborators or must stay internal.
Schema documentation is itself valuable data; see data dictionaries and schema documentation as AI grounding data and, for multi-system joins across ERP, CRM and ticketing, packaging linked records from multiple business systems.
How SourceX handles relational extracts
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing and ongoing purchases; nothing is held in stock and a request does not guarantee a match. Relevant categories include support and sales histories, engineering records and finance workflows; owner pages cover sales CRM pipeline histories and supply chain and logistics datasets. Buyers describe the data they need, and SourceX looks for US businesses that hold it, with every release approved by the supplying company.
Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe your schema and history requirements to SourceX using the specification table above. For the wider category map, start at the tabular, time-series and transactional data hub or the AI data guide index.
Request multi-table relational data for your models
SourceX sources operational datasets from US companies, works through Find, Assess, Agree, Transact and Manage, and delivers only after an executed agreement and supplier approval through private, access-controlled workflows. Describe the tables, keys and history you need at sourcex.si/buyers.
Frequently asked questions
Can I train a relational model on a flattened feature export?
You can train a tabular model on it, but not a relational deep learning model in the RelBench sense, which learns from the raw tables and their keys instead of engineered aggregates [1]. A flattened export also hides how features were computed, which makes leakage harder to audit.
How many databases do I need for relational foundation model pretraining?
There is no settled number. As of October 2026, public work points to collections of 70+ real databases for pretraining [2], so a single licensed database is usually more valuable as a held-out evaluation schema or as domain-specific fine-tuning data than as the bulk of a pretraining mix.
What file format should a multi-table delivery use?
One Parquet file or partitioned directory per table, plus machine-readable DDL and a data dictionary, is a practical default because Parquet keeps column types and loads per table [8]. Event-centric process data can arrive as an OCEL 2.0 SQLite log [9].
Sources
- arXiv (Robinson et al.), "RelBench: A Benchmark for Deep Learning on Relational Databases" (2024). https://arxiv.org/pdf/2407.20060
- arXiv, "RelBench v2: A Large-Scale Benchmark and Repository for Relational Data" (2026). https://arxiv.org/html/2602.12606v1
- ICLR, "RelBench v2 (ICLR 2026 listing)" (2026). https://iclr.cc/virtual/2026/10019682
- Silicon Saxony, "SAP advancing enterprise AI research with first real ERP dataset". https://silicon-saxony.de/en/sap-advancing-enterprise-ai-research-with-first-real-erp-dataset/
- arXiv (Nunez von Voigt et al.), "Quantifying the Re-identification Risk of Event Logs for Process Mining" (2020). https://arxiv.org/pdf/2003.10707
- California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- The Apache Software Foundation, "Apache Parquet Documentation". https://parquet.apache.org/docs
- OCEL standard authors (arXiv), "OCEL (Object-Centric Event Log) 2.0 Specification" (2023). https://arxiv.org/pdf/2403.01975
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.