Schemas, packaging and delivery
Packaging Linked Records from Multiple Business Systems
Quick answer
To deliver relational data for machine learning from several business systems, ship normalized tables with declared primary and foreign keys, an entity-relationship diagram with cardinalities, and a crosswalk table wherever two systems use different IDs for the same customer or case. Add per-case JSON bundles only as a derived view. Every table must share one pseudonymization scheme so joins still work after de-identification, and the delivery should pass referential-integrity checks before it leaves the source.
By SourceX Editorial · Updated
This page covers the packaging pattern itself. For the data category, see enterprise workflow datasets and agent trajectories; for broader delivery mechanics, start at the dataset delivery hub; for how tables relate to free text, see structured vs unstructured AI training data.
Why linked records lose value when flattened
Flattening several systems into one wide file throws away the structure that agent training and relational models need. A common practice is to materialize a single joined training table and export it, which discards relational information that a flat file cannot carry [4]. Relational deep learning, by contrast, treats tables as a graph of entities connected by keys, so the keys are the signal rather than overhead [3].
The research community now benchmarks directly on multi-table databases. RelBench defines predictive tasks over relational databases and trains graph neural networks across linked tables [1], and a larger RelBench v2 (comprising over 22 million rows across 29 tables) has followed [2]. SAP's SALT release, described as anonymized linked business tables from a customer's ERP system, shows the same shape coming from a real enterprise source [5]. If you plan to use these methods, see multi-table relational data for relational foundation models.
For agent work the problem is sharper. A support agent that must decide on a refund needs the ticket thread, the CRM account tier, the open invoice and the order line in the ERP, linked in time order. A flattened row per ticket typically drops the one-to-many invoice history or duplicates it across rows, and both failures teach the model wrong facts.
Normalized tables or per-case bundles: choosing the delivery shape
Ask for normalized tables as the system of record and treat denormalized bundles as a convenience layer built from them. Normalized tables preserve cardinality, avoid duplication and let you rebuild any view. Bundles are easier to feed into SFT, evaluation harnesses and RL environments built from business workflows, but each bundle freezes one choice of join and time window.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Criterion | Normalized tables (Parquet + ERD) | Per-case JSON bundles (JSONL) |
|---|---|---|
| Best for | Relational models, feature work, re-slicing | SFT examples, eval cases, agent episodes |
| Cardinality | Explicit in schema and ERD | Implicit in nesting; easy to lose |
| Duplication | None by design | Shared entities repeat in every case |
| Leakage control | Filter by timestamp at join time | Must be cut correctly when built |
| Schema change | Add columns per table | Rebuild every bundle |
| Integrity check | Foreign-key orphan counts | Per-bundle completeness only |
A practical default is both: Parquet tables as the canonical layer, plus a JSONL bundle file generated by a published script, with each bundle carrying the primary keys of every record it contains. Parquet writes file metadata, including the locations of all column chunks, in the file footer, so readers can load only the columns they need [7]. For format trade-offs, see Parquet vs JSONL for licensed training data.
Crosswalk tables for IDs that differ between systems
A crosswalk table is the most reliable way to join records when the CRM, help desk and billing system each assign their own identifier to the same customer. Salesforce account IDs, Zendesk organization IDs and NetSuite customer internal IDs rarely match, and email-domain or name matching produces silent false joins. The crosswalk should be built at the source, where the supplier can see the original keys, and delivered as its own table.
Specify these columns for every crosswalk row: a canonical entity key, the source system name, the source-system key (pseudonymized), the match method (exact_key, integration_sync, deterministic_rule, probabilistic, manual), a match confidence where applicable, and valid-from and valid-to dates. The dates matter because accounts merge, split and get reassigned; a crosswalk without validity windows will attach a 2023 ticket to a 2025 account owner. For how to audit the result, see verifying record linkage across systems.
Request a separate link table for case-level relationships, not just customers. A ticket that references an invoice number, an opportunity that converts to a sales order, or an incident that spawns a change request each need an explicit edge with both keys and the event time. The cross-system workflow records page shows how these edges form one task timeline.
Pseudonymous keys that stay consistent across every table
Pseudonymization must be applied once, with one keyed function, across all tables and the crosswalk, or the joins break. If the help desk export hashes customer emails with one salt and the billing export uses another, the same person becomes two entities. A keyed HMAC over the original identifier, with the key held only by the supplier, gives stable tokens across tables and across later deliveries; the stable record IDs and join keys page covers versioning those keys.
Consistent keys also raise linkage risk, because a single token now connects ticket text, payment history and CRM notes. California defines deidentified information as information that cannot reasonably be used to infer information about, or otherwise be linked to, a particular consumer, with specific conditions on how the holder treats it (Civil Code section 1798.140(m)) [8], so the linked package, not each table alone, is what should be assessed. Free-text fields need their own pass: automated tools such as Presidio detect and replace PII but state that they cannot guarantee finding all of it [9]. Compare the standards in deidentified data under US state privacy laws.
Entity-relationship diagram and cardinality documentation
An entity-relationship diagram turns a folder of files into a dataset another engineer can join correctly on the first try. It should name every table, its primary key, each foreign key and the cardinality of every relationship (one-to-one, one-to-many, many-to-many through a link table), plus which side may be null. Ship it as both an image and a machine-readable schema.
Machine-readable metadata keeps the diagram and files in sync. Croissant, a schema.org-based JSON-LD vocabulary from the MLCommons working group, describes a dataset's file resources and record structure for ML tools [6]; pair it with a column dictionary that records type, units, time zone, null semantics and source field name. A completed technical delivery specification is the right place to fix these requirements before extraction.
Illustrative example: invented to show structure; it does not describe an available dataset.
tables:
account: { pk: account_key, source: crm }
ticket: { pk: ticket_key, source: helpdesk,
fk: { account_key: account.account_key }, null_fk: allowed }
ticket_event: { pk: event_id, fk: { ticket_key: ticket.ticket_key } }
invoice: { pk: invoice_key, source: billing,
fk: { account_key: account.account_key } }
ticket_invoice: { pk: [ticket_key, invoice_key], link_basis: "invoice number in ticket field" }
id_crosswalk: { pk: [source_system, source_key],
cols: [account_key, match_method, valid_from, valid_to] }
relationships:
- account 1..* ticket
- ticket 1..* ticket_event
- account 1..* invoice
- ticket *..* invoice via ticket_invoice
time: { zone: UTC, event_time_col: occurred_at }
Referential-integrity checks to run before and after delivery
Every multi-system delivery should ship with an integrity report, and you should rerun it on intake. Exports from different systems are usually pulled at different moments with different filters, so orphans are the norm rather than the exception. The report turns that into a known, documented number instead of a surprise during training.
Illustrative example: invented to show structure; it does not describe an available dataset.
Integrity checklist for a linked delivery:
- Primary-key uniqueness: zero duplicate keys per table, including after pseudonymization (hash collisions or reused source IDs).
- Orphan rate per foreign key: count child rows whose parent key is missing, by relationship, with the reason (out-of-window, deleted, filtered).
- Cardinality conformance: relationships declared one-to-one have no fan-out; observed maximum children per parent are reported.
- Crosswalk coverage: share of entities in each system that resolve to a canonical key, broken out by match method.
- Temporal sanity: no child event before its parent was created; no ticket linked to an invoice dated after the ticket was closed unless documented.
- Extraction alignment: each table's snapshot timestamp and filter, so you can see why counts differ.
- Bundle reproducibility: rebuilding JSONL bundles from tables with the published script yields identical case counts and checksums.
Pair the report with a file manifest and hashes, covered in dataset manifests and checksums.
Time windows, leakage and recurring deliveries
Linked records carry future information unless you cut every table at the same reference time. For evaluation and outcome-labeled tasks, a ticket bundle must exclude refunds, escalations or account changes that happened after the decision point, or the model learns the answer from the context. Request an occurred_at and, where systems record it, an updated_at on every row, so you can reconstruct the state as of any moment.
Ongoing purchases add drift: fields get renamed in the CRM, a billing migration changes ID formats, a help desk adds custom fields. Keep the canonical keys and crosswalk stable across releases and track changes per table, as described in handling schema changes across recurring deliveries.
How SourceX approaches multi-system datasets
SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, and finance and legal workflows, and manages the licensing process; nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and the supplying company approves every release. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Buyers can describe the linked records and keys they need on the SourceX buyer page.
Request linked multi-system records for agent training
If your project needs ticketing, CRM, billing or ERP records linked by consistent keys, describe the systems, relationships and time windows rather than specific companies. SourceX looks for US businesses that hold the described data and works through assessment and licensing before anything is delivered through a private, access-controlled workflow. Describe the linked records you need.
Sources
- arXiv, "RelBench: A Benchmark for Deep Learning on Relational Databases" (2024). https://arxiv.org/pdf/2407.20060
- arXiv, "RelBench v2: A Large-Scale Benchmark and Repository for Relational Data" (2026). https://arxiv.org/abs/2602.12606
- Kumo AI, "Relational Deep Learning". https://kumo.ai/research/relational-deep-learning-rdl/
- Academia.edu, "Machine Learning over Static and Dynamic Relational Data". https://www.academia.edu/128861538/Machine_Learning_over_Static_and_Dynamic_Relational_Data
- SAP News Center, "SAP SALT: Real ERP Dataset for Enterprise AI Research" (2025). https://news.sap.com/2025/04/sap-salt-real-erp-dataset-enterprise-ai-research/
- arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- The Apache Software Foundation, "File Format" (Apache Parquet). https://parquet.apache.org/docs/file-format/
- California Legislature, "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- Microsoft (presidio project), "Presidio - Data Protection API". https://data-privacy-stack.github.io/presidio/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.