Skip to content

Tables, time series and transactional data

Training Data for Business-Table Autocompletion and Field Recommendation

Quick answer

To train a model that suggests values for ERP, CRM or finance form fields, you need historical business documents with their final field values, the linked master-data tables as they stood when each document was entered, and the record of what users changed. Labels must be time-respecting, so features use only information available at entry time, and fields filled later in the workflow must be excluded. Public benchmarks such as SAP's SALT and the RelBench v2 autocomplete tasks show the task framing; production accuracy depends on licensed histories from systems like yours.

By SourceX Editorial · Updated

What the autocompletion task actually predicts

Field autocompletion predicts one target field of a business document from the other fields already entered and from linked tables. SAP's SALT dataset, short for Sales Autocompletion Linked Business Tables, frames it on ERP sales orders, with targets such as plant and shipping condition predicted from the order header, items and linked customer data [1]. The same formulation covers a general ledger (GL) account on a supplier invoice line, a cost center on an expense item, a payment term on a new customer, or an opportunity stage and product family in a CRM.

RelBench v2 generalizes this as autocomplete tasks: mask an attribute in a row of a relational database and predict it from the rest of the row plus related rows, subject to time constraints [2]. That builds on the original RelBench benchmark of predictive tasks over multi-table databases [5]. Commercial tooling follows the same loop. SAP's Data Attribute Recommendation service, covered in SAP developer tutorials, was built around uploading historical records against a defined dataset schema, training a model and serving recommendations [3]; check current availability before planning around it.

The practical consequence for buyers is that the "label" is just a column that already exists in the source system. The hard part is not annotation. It is reconstructing what the user could see at entry time and separating values a person chose from values the system derived.

The four layers a training extract needs

A usable dataset has four layers: the document records, the linked master data as of entry time, the field-level change history, and a dictionary that says how each field is populated. Most first-pass extracts deliver only the first, which is enough for a demo and not enough for a model you can trust on a live form.

  1. Document header and line tables. In SAP SD this is the sales document header and item tables (VBAK and VBAP); in an accounts-payable flow it is the invoice header and lines; in Salesforce it is Opportunity and OpportunityLineItem. Keep the stable document key, the line number and the created-at timestamp.
  2. Point-in-time master data. Customer, vendor, material and chart-of-accounts tables change. A customer's default shipping condition or a vendor's default GL account at entry time is a strong feature; today's value is a leak. Ask for effective-dated versions or change logs (for example SAP change documents in CDHDR/CDPOS, or Salesforce field history tracking) rather than a current snapshot.
  3. Field-level change and override history. Final values hide what the user had to fix. The most valuable signal is whether a field was defaulted by the system, typed, or overwritten, and by whom (role, not identity). This is a working hypothesis rather than a published finding: a model trained only on final values learns to reproduce defaults, including defaults that users routinely correct.
  4. Field population dictionary. For each candidate target and feature, record whether it is entered manually, defaulted from master data, derived by a rule or determination procedure, or filled by a later step such as delivery, billing or approval. The data dictionaries and schema documentation guide covers how to request this layer.

If you are licensing the underlying tables, the grain and key decisions are covered in what to specify when licensing tabular data.

Building time-respecting labels without leakage

Every training example should be built from a cutoff timestamp equal to the moment the target field was first entered, and every feature must be observable at or before that cutoff. RelBench v2 makes this explicit for its autocomplete tasks by respecting time constraints when constructing inputs [2]. Temporal and target leakage, where features carry information unavailable at prediction time, inflates offline accuracy and collapses in production; time-based splits are the standard mitigation [4].

Business documents leak in predictable ways:

  • Downstream documents. Delivery, goods issue, billing and payment records link back to the order. Their existence or values reveal the plant, shipping point or GL account the user eventually chose.
  • Derived fields. In SAP, shipping point and route are often determined from plant and shipping condition; using them as features to predict plant is circular.
  • Status and workflow fields. Approval status, block reasons and "last changed by" timestamps are set after entry.
  • Current master data. A customer record updated after the order often reflects the order itself.
  • Aggregates computed over the full history. "Most frequent plant for this customer" must be computed only over documents created before the cutoff.

Split train, validation and test by entry date, not at random, and hold out at least one later period. If you need to measure generalization to new customers or materials, add a second split that holds out entities first seen after the training window.

Label spaces, long tails and how much history you need

Autocompletion targets range from a handful of classes to thousands of codes, and the long tail decides how much data you need. A shipping condition may have a dozen values; a GL account or material group in a large chart of accounts can run to thousands, many used a few times a year. A reasonable working assumption is that long-tail targets need either several companies' histories or many years from one company before rare codes have enough examples.

Report metrics that match the user experience. Top-1 accuracy matters if the field auto-fills; top-k recall matters if you show a dropdown of suggestions. Track coverage at a confidence threshold, since a recommender that abstains on uncertain lines is usually preferred over one that fills them wrongly.

Multi-company data adds a mapping problem. Charts of accounts, plant codes and payment terms are company-specific, so you either train per-tenant heads, map codes to a shared taxonomy, or treat codes as entities with learned embeddings. The guide on how much tabular data you need covers sizing in rows, tables and time span.

Request specification for an autocompletion dataset

A precise request names target fields, the entry event that defines the cutoff, the linked tables, and the change history you expect. The template below is a starting point for an ERP sales-order use case; swap the field names for AP invoices, expense lines or CRM opportunities.

Illustrative example: invented to show structure; it does not describe an available dataset.

ElementWhat to specifyWhy it matters
Source system and modulee.g., ERP sales and distribution; AP invoice entry; CRM opportunitiesField semantics and defaulting logic differ by module
Target fieldse.g., plant, shipping condition, payment terms, GL accountEach target needs its own population rule and cutoff
Entry eventDocument create timestamp and first-save timestamp per fieldDefines the time cutoff for features
Document grainOne row per header and one per line, with stable keysLine-level targets need line context
Linked tablesCustomer, material, vendor, chart of accounts, effective-datedPoint-in-time features without leakage
Change historyOld value, new value, timestamp, changed-by role, change source (user, rule, interface)Separates defaults from user choices; yields correction labels
Population dictionaryManual, defaulted, derived or filled-later flag per fieldLets you exclude leaky features
Time spanContinuous months, including at least one full fiscal yearSeasonality and year-end behavior
Code listsFull value lists with descriptions and validity datesHandles retired and new codes
De-identificationCustomer and user names, emails, phones and account numbers removed or replaced; method documentedPrivacy review and allowed-use scoping

An illustrative record after preparation might look like this:

{
  "doc_id": "SO-000184233",
  "line_no": 20,
  "entry_ts": "2025-03-14T09:42:11Z",
  "customer_key": "C-7f3a",
  "material_key": "M-19c0",
  "order_qty": 120,
  "sales_org": "1000",
  "customer_default_shipcond_asof_entry": "02",
  "target_field": "plant",
  "final_value": "1200",
  "system_default_value": "1100",
  "user_overrode_default": true,
  "excluded_features": ["shipping_point", "delivery_doc_id", "billing_doc_id"]
}

Correction history as a second training signal

Override and correction events are the closest thing to human feedback a business table holds. When a user replaces a defaulted value, the pair (default, chosen value) tells you where the current rule is wrong; when a value is corrected days later, often after a delivery or posting fails, the original entry is a negative example.

Use these events in three ways. Weight examples where users overrode a default, because those are where a recommender adds value. Build an evaluation slice from corrected records to measure whether the model would have avoided the error. Train a separate "is this value likely to be changed" classifier that can flag fields for review instead of filling them.

The neighboring problem of cleaning dirty tables with known fixes is covered in dirty tables with ground-truth corrections. For extraction from scanned documents rather than entry-time prediction, see document extraction evaluation ground truth.

Choosing among public benchmarks, synthetic data and licensed histories

Public benchmarks are right for method development, but they rarely match your fields, codes or defaulting rules. SALT gives real ERP sales tables with a defined autocompletion framing [1], and RelBench v2 offers autocomplete tasks across several relational databases with time-aware construction [2][5]. Neither tells you how your customers' configuration and user behavior shape the target distribution.

Synthetic tables help with schema testing and privacy-safe demos, but they cannot reproduce the override patterns and long-tail codes that make the task hard; the synthetic tabular data evaluation guide explains why. Licensed operational histories from companies that run the same module are the main route to production-grade training and evaluation sets. For broader context on relational learning over these tables, see multi-table relational data for relational deep learning and ERP transaction and master data for AI training.

Document whatever you assemble with a dataset card that records source systems, time span, target definitions, exclusions and known biases [6]. Finance-focused teams can also review the finance operations AI use case and the accounting reconciliations dataset category, which describe adjacent workflows where field recommendations are common. If your model needs histories from specific ERP, CRM or finance workflows, you can describe the data you need to SourceX.

Diligence questions before you license business-table histories

Ask suppliers questions that expose leakage risk, rights and preparation before you commit to a schema. Data from operational systems carries customer, vendor and employee information, so privacy and allowed use must be settled before features are engineered.

  • Which fields are user-entered, defaulted, derived or filled later, and is that documented per field?
  • Are master-data tables effective-dated, or can change logs reconstruct the state at entry time?
  • Does the extract include field-level change history with change source, and how are user identities replaced?
  • Were names, emails, phones and account numbers removed or replaced, which method was used, and was a sample checked?
  • Does the license define the records covered, permitted uses (training, evaluation or both), term and delivery method?
  • Were any fields populated by a vendor's own recommendation feature, which would make your labels partly machine-generated?

Sourcing autocompletion training data through SourceX

SourceX sources operational datasets from US companies on request, including sales histories, finance and legal workflows and other business records, and manages licensing and ongoing purchases. Each dataset is rights-reviewed, has personal details removed or replaced with the method recorded, and is delivered under a license defining records, uses, term and delivery; a request does not guarantee a match. Explore more in the tabular and transactional data buyer's guide or the AI data hub, then request field autocompletion training data from SourceX.

Sources

  1. Silicon Saxony, "SAP: Advancing Enterprise AI Research with First Real ERP Dataset". https://silicon-saxony.de/en/sap-advancing-enterprise-ai-research-with-first-real-erp-dataset/
  2. arXiv, "RelBench v2: A Large-Scale Benchmark and Repository for Relational Data" (2026). https://arxiv.org/html/2602.12606v1
  3. SAP, "Machine Learning tutorials (SAP Developers tag page)". https://developers.sap.com/tags/topicmachine-learning/
  4. Machine Learning Mastery, "3 Subtle Ways Data Leakage Can Ruin Your Models (and How to Prevent It)". https://machinelearningmastery.com/3-subtle-ways-data-leakage-can-ruin-your-models-and-how-to-prevent-it/
  5. arXiv, "RelBench: A Benchmark for Deep Learning on Relational Databases" (2024). https://arxiv.org/pdf/2407.20060
  6. Hugging Face, "Create a dataset card (Datasets library documentation)". https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data