Skip to content

Tables, time series and transactional data

Timestamped Business Event Sequences for Temporal and Sequence Models

Quick answer

Event sequence data for machine learning is a log of what happened to each business entity, in order: one row per event, keyed by entityid, with an eventtype, an eventtime, an ingesttime and typed attributes. It trains next-event, time-to-event and temporal graph models only if the buyer specifies the entity grain, both clocks, deduplication rules and a time-based split. Most failures trace to late events, retry duplicates and random splits that leak the future into training.

By SourceX Editorial · Updated

What counts as a business event sequence

A business event sequence is an ordered stream of state changes tied to a durable entity, such as a customer, account, order, device or shipment. Typical sources are order management (order_placed, order_shipped, order_returned), billing (invoice_issued, payment_failed, payment_retried, subscription_canceled), identity (login_succeeded, mfa_failed, password_reset) and CRM or ticketing status changes. The model target is usually the next event type, the time until a specific event, or a link between two entities at a future time.

This page covers entity streams used as predictive inputs. Process-mining case logs, where each trace is one execution of a workflow, are a different shape and use; see object-centric event logs (OCEL 2.0) and the owner page on enterprise workflow datasets and agent trajectories. The two overlap at the raw-table level, which is why the entity grain needs stating before anything else.

The minimum event schema to request

The minimum usable schema is six fields: entity key, event type, event time, ingest time, attributes and source system. Without ingest_time you cannot reproduce what a model would have known at prediction time, and without source_system you cannot explain duplicates that arrive from two pipelines. Ask for a controlled vocabulary for event_type, with a mapping table if the supplier renamed types over the years.

Use a delivery format that preserves types. Parquet partitioned by event date is common for large streams; JSON Lines works for nested attributes as long as each line is one valid UTF-8 JSON object with no blank lines [7]. Ask for a machine-readable data card, for example in Croissant, the schema.org-based JSON-LD vocabulary that describes files and record structure [8].

Illustrative example: invented to show structure; it does not describe an available dataset.

{"event_id": "e-000184223", "entity_id": "acct_7f3a", "entity_type": "account", "event_type": "payment_failed", "event_time": "2025-03-14T09:12:44.512Z", "ingest_time": "2025-03-14T09:13:02.001Z", "source_system": "billing_v2", "sequence_no": 41, "attributes": {"amount_minor": 4900, "currency": "USD", "decline_code": "insufficient_funds", "retry_attempt": 0}, "related_entities": [{"type": "invoice", "id": "inv_19c2"}]}
FieldWhy it mattersFailure if missing
event_idStable key for deduplicationRetries counted as new events
entity_id, entity_typeDefines the sequence grainMixed grains in one sequence
event_time (UTC, ms)Ordering and inter-event gapsTies and timezone shifts corrupt intervals
ingest_timePoint-in-time feature reconstructionSilent use of late-arriving data
sequence_noTie-breaking within one entityNon-deterministic order on equal timestamps
source_systemLineage and duplicate diagnosisUnexplained double events after migrations
related_entitiesTemporal graph and relational linksGraph models cannot be built

Event time, ingest time and late events

Every event stream has two clocks, and a model that ignores the gap between them will look better offline than it performs in production. Event_time says when the business change happened; ingest_time says when the warehouse saw it. Mobile apps buffer events offline, payment processors post settlements a day later, and backfills after an outage can land weeks of history with the same ingest_time.

Ask the supplier for the distribution of ingest_time minus event_time per source system, plus a list of known backfill windows. Decide in writing whether features are computed as of event_time or as of ingest_time; for production parity, the safer default is ingest_time. Out-of-order arrival also matters for streaming evaluation, so request the original arrival order or keep ingest_time at millisecond resolution.

Deduplication, retries and status flapping

Duplicates in business streams are usually retries, not errors, and they should be resolved by documented rules rather than by dropping rows with identical timestamps. Webhook redelivery, at-least-once message queues and client retries all create near-identical events with different event_ids. A payment_retried event is meaningful, while a duplicated payment_failed from a redelivered webhook is not.

Request the idempotency key or upstream message ID where it exists, and ask how the supplier marks retries. Also ask about status flapping, where a record toggles between two states within seconds because of a sync loop; this can dominate next-event targets if left in. Record the deduplication rule in the dataset card so evaluation sets are cleaned the same way as training sets.

Sessionization and sequence windows are modeling choices

Session boundaries, inactivity thresholds and sequence truncation are your decisions, not properties of the data, so request raw events and document the choices you apply. A 30-minute inactivity gap suits web clickstreams; a 90-day gap may suit B2B purchasing. If the supplier pre-sessionized the data, ask for the threshold used and whether the raw stream is also available.

The same applies to censoring for time-to-event work. A temporal point process or survival model needs to know when observation of each entity started and ended, so request first_seen, last_seen and an extraction cutoff timestamp. Without them, an account that simply left the extract looks identical to one that churned. For subscription targets, see subscription, billing and usage histories for churn prediction.

Streams, tables and temporal graphs are interchangeable views

One event stream can be modeled as sequences, as relational tables or as a temporal graph, and the choice should follow the model, not the delivery format. RelBench v2 converts temporal graph benchmark event streams into relational schemas, showing that the same events can feed relational deep learning [1]. If you plan that route, ask for entity tables (accounts, products, merchants) alongside the event table, keyed consistently; see multi-table relational data for relational deep learning.

For temporal graph models, each event with two related entities becomes a timestamped edge, and evaluation should rank the true future link against many candidates rather than score it against a single random negative, which tends to inflate results [9]. Business data is usually heterogeneous: customers, merchants and products are different node types with different edge semantics. Ask for related_entities on every event so you can build the bipartite or heterogeneous graph yourself.

Standard exchange formats exist mainly on the process-mining side. XES (IEEE 1849-2023) is an XML format for event logs and streams [6], and OCEL 2.0 offers relational SQLite, XML and JSON formats for events linked to multiple objects [5]. They are useful if a supplier already exports them, but most predictive teams will flatten to one event table plus entity tables.

Splitting by time so evaluation is honest

Event-sequence evaluation must split by time, and usually by entity as well, because random row splits leak future behavior into training. RelBench uses temporal splits over timestamped relational events as its standard design [2]. In predictive process monitoring, random splits were shown to leak information across cases, and inconsistent splits made published results hard to reproduce [3].

A defensible design uses a training cutoff, a validation window and a test window, with features computed only from events whose ingest_time precedes each prediction timestamp. For next-event tasks, also hold out a set of entities never seen in training to measure cold-start behavior. For held-out evaluation built from real records, see building a golden evaluation dataset from business records.

Illustrative example: invented to show structure; it does not describe an available dataset.

SplitEvent window (event_time)EntitiesLabels from
Train2023-01-01 to 2024-12-3180% of entitiesEvents up to 2024-12-31
Validation2025-01-01 to 2025-03-31Same 80%Events in window
Test (temporal)2025-04-01 to 2025-06-30Same 80%Events in window
Test (cold-start)2025-04-01 to 2025-06-30Held-out 20%Events in window

Privacy and re-identification in event streams

Event sequences re-identify people more easily than static tables, because the order and timing of events is itself a fingerprint. Research on process-mining logs defines individual uniqueness measures and shows that event logs carry real re-identification risk even without names [4]. Removing emails and account numbers is necessary but not sufficient when exact timestamps, amounts and locations remain.

Ask the supplier how identifiers were replaced (salted hashing, tokenization or random surrogates) and whether the same surrogate is stable across tables, since joins depend on it. Ask whether timestamps were shifted or coarsened, and by how much, because that directly limits inter-event-time modeling; see what date shifting does to temporal models. Free-text attributes, such as support notes attached to status changes, need separate review.

Buyer request checklist for event sequence data

A useful request describes the entity, the events, the window and the evaluation plan, not a company. You can paste this into an internal spec or an external data request.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Entity grain: account, order, device or shipment; one primary entity_type per sequence.
  • Event vocabulary: required event types, with rename mappings over time.
  • Window and volume: event_time range, approximate entities and events per entity you need.
  • Clocks: event_time and ingest_time in UTC, millisecond precision, plus known backfill windows.
  • Lineage: source_system per event, idempotency or message IDs, deduplication rule applied.
  • Observation bounds: first_seen, last_seen and extraction cutoff for censoring.
  • Related entities: IDs needed for relational or temporal graph views, with entity tables.
  • De-identification: method for identifiers, timestamp treatment, free-text handling.
  • Format and documentation: Parquet or JSON Lines, data dictionary, Croissant or equivalent card.
  • Allowed uses: training, evaluation or both, and whether derived models may be deployed commercially.

For the general spec on grain, keys and history, see what to specify when licensing tabular data, and for the wider cluster, the tabular, time-series and transactional data buyer's guide.

How SourceX sources business event streams

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records and finance workflows, and it does not hold event data in stock. You describe the entity grain, event vocabulary and window; SourceX looks for US businesses that hold matching data, and every release is approved by the supplying company. A request does not guarantee a match. You can describe your event sequence requirements on the buyers page.

Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows, never email attachments, and only after an executed agreement and supplier approval.

Request event sequence data for your sequence models

If you need entity-keyed business event streams for next-event, time-to-event or temporal graph models, describe the data you need and SourceX will look for US companies that hold it and manage the licensing process. Pricing and allowed uses are agreed per deal, and nothing is contracted until a supplier agrees. Start a data request on the SourceX buyers page.

Sources

  1. arXiv, "RelBench v2 (arXiv:2602.12606)" (2026). https://arxiv.org/html/2602.12606v1
  2. arXiv, "RelBench: A Benchmark for Deep Learning on Relational Databases (arXiv:2407.20060)" (2024). https://arxiv.org/pdf/2407.20060
  3. arXiv (Weytjens and De Weerdt), "Creating unbiased public benchmark datasets with data leakage prevention for predictive process monitoring (arXiv:2107.01905)" (2021). https://export.arxiv.org/abs/2107.01905
  4. arXiv (Nunez von Voigt et al., CAiSE 2020), "Quantifying the Re-identification Risk of Event Logs for Process Mining" (2020). https://arxiv.org/pdf/2003.10707
  5. OCEL standard authors (arXiv:2403.01975), "OCEL (Object-Centric Event Log) 2.0 Specification" (2023). https://arxiv.org/pdf/2403.01975
  6. IEEE Standards Association, "IEEE 1849-2023 Standard for eXtensible Event Stream (XES)" (2023). https://standards.ieee.org/ieee/1849/10907
  7. jsonlines.org, "JSON Lines". https://jsonlines.org/
  8. arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  9. arXiv (Huang et al.), "Temporal Graph Benchmark for Machine Learning on Temporal Graphs" (2023). https://arxiv.org/abs/2307.01026

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data