Agent, workflow and domain-reasoning data
Cross-system workflow records: linking ERP, CRM, ITSM and email into one task timeline
Quick answer
Cross-system workflow records are the events of one piece of work, such as a customer escalation that becomes a credit memo, pulled from each system it touched and merged into a single ordered timeline. For agent training and evaluation, the value is in the joins: shared keys, consistent pseudonyms, corrected clocks and a coverage report showing how many cases link end to end. Buy the linkage method and its error rates, not just the tables.
By SourceX Editorial · Updated
Why single-system exports underperform for multi-application agents
Single-system exports teach an agent one application, while real tasks fail at the handoffs between applications. Integration vendors argue that operational agents often stall because they cannot reach context held in other systems, so routing and verification decisions rest on incomplete records [2]. Research prototypes now test coordinating agents on serial tasks that span several heterogeneous enterprise systems [1].
Public benchmarks show the gap. OSWorld includes multi-application workflows scored by execution checks [3], WebArena reports end-to-end success far below human levels on self-hosted sites [4], and tau-bench pairs agents with databases, policy documents and simulated users [5]. None of them contains your customers' real handoff patterns: a Salesforce case that spawns a ServiceNow incident, an SAP return order, and an email thread with the carrier. That is what licensed linked records add.
For the broader map of agent data types, start at the agent training data hub; for the product-level view, see enterprise workflow datasets and agent trajectories and the guide to connecting ticket histories with documentation.
The linking keys that actually join business systems
Records join reliably only on keys that one system writes into another; fuzzy matching on names and dates should be a fallback with a measured error rate. Ask each supplier which of these keys exist, where they live and how often they are populated.
- Business document numbers. Sales order, purchase order, delivery and invoice numbers (for example SAP VBELN and EBELN, or NetSuite tranid) are the strongest joins in order-to-cash and procure-to-pay chains.
- Case and ticket IDs. CRM case numbers and ITSM incident or request numbers (ServiceNow INC/RITM, Jira issue keys) often appear in each other's reference or correlation fields when an integration exists.
- Email thread identifiers. Message-ID, In-Reply-To and References headers rebuild reply chains [8]; ticket numbers in subject lines are a weaker secondary key.
- Customer, vendor and employee master IDs. Account IDs and vendor numbers join across systems only if master data is synchronized; mismatched masters are the most common silent failure.
- Integration correlation IDs. Middleware logs (MuleSoft, Boomi, Workato run IDs) sometimes carry the only key connecting two systems.
The process-mining community already has a format for this shape of data. OCEL 2.0 lets one event relate to many objects (order, invoice, case) and records object-to-object relations and attribute changes over time [6], and the public BPI Challenge 2019 purchase-order log has been converted to object-centric form [7]. Asking suppliers for an OCEL-like structure is more useful than asking for one flat case ID, because real work fans out and merges. See process mining event logs for agents for the event-log side and verifying record linkage across systems for test methods.
Pseudonymize once, with one key map, before any join breaks
Identities must be replaced consistently across every system, or the timeline falls apart after de-identification. If the CRM export hashes a customer email one way and the email archive redacts it another way, the join key disappears. The usual pattern is one keyed tokenization (for example HMAC with a supplier-held secret) applied to every identifier type in every source, so the same person or account maps to the same token everywhere.
Consistency cuts both ways. A stable token that appears across five systems builds a richer profile than any single export, which raises re-identification risk. Under California's CCPA, information is "deidentified" only if it cannot reasonably be linked to a particular consumer and the business meets conditions that include contractual commitments from recipients [9]. Ask suppliers to document which fields were tokenized, which were dropped, how free text (email bodies, ticket comments, notes) was scrubbed, and what residual risk review was done on the joined output, not just on each table.
Clock skew, time zones and event ordering
Merged timelines are only as trustworthy as their timestamps, and every system records time differently. Typical problems include ERP posting dates with no time component, CRM fields stored in UTC but exported in the user's locale, email Date headers set by the sender's client, and ITSM records that store both opened_at and sys_created_on with different meanings.
Before merging, require each source to declare its time zone, precision and which field means "when it happened" versus "when it was recorded." Normalize to UTC in ISO 8601 with an explicit offset, keep the original value, and flag events whose order is ambiguous because they share a date but lack a time. For agent evaluation, ambiguous order matters: a grader that checks "credit issued after approval" will produce false failures if posting dates round to midnight. Long waits and handoffs are measured from these same fields; see long-horizon task records.
A coverage report: the one document that tells you if linkage worked
The coverage report states what share of cases link fully across the systems in scope, and buyers should treat it as a deliverable. Without it, you cannot tell whether a 40-event timeline is complete or whether half the work happened in a system that was not exported.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Metric | Definition | Why it matters |
|---|---|---|
| Cases in scope | Root objects (e.g., CRM cases opened in period) | Denominator for every rate below |
| Fully linked share | Cases with at least one event in every in-scope system | Upper bound on usable end-to-end tasks |
| Join method mix | Share joined by hard key vs. fuzzy match | Fuzzy joins need sampled precision |
| Sampled join precision | Manually checked joins that were correct | Estimates false links in training data |
| Orphan events | Events not attached to any case | Signals missing keys or out-of-scope systems |
| Ambiguous ordering | Events with date-only or tied timestamps | Affects step-order labels and graders |
| Redaction impact | Joins lost after pseudonymization | Shows whether de-identification broke linkage |
Pair the report with a data dictionary for the delivery and machine-readable provenance; Croissant-RAI is one vocabulary for recording how records were prepared [10]. Packaging choices for linked tables are covered in packaging linked records from multiple systems.
An illustrative linked timeline record
A good unit of delivery is one case with its objects, events and links, not separate dumps per system. The record below shows the minimum fields an agent researcher needs to build a trajectory, an environment seed or a graded evaluation task.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"case_id": "case_7f3a",
"objects": [
{"id": "crm_case:tok_91c2", "type": "crm_case", "system": "crm"},
{"id": "itsm_inc:tok_44be", "type": "incident", "system": "itsm"},
{"id": "erp_so:tok_0d17", "type": "sales_order", "system": "erp"},
{"id": "erp_cm:tok_5a80", "type": "credit_memo", "system": "erp"},
{"id": "thread:tok_e2f9", "type": "email_thread", "system": "mail"}
],
"events": [
{"ts": "2025-03-04T14:02:11Z", "ts_precision": "second", "system": "crm",
"activity": "case_opened", "objects": ["crm_case:tok_91c2", "erp_so:tok_0d17"],
"actor": "agent_tok_12"},
{"ts": "2025-03-04T14:20:00Z", "ts_precision": "second", "system": "itsm",
"activity": "incident_created", "objects": ["itsm_inc:tok_44be", "crm_case:tok_91c2"],
"link_method": "correlation_field"},
{"ts": "2025-03-06", "ts_precision": "day", "system": "erp",
"activity": "credit_memo_posted", "objects": ["erp_cm:tok_5a80", "erp_so:tok_0d17"],
"link_method": "document_reference"}
],
"outcome": {"resolved": true, "label_source": "crm_status_closed"}
}
Note the day-precision ERP event and the per-event link_method; both should be visible to the buyer rather than silently smoothed. Outcome labels deserve their own definition work; see task success labels for agent trajectories.
Sourcing questions to send before any sample
Most linkage problems are visible in a supplier's answers before any data moves. Use this checklist in your first request.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Which systems and versions are in scope (e.g., Salesforce Service Cloud, SAP S/4HANA, ServiceNow, Microsoft 365 mail), and which ones touched the process but are excluded?
- What is the root object, and which hard keys connect each system to it?
- What share of cases link fully, and how was fuzzy-join precision sampled?
- Was one pseudonymization key map applied across all sources, and how were free-text fields handled?
- What time zone and precision does each timestamp field carry?
- Do the supplier's agreements with its software vendors or customers restrict exporting or licensing these records? Some SaaS and outsourcing contracts limit data use; confirm with counsel.
- Are attachments (invoices, PDFs, screenshots) included and linked to events, or only referenced?
- Will the same extraction be repeatable for later periods, so train and held-out evaluation sets come from the same pipeline?
Licensing scope for replay, environments and derived benchmarks is its own question; see license terms for agent workflow data. If you are weighing real records against commissioned or generated data, compare options in licensed vs. commissioned vs. synthetic trajectories.
How SourceX handles cross-system record requests
SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing and ongoing purchases. Data is sourced on request rather than held in stock, so a request for linked CRM, ERP and ticket records starts a search for US businesses that hold them, with no guarantee of a match. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked (no method is perfect), and the supplying company approves every release. You can describe the linked records you need and the systems involved.
Request linked cross-system workflow data
SourceX looks for US businesses that hold the multi-system records you describe and runs the process from finding and assessing the data through agreeing a license that defines records, uses, term and delivery. Nothing is contracted until the supplier agrees, and terms are set per deal. Start by describing your systems, linking needs and use case at SourceX for buyers.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Frequently asked questions
Can I build cross-system timelines from separate single-system purchases?
Only if every source shares keys and the same pseudonymization map. Buying CRM and ERP data from different companies cannot produce joined timelines, and buying them from one company at different times often breaks token consistency.
Is OCEL required as the delivery format?
No. OCEL 2.0 is a useful reference model for many-to-many object relations [6], but relational tables with an explicit event-object link table carry the same information. What matters is that one event can reference several objects.
How do linked records differ from approval records?
Approval records focus on one decision point and its rationale; linked records follow the whole task across systems. For the decision-focused variant, see approval and rejection records.
Sources
- arXiv, "Agentic Nesting: A New Methodology for Existing Enterprise Application Integration and Services" (2026). https://arxiv.org/pdf/2608.05159
- Airbyte, "AI operations agents: workflow use cases and tooling". https://airbyte.com/agentic-data/ai-operations-agents-workflow-use-cases-tooling
- Xie et al., arXiv (NeurIPS 2024), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
- Zhou, Xu et al., arXiv, "WebArena: A Realistic Web Environment for Building Autonomous Agents" (2024). https://arxiv.org/abs/2307.13854v4
- Sierra Research, arXiv, "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- Berti, Koren, Adams, van der Aalst et al., arXiv, "OCEL (Object-Centric Event Log) 2.0 Specification" (2024). https://arxiv.org/pdf/2403.01975
- 4TU.ResearchData, "BPI Challenge 2019 (OCEL)". https://data.4tu.nl/datasets/46a7e15b-10c7-4ab2-988d-ee67d8ea515a
- IETF, "RFC 4021: Registration of Mail and MIME Header Fields" (2005). https://datatracker.ietf.org/doc/rfc4021
- California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- Jain et al. (MLCommons), arXiv, "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.