Enterprise workflow datasets: task execution histories for AI agents
An enterprise workflow dataset is a set of task execution histories: connected records that follow one business task from the incoming request through the context available at the time, the actions taken in tools, decisions, handoffs and the outcome, plus QA where available. SourceX builds these trajectories from the records and audit logs of systems such as Zendesk, Salesforce, Jira and NetSuite, de-identifies them and licenses them for an agreed use. They are system-level tool-use histories, not screen recordings.
Dataset manifest
Sourced to your spec- What it is
- Linked task trajectories from request to outcome, across every tool the work touched
- Typical systems
- Zendesk, Salesforce, Jira, ServiceNow, NetSuite, GitHub, Slack, internal tools
- Typical history
- Varies by partner; detailed action logs may cover less time than records
- Modality
- Structured event sequences with ticket, note and message text
- Delivery formats
- Agreed per order; JSONL trajectories or flat event tables keyed by task
- Preparation
- Records linked across systems; people and accounts de-identified; timestamps and outcomes kept
- Licensing
- Rights confirmed per underlying system; non-exclusive, or exclusive snapshot where agreed
- Availability
- Depends on partners whose systems share linkable IDs; not guaranteed
What a delivery contains
Fields vary by source system and are fixed per order. A typical delivery includes:
| Field | Type | What it holds |
|---|---|---|
| trajectory_id | string | Pseudonymous ID for one task instance, from the triggering request to the final outcome. |
| task_type | enum | The partner's workflow category, such as warranty claim, access request or vendor onboarding. |
| input | object | The request or event that started the task, with channel, timestamp and de-identified content. |
| context | array | What was available while the work happened, such as knowledge-base searches and record values as they stood at that moment. |
| steps | array | Ordered actions in tools, each with timestamp, actor role, system, action and target record. |
| steps[].action | string | Normalized action name, such as apply_macro or create_rma, with the system's raw event name kept alongside. |
| steps[].state_change | object | Before and after values of the fields an action changed, taken from audit logs or field history. |
| decisions | array | Decision points with the options open, the path chosen, who chose it and the recorded basis. |
| handoffs | array | Transfers between people, teams or queues, with the wait until the receiver's first recorded activity. |
| exceptions | array | Departures from the standard path, such as policy overrides, tool errors, rework or missed SLAs. |
| outcome | object | Final status and code, elapsed time, and later signals such as reopens, returns or disputes. |
| qa | object | QA scores by rubric dimension and reviewer comments, where the partner reviews this workflow. |
| provenance | object | Source logs, join keys and link method for each step, so every event traces back to a system record. |
Example record
{
"trajectory_id": "trj_5e19b2",
"task_type": "warranty_replacement",
"input": { "t": "2024-02-06T15:02:11Z", "channel": "email", "system": "zendesk",
"text": "My [PRODUCT_1] shows error E42 and won't power on. Serial [SERIAL_1]." },
"context": [
{ "t": "2024-02-06T15:09:30Z", "type": "kb_search", "system": "help_center",
"query": "E42 power fault", "opened": ["kb_e42_reset"] },
{ "as_of": "2024-02-06T15:02:11Z", "type": "record_state", "system": "netsuite",
"record": "serial_registration", "values": { "warranty_end": "2024-01-25" } }
],
"steps": [
{ "t": "2024-02-06T15:10:48Z", "actor": "agent_t1", "system": "zendesk",
"action": "apply_macro", "raw_event": "macro_applied", "target": "e42_reset_steps" },
{ "t": "2024-02-06T15:24:16Z", "actor": "customer", "system": "zendesk",
"action": "reply", "text": "Did the reset. Same error." },
{ "t": "2024-02-06T15:26:52Z", "actor": "agent_t1", "system": "zendesk",
"action": "escalate", "state_change": { "group": ["tier_1", "warranty"] } },
{ "t": "2024-02-07T09:14:02Z", "actor": "warranty_specialist", "system": "netsuite",
"action": "create_rma", "state_change": { "rma_status": [null, "approved"] } },
{ "t": "2024-02-07T09:15:20Z", "actor": "warranty_specialist", "system": "carrier_portal",
"action": "create_return_label", "target": "rma_7731" },
{ "t": "2024-02-07T09:16:45Z", "actor": "warranty_specialist", "system": "zendesk",
"action": "set_status", "state_change": { "status": ["open", "solved"] } }
],
"decisions": [
{ "t": "2024-02-07T09:11:37Z", "by": "warranty_specialist",
"options": ["deny", "repair", "replace"], "chosen": "replace", "basis": "goodwill_policy_v3" }
],
"handoffs": [ { "from": "agent_t1", "to": "warranty_specialist", "wait_min": 1065 } ],
"exceptions": [ { "type": "policy_override", "detail": "warranty ended 12 days before the request" } ],
"outcome": { "code": "replacement_shipped", "elapsed_min": 1095, "reopened": false,
"return_inspection": "fault_confirmed" },
"qa": { "rubric": "2024.1", "process": 5, "policy": 4, "communication": 4,
"note": "Right goodwill call. State the return deadline in the reply." },
"provenance": { "link_method": "shared_ids", "join_keys": ["ticket_id", "serial", "rma_id"],
"sources": ["zendesk_ticket_audits", "netsuite_system_notes", "carrier_api_log"] }
}Synthetic record for illustration. Field names, structure and format are agreed per order.
What AI teams use it for
Fine-tune agents on multi-step tool use
Trajectories show which tools experienced staff open, in what order and what they change, under the partner's real policies and permission boundaries.
Build outcome-graded evaluation tasks
A trajectory's input and point-in-time context become the task, and the recorded end state becomes the reference an agent's result is graded against, on work no model has seen online.
Train reward models on business results
QA scores, reopens and downstream results label which paths worked, a reward signal grounded in what happened after the task closed.
Teach escalation and approval judgment
Decision points and handoffs show when a case needed a specialist or a sign-off, a judgment an autonomous agent has to make without being told.
Build realistic simulated environments
Field-level state changes show how business systems respond to each action, which helps when building tool mocks and RL environments that behave like production systems.
Use-case guides: Enterprise and computer-use agents, Private evaluation sets, Customer support agents, Finance and accounting agents
What makes this data valuable
Cross-system links
Steps joined across tools through shared identifiers, with the link method recorded.
Point-in-time context
Record values as they stood when each step happened, not as they look today.
Exception paths
Overrides, rework, approvals and reopened cases, where scripted policies run out.
Outcome horizon
Results measured after closing, such as reopens, returns or disputes, not only the final status.
Field-level changes
Before-and-after values from audit logs show exactly what each action did.
Human feedback
QA reviews, approvals and reviewer notes that judge how the work was done as well as its result.
Final states hide the work
Most exports from business systems describe how records ended: a ticket marked solved, an invoice paid, an opportunity closed. The path that produced that state is what an agent has to learn. In support, that path runs from the customer's request (input), through the ticket and knowledge searches (context), the agent's steps in each tool (actions) and the choice to escalate or resolve (decision), to the resolution and its QA score (outcome). It lives in ticket audits, field-history tables, system notes, changelogs and event streams, spread across every tool the task touched. A single refund can leave traces in the help desk, the billing system, the payment processor's logs and a chat thread with finance. No one system holds the trajectory.
Rebuilding it means extracting events from each system, mapping them to a shared action vocabulary, attaching each event to the right task and recomputing what each record looked like at every step. That last part is easy to get wrong. If context is filled from today's records instead of the state at the time, a model learns to act on information the employee did not have. Many systems also log changes but not views, so what someone looked at is often inferred from what they linked or cited, while what they changed is observed directly.
Where reconstruction goes wrong
- Automation noise. Triggers, workflow rules, integrations and bots write events that look like actions. Unless they are flagged, they inflate trajectories and teach a model to perform steps nobody performed.
- Clock skew. Systems stamp time differently, some at the moment of action and some at sync time, so ordering across systems has to be checked rather than assumed.
- Merges, splits and migrations. Merged tickets, duplicate accounts and a system migration partway through the history break chains unless old and new IDs are mapped.
- Off-system work. Phone calls, side conversations and personal spreadsheets leave no event. A good delivery marks these gaps so that two distant steps do not look adjacent.
That is why SourceX scopes a workflow dataset from one workflow and the systems it crosses, not from a list of exports: where each stage of the task is recorded, which identifiers connect the systems, how far back detailed logs go and what counts as the outcome. A sample of trajectories is built end to end for your review, so the action vocabulary, the link-confidence threshold and the outcome definition are agreed before the full dataset is prepared and licensed.
What to check before licensing
- Confirm rights for every system in the chain, not only the main one. A trajectory that passes through a client's portal or an outsourcer's tooling can carry data the partner does not own.
- Ask for chain-completeness figures on the sample: how many trajectories run unbroken from input to outcome, and how many events could not be attached to any task.
- Ask how the outcome is defined and when it is measured. A case marked solved that reopens two weeks later, or an invoice paid and later disputed, should not count as a success.
- Review the action vocabulary and its mapping from raw events. Check that rare but important actions, such as overrides or refunds above a limit, were not collapsed into generic updates.
- Get the policy and SOP versions in force across the covered period, so a change in behavior can be traced to a rule change instead of being learned as noise.
- Count how many trajectories follow the standard path and how many take an exception path, per task type. Routine paths usually dominate, and exceptions are where agents fail.
- Settle in the license whether simulated environments or tool mocks built from the data count as derivatives, and what happens to them when the license ends.
How licensing works through SourceX
- 1
Define
Send the domain, modality, volume, format, timeline and permitted use you need.
- 2
Source
SourceX identifies businesses that hold matching data and are open to licensing it.
- 3
Qualify
Fit, rights and quality are checked, and you review samples before committing.
- 4
License
Scope, permitted use, exclusivity, price and obligations are agreed in writing.
- 5
Deliver
Approved data is prepared, de-identified where required and transferred securely.
Questions buyers ask
Are these screen recordings of employees at work?
No. They are system-level action histories reconstructed from the records, audit logs and event streams of business systems: which action was taken, in which tool, by which role, on which record, when, and what changed as a result. There are no pixels, cursor movements or keystrokes. For computer-use agents they supply task structure, decision points and outcomes; GUI-level demonstrations are a different kind of data.
How are records from different systems linked into one trajectory?
Through shared identifiers wherever they exist, such as a ticket ID stored on an order, an RMA number quoted in a ticket or an account ID carried from CRM to billing. Where no shared key exists, events can be matched on time windows and entities, which is less certain. Links can carry their method, so you can keep only deterministic links or weight the rest.
How does a task execution history differ from an agent trajectory?
The structure is the same, but the actors are people. An agent trajectory records a model's observations, tool calls and results; a task execution history records the request, the context people consulted, their actions in tools, decisions, handoffs and outcome. Converting one into the other means mapping actions onto your tool schema and context onto observations, which is why normalized action names and raw event names are both kept.
Can these histories be used for reinforcement learning?
Yes, as offline data. Outcomes, QA scores and later signals such as reopens can serve as rewards, and recorded state changes help build simulated environments. What the data cannot show is what would have happened after a different action, so it suits imitation learning, reward modeling, offline RL and evaluation better than online exploration.
Which workflows can be sourced?
Workflows a partner runs through systems with usable records, subject to rights. Typical patterns are support (request, ticket, knowledge search, agent actions, escalation, resolution, QA score), sales (research, outreach, calls, proposal, negotiation, won or lost), engineering (issue, review, tests, deployment, incident follow-up), finance (transaction, categorization, reconciliation, approval, exceptions) and recruiting (requirement, sourcing, screening, interviews, placement). Availability depends on partners.
How far back do task execution histories go?
That differs from partner to partner, and the limit is often the action log rather than the records. A company may keep years of closed tickets or invoices while detailed audit logs cover a shorter span, depending on the system, its settings and past migrations. Ask for the period covered by detailed logs in each system, so you can see where trajectories thin out into final states.
How is personal data handled when records are linked?
It is de-identified with one pseudonym scheme across all systems, so the same customer, employee or account keeps one ID at every step; separate schemes per system would break the links that make a trajectory. Free text in notes, messages and comments is de-identified too, and actors appear as roles with stable pseudonymous IDs.
Related datasets
- SOPs, playbooks and internal knowledge bases
Written procedures with page history, ownership and links to execution records
- Human feedback and QA-scored work
Work items with scores, verdicts and corrections from the people who reviewed them
- Customer support ticket datasets
Resolved support cases with full threads, internal notes and outcomes
- Accounting and reconciliation workflows
Coded transactions, reconciliations and close records with reviewer corrections and approvals
- Software engineering histories (issues, PRs, reviews)
Issues linked to commits, pull requests, code review, CI runs, deploys and incidents
Evaluating this data for procurement?
Diligence packets are prepared per dataset. Rights, privacy processing and quality differ between datasets.
Request dataset diligenceTell us what your models need
Send your spec — domain, volume, format, timeline and permitted use — and SourceX will match it against partner data and come back with what can be licensed.
Updated 3 October 2026. Own data like this? See how companies license it to AI developers.