Training data for enterprise and computer-use agents
Enterprise and computer-use agents learn from real task trajectories: the request that started a piece of work, the records and files consulted, each action in each system, the approvals and handoffs, and the final outcome, plus the SOP that governed it. SourceX sources these as workflow histories reconstructed from business systems' audit logs and event streams, not screen recordings, together with the SOPs, documents, approved email and chat, and review outcomes around them.
Dataset types to start with
- Enterprise workflow and task execution histories
Each task's trigger, context, actions in each system, decisions, handoffs and outcome arrive as one connected record, which supervises planning and tool sequencing and gives a recorded end state to check an agent against.
- SOPs, playbooks and internal knowledge bases
Written procedures are the instructions an enterprise agent will be handed. Set against execution histories, they show which steps people skip, reorder or add when real cases do not fit the document.
- Enterprise document archives
Tasks consume and produce files: the invoice to match, the form to complete, the template to fill. Archives with version history supply them in native formats.
- Approved workplace email and chat exports
Many tasks start as an email or chat message and pause for clarifications and approvals. Approved exports, scoped to the teams that run the process, carry the requests and reasons that system logs leave out.
- Human feedback and QA-scored work
Approvals, rejections, reviewer corrections and rework flags grade whether a task was done right, not only whether it finished, which makes trajectories usable as rewards and as eval rubrics.
- IT service management and incident histories
Access requests, provisioning, standard changes and incidents are well-bounded tasks with clear completion states and approval chains, a practical first domain for IT and operations agents.
Why this data is hard to get
Public agent tasks are staged
Public agent benchmarks mostly run on replica websites or freshly installed apps, with tasks written for the benchmark. They test operating an interface, not handling the exceptions, approvals and half-finished records of real back-office work.
No single system holds the task
A vendor onboarding touches email, a procurement tool, the ERP and a shared drive. Each export shows a fragment, and joining them needs internal IDs and identity mappings that exist only inside the company.
Logs record changes, not reasons
Audit trails say what changed and when, rarely why. The reasons sit in email, chat and comments, so a log-only history teaches sequences without the criteria that chose them.
Action-level detail expires first
Many systems keep audit and event logs for less time than the records themselves, so older work may survive only as end states, without the steps that produced them.
Activity data is employee data
Action logs show who did what. Licensing them raises employee privacy and notice questions, and in some jurisdictions works council consultation, on top of the customer and vendor data inside the records.
What a workflow history contains, and what it does not
SourceX's workflow histories are system-level action records reconstructed from business systems' logs, not screen recordings. The raw material is logging that business software already does: audit trails, record version histories, event streams and API logs in ERP, CRM, ticketing and approval tools. Events from several systems are joined by case, request or object IDs and ordered by time into one trajectory per task: the trigger, records consulted where views are logged, each create, update and approval with its parameters, the handoffs and the end state.
Reconstruction is where quality is won or lost. Clock skew between systems, batch jobs that touch many records at once, service accounts acting for people and steps done offline all have to be handled, or the trajectory will not reflect what a person did. Ask how task boundaries were drawn and which events were dropped or inferred.
Using system-level records for computer-use agents
A computer-use agent acts through a screen, but its task is defined by business systems. System-level histories supply what public benchmarks lack: a realistic task mix, exceptions included, a reference path at the level of records and decisions, and a recorded end state to check the agent's work against. One way to use them is to rebuild the relevant records in a sandbox, give the agent the original trigger, let it operate the applications, and compare the resulting state with the recorded one. Screen-level demonstrations, if you need them, can be generated in that environment. Agents that call APIs and tools directly can train on the records with less conversion, since system-level actions map closely to API operations.
Scoping a request around one process
Start with one process family rather than enterprise work in general: procure-to-pay, employee onboarding or access requests, for example. Name the systems involved, whether email and chat context is required, the exception types you need covered, and the period. Plan to hold out by time or by business unit, so the agent is tested on process variants it has not seen. SourceX looks for businesses that run that process; each candidate's manifest describes the systems, the logs available and how far back action-level detail goes, and you review sample trajectories before licensing.
What good data looks like
- Each trajectory has a defined trigger, start, end and outcome, joined across systems by case, request or object IDs.
- Every action records the system, object, action type, timestamp and parameters or before-and-after values, with its source log documented.
- Handoffs, approvals, rejections and rework loops appear as explicit events rather than gaps between timestamps.
- Each trajectory links to the SOP version in force and to the documents it read or produced.
- Each person keeps one pseudonym across every system, with role and team retained so handoffs stay readable.
- Rejected requests, corrections, cancellations and out-of-policy cases are included, with their share reported.
Questions buyers ask
Are SourceX workflow histories screen recordings?
No. They are system-level action records reconstructed from logs that business software already keeps. Each step names the system and record touched, the action, the values that changed, the time, and a pseudonymized actor with their role. They contain no screenshots, video, cursor movements or keystrokes: they record what was done in each system, not how the screen looked, so interface-level skills have to come from your own environment or other data.
Which business processes can workflow data cover?
Any process run through systems that keep usable logs, such as procure-to-pay, order-to-cash, onboarding, IT access requests, claims intake or dispatch. Supply depends on which businesses run the process at volume, log it in enough detail and agree to license the records, and log retention limits how far back action-level detail reaches.
Can the emails and chats behind a task be included?
Yes, where the business approves an export and the license covers it. Messages are usually limited to the teams and channels that run the process and linked to trajectories by case or request ID. They carry more third-party personal data and confidential content than system logs, so they get heavier de-identification, and privileged or unrelated threads are left out.
How should an agent be scored on historical tasks?
Score the end state and the constraints rather than step-by-step imitation. Real tasks often have several valid paths, so penalizing every deviation from the recorded sequence rewards memorization. Check whether the agent reached an equivalent final state in each system, respected approval rules and the SOP, and escalated where a human had to. Review outcomes tell you which recorded runs deserve to be references.
How are employees' identities handled in action logs?
Each person becomes one pseudonym that stays the same in every system, with role and team kept so handoffs and approval chains still make sense. Names and signatures in comments, free-text fields and attached files are de-identified as well. Because the logs describe employees' work, the partner's review of its right to license them can include employee notice and, in some jurisdictions, works council consultation.
Tell us what you are building
Describe the model or agent, the tasks it must handle, and the volume, format and permitted use you need. SourceX will match it to partner data.
Updated 3 October 2026.