Skip to content

Agent, workflow and domain-reasoning data

AI agent training data: the workflow, action and decision data agents learn from

Quick answer

AI agent training data is the record of real work that an agent learns to repeat or is graded against. It comes in six kinds: screens with labeled elements, GUI and tool-call trajectories, process event logs, decisions with their rationale, errors with their fixes, and seed state for practice environments. Public benchmarks cover websites and desktop apps; enterprise systems of record, approval waits and exceptions are the gap. Filling it usually means rebuilding trajectories from business logs and licensing them carefully.

By SourceX Editorial · Updated

Definitions live in the glossary entries for agent trajectory, computer-use data and tool-use data; why AI agents need business data covers the motivation. For data types outside agent work, start from the AI data sourcing guide for buyers.

Six kinds of agent data, sorted by what they teach

Agent data is easiest to scope by the skill it teaches: grounding, action, process, judgment, recovery or environment. Each kind comes from different systems in a different record unit, and most programs need several.

KindWhat it teachesWhere it originatesRecord unitGo deeper
GroundingLocating on-screen elements in the target appsScreenshots paired with accessibility trees, DOM snapshots or element boxesOne screen with labeled elementsGUI grounding data from enterprise applications
ActionThe next click, keystroke or tool call toward a goalScreen recordings with input events, browser action logs, API call logsStep: observation, action, resultComputer-use step records, API call logs as tool-use data
ProcessStep order across systems, handoffs and waiting timeEvent logs in XES [1] or OCEL 2.0 [2], ticket and case histories, field-level audit trailsCase or object-centric event logProcess mining event logs, ticket histories as trajectories
JudgmentWhich option to take under policy, and whyApprovals, rejections, underwriting and claims decisions with reason codes and notesInputs at decision time, choice, rationale, later outcomeDecision records with rationale, approval and rejection records
RecoveryNoticing an error and repairing itReopened cases, reversals, credit memos, exception queuesFailure plus corrective stepsException handling records, rework and reversals
EnvironmentA realistic world to act in and a test of successSeed records, state snapshots, app configuration, user-simulator dataInitial state plus success criteriaSeed data for agent sandboxes, user simulator data

A tool-calling agent behind an API can skip grounding; a computer-use agent for claims intake may need all six. SourceX's capability pages list business records for computer-use agents, tool use and function calling, long-horizon task agents, agentic workflow planning, IT operations agents and email and calendar agents.

Train, build or test: three jobs for the same records

The same business records can train a policy, furnish the environment and grader that policy practices in, or evaluate it on held-out tasks. Each job needs different fields, so fix the job before writing the specification.

Training a policy. Supervised fine-tuning on demonstrations, and reinforcement learning (RL) warm-starts, need step-level observation-action pairs plus a stated goal. Mind2Web is a public example: open-ended tasks paired with crowdsourced action sequences on real websites [3]. Ask for coverage across task variants, not hundreds of repeats of one path; see data for post-training teams.

Building environments and graders. RL and repeatable evaluation need a world the agent can act in and an automatic check of the result. τ-bench builds that world from databases and APIs, domain policy documents and simulated user scenarios, grades the database state at the end of an episode, and reports pass^k, the probability of succeeding on all k repeated trials of a task [4]. OSWorld supports task setup and execution-based evaluation on real operating systems [5].

Business data supplies the realistic parts: seed records, app configuration, the policy in force and the end state that counts as done. See configuration data for enterprise app replicas and the RL environment entry.

Evaluating on held-out tasks. Test tasks must stay out of every training set, including the model vendor's. In a February 2026 post, OpenAI said SWE-bench Verified had become increasingly contaminated and that it had stopped reporting scores on it [6]. SWE-Bench Pro keeps a held-out set private and adds a commercial set from startup codebases whose code is not released [7]: private company data used to keep part of an agent benchmark out of public training corpora. Agent evaluation task suites and human baseline data cover design.

Assign each case to training, environment or test by case ID and time period before anyone uses it. Splitting at the step level leaks, because steps from one case land on both sides.

What public agent benchmarks cover, and what enterprise work adds

Public agent benchmarks cover consumer websites, self-hosted web apps, desktop operating systems and simulated service desks, but say little about configured back-office systems, approval waits measured in days, or the exceptions enterprise agents meet. Use them as baselines and task templates, and source data for the gaps.

BenchmarkWhat it coversWhat enterprise buyers still have to source
Mind2Web [3]Over 2,000 tasks from 137 real websites in 31 domains, with crowdsourced action sequencesLogged-in internal applications; tasks spanning several sessions
WebArena [8]Self-hosted sites for e-commerce, forums, collaborative software development and content managementERP, claims, ITSM and other systems of record; role-based permissions
OSWorld [5]369 tasks in the 2024 release, with real desktop and web apps, file I/O and multi-app workflows, in an environment supporting Ubuntu, Windows and macOSCustomized enterprise clients, virtual desktops, terminal applications
τ-bench [4]Simulated retail and airline domains with APIs, written policies and simulated usersReal policy versions over time; the real mix of exceptions
Spider 2.0 [9]632 data-workflow problems on databases from real applications, often over 1,000 columns, on systems such as BigQuery and SnowflakeThe business request behind a query and the action taken after it

Benchmarks are revised, so scores from different releases are not comparable, and public tasks leak into training data, as the SWE-bench Verified case shows [6]. Computer-use evaluation tasks with verifiable end states covers building your own, and legacy desktop and terminal interaction data covers interfaces no benchmark models.

Why business records rarely arrive as clean trajectories

Business systems log state changes, not the steps a person took to cause them. A trajectory has to be reconstructed by joining audit trails, event logs, tickets and messages on shared object IDs and ordering them by time.

An ERP audit trail shows that an invoice's payment block changed at 14:07 and who changed it, not the screens checked first, the value copied from an email or the abandoned attempt. Clock skew between systems reorders steps, and scheduled jobs and RPA bots write under service accounts that look like people. Field-level audit trails, cross-system workflow records and RPA bot logs cover each source.

Process-mining formats help. IEEE 1849-2023 (XES) defines an XML format for event logs and event streams and supersedes the 2016 edition [1]. OCEL 2.0 links events to several objects, records changes to object attributes over time and qualifies object-to-object relationships, with SQLite, XML and JSON exchange formats [2]. That object-centric model fits back-office work, where one order-to-cash case touches an order, deliveries, invoices and payments; see procure-to-pay and order-to-cash records.

Require every step to state its evidence: observed in an audit trail, observed on screen, or inferred. Inferred steps help process supervision but are weak action labels. SourceX's workflow task histories page describes records rebuilt from system audit logs.

Illustrative example: invented to show structure; it does not describe an available dataset.

task:
  task_id: ap-exc-0447
  family: invoice_price_variance_resolution
  goal: "Release an invoice blocked by a purchase-order price variance"
  systems: [erp_accounts_payable, erp_purchasing, email, itsm_queue]
  outcome: {status: posted_after_po_change, label_source: erp_document_status, reopened_within_30d: false}
  timing: {active_minutes: 38, wall_clock_hours: 52}   # most of the elapsed time is waiting for the buyer
steps:   # steps 1-2 (invoice receipt and automatic match) omitted
  - {n: 3, ts: "2025-03-12T14:07:55Z", actor_role: ap_specialist, system: erp_accounts_payable,
     action: {type: field_update, object: invoice_4471, field: payment_block, from: none, to: price_variance},
     evidence: audit_trail}
  - {n: 4, ts: "2025-03-12T14:11:20Z", actor_role: ap_specialist, system: email,
     action: {type: send_message, to_role: purchasing_buyer, template: price_query},
     evidence: mail_log}
  - {n: 5, ts: "2025-03-14T09:40:02Z", actor_role: purchasing_buyer, system: erp_purchasing,
     action: {type: field_update, object: po_88213_line_2, field: net_price, from: 41.20, to: 43.05},
     rationale: "Supplier increase accepted under contract amendment", evidence: audit_trail}
  - {n: 6, ts: "2025-03-14T09:52:10Z", actor_role: ap_specialist, system: erp_accounts_payable,
     action: {type: release_block, object: invoice_4471},
     evidence: inferred_from_status_change}   # no screen capture exists for this step
privacy: {people: role_pseudonyms_consistent_across_systems, supplier_names: tokenized}

Rights and privacy: what a screen or log reveals

Agent data captures more parties than most datasets: customers in the records, employees doing the work, vendors whose software is on screen and, in deployment logs, your own users. Each needs a rights answer before delivery.

What the data capturesThe question to settleGo deeper
Customers or patientsIdentifiers sit in pixels, window titles and autocomplete lists, not only in fields. Health data needs HIPAA de-identification: Safe Harbor removal of 18 identifiers, or an expert's determination that identification risk is very small [10]PII redaction for screen recordings and trajectories
Employees doing the workScreen, keystroke and call capture at work raises notice and consent questions that vary by state; California prohibits recording confidential communications without all parties' consent [11]. Keystroke logs can also capture passwordsCollecting computer-use demonstrations at work
Vendors' software on screenScreenshots and environment replicas reproduce third-party interfaces whose license terms may restrict reuseReplicating third-party software, third-party content in screen recordings
Users of your deployed agentFTC staff warned in 2024 that model-as-a-service companies may be liable if they break promises not to use customer data for training [12]Training on agent deployment logs

Documentation duties follow some agents. From 1 January 2027, Colorado's SB26-189, signed in May 2026, requires developers of automated decision-making technology that materially influences consequential decisions, such as lending, insurance, employment or health care, to give deployers documentation that includes the categories of training data [13]. As of October 2026, the Attorney General's implementing rules are an interim draft [14].

In the EU, providers of general-purpose AI models must publish a sufficiently detailed summary of the content used for training under Article 53(1)(d) of the AI Act [15]. The compliance hub and license terms for agent data go further.

For the data it sources, SourceX runs a rights review that checks the business owns or may share the records and that required consents are in place. Names, emails, phone numbers and account numbers are removed or replaced before delivery and the method is recorded, though no de-identification method is perfect; health records must meet HIPAA Safe Harbor or Expert Determination before they are considered for a license. This section is general information, not legal advice.

Scoping your first agent data request

A workable request names the task family, the systems, the observation and action format, how success is labeled, and every use the license must cover.

  1. Task family and goal: for example, "release an invoice blocked by a price variance," with in-scope and out-of-scope variants.
  2. Systems and surfaces: application, version, customization level, and whether it runs in a browser, desktop client, virtual desktop or terminal.
  3. Observations: pixels, accessibility tree or DOM, API request and response payloads, or event log only.
  4. Actions: clicks and keystrokes, tool calls with arguments, or business-level actions such as approve or reassign.
  5. Outcome label and its source: document status, reopen flag or QA review; see task success labels.
  6. Horizon: steps per task, active versus wall-clock time, and handoffs; see long-horizon task records.
  7. Exceptions: a minimum share of cases off the happy path.
  8. Privacy: pseudonyms consistent across steps and systems, so the agent can follow one customer through a case.
  9. License uses: training, environment replication, derived tasks and benchmark publication, plus split rules.
  10. Delivery: JSONL with one episode or step per line (UTF-8, one JSON value per line [16]), OCEL or XES for event logs, and a datasheet covering motivation, composition and collection process [17].

Writing an agent data specification expands each item, evaluating an agent data sample covers acceptance, and cost drivers for trajectory data explains pricing variables. For the sourcing route, compare licensed, commissioned or synthetic trajectories.

SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases. The kinds of data it sources include support and sales histories, engineering records, documents, finance and legal workflows, and new recordings of hands-on work. Datasets are sourced on request rather than held in stock, so a request does not guarantee a match, and every release is approved by the supplying company.

SourceX does not train AI models. To start, describe the workflow records your agent needs.

Agent data guides by interface and decision type

Choose the guide that matches how your agent touches software and what it decides.

Mistakes that waste an agent data budget

Each of these gaps is visible in the specification before signing and costs more to fix after delivery.

  • Screen video with no input events, task boundaries or outcomes. It is raw material, not a trajectory, and inferring actions from pixels afterwards is harder than logging input events at capture.
  • "Closed" taken as success. Auto-close rules, reopened cases and later reversals make status fields poor labels unless checked against what happened next.
  • Exceptions filtered out for clean demonstrations. The rework and edge cases removed are the ones agents fail on in production.
  • A training-only license. Building an environment replica, deriving tasks or publishing a benchmark may fall outside a grant that names only training, so name each use.

Building agents that need real work histories?

Describe the task family, the systems involved, the outcome labels you need and the uses the license must cover. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Specify your workflow dataset.

Guides in this section

Sources

  1. IEEE Standards Association, "IEEE 1849-2023 - IEEE Standard for eXtensible Event Stream (XES) for Achieving Interoperability in Event Logs and Event Streams" (2023). https://standards.ieee.org/ieee/1849/10907
  2. OCEL standard authors (ocel-standard.org), arXiv:2403.01975, "OCEL (Object-Centric Event Log) 2.0 Specification" (2023). https://arxiv.org/pdf/2403.01975
  3. Deng, Su et al., The Ohio State University (arXiv:2306.06070; NeurIPS 2023), "Mind2Web: Towards a Generalist Agent for the Web" (2023). https://arxiv.org/abs/2306.06070v1
  4. Yao et al., Sierra (arXiv:2406.12045), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  5. Xie et al., XLANG Lab, HKU and collaborators (arXiv:2404.07972 v2; NeurIPS 2024), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
  6. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  7. Deng et al., Scale AI (arXiv:2509.16941), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
  8. Zhou, Xu et al., Carnegie Mellon University (arXiv:2307.13854 v4), "WebArena: A Realistic Web Environment for Building Autonomous Agents" (2024). https://arxiv.org/abs/2307.13854v4
  9. Lei et al., XLANG Lab, HKU and collaborators (arXiv:2411.07763; ICLR 2025), "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows" (2024). https://www.arxiv.org/pdf/2411.07763
  10. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  11. California Legislature (California Legislative Information), "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632
  12. Federal Trade Commission, Office of Technology (Tech@FTC staff blog), "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  13. Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
  14. Colorado Attorney General, "Colorado Automated Decision-Making Technology & Chatbot Safety Rulemaking" (2026). https://coag.gov/ai/
  15. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  16. jsonlines.org, "JSON Lines". https://jsonlines.org/
  17. Gebru et al. (arXiv:1803.09010; Communications of the ACM 2021), "Datasheets for Datasets" (2021). https://arxiv.org/pdf/1803.09010

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data