Agent, workflow and domain-reasoning data
Writing an agent data specification: tasks, actions, observations, coverage and outcomes
Quick answer
An agent training data requirements specification tells a supplier which tasks you need, what the agent would have observed at each step, which actions it could take, what system state it acted on and how each episode ended. Set coverage as minimum episode counts per application, process variant and exception type rather than one total. Then say whether you want historical extracts, new capture or both, and state privacy tests for screenshots and typed text, the delivery format, the sample and the uses you need licensed.
By SourceX Editorial · Updated
What an agent spec adds to an ordinary data request
An agent spec keeps everything a normal supplier request contains (scope, format, rights, delivery) and adds sections that describe sequences rather than records: tasks, observations, actions, environment state and outcomes. Start from the general structure in SourceX's guide to writing a data request for suppliers or the data request builder for AI teams, and attach the agent sections below as a technical annex. In a formal tender, that annex becomes the requirements section of your agent data RFP, and the AI training data RFP template covers the commercial sections around it.
Each section has a natural owner on the buyer side and a predictable gap when nobody owns it.
| Spec section | Owner on the buyer side | Typical gap when left vague |
|---|---|---|
| Task inventory | Applied research lead | Instructions written by annotators after the fact, presented as user intent |
| Observation fields | Agent or modeling engineer | Screenshots downscaled with no record of the original resolution |
| Action space | Agent or modeling engineer | Free-text descriptions of what the worker did instead of typed actions with arguments |
| Environment context | Platform or evaluation engineer | No application version or start state, so episodes cannot be replayed |
| Outcome labels | Evaluation lead | "Ticket closed" used as success without checking for reopens |
| Coverage targets | Research lead with the product owner | One total episode count, successes only |
| Privacy acceptance | Privacy reviewer | Redaction tested on text fields but not on pixels or typed input |
| Permitted uses | Counsel | Training licensed; evaluation and environment building never mentioned |
For the kinds of agent data these sections describe and where each comes from, start at the agent training data hub.
Task inventory: instructions, start conditions and real frequency
The task inventory is a numbered list of the jobs the agent must learn, each with a natural-language instruction, the applications involved, the start condition and how often the task occurs in real operations. Ask for instructions at two levels: a goal-level instruction ("issue a partial refund for the damaged item on order 4471") and a step-level instruction for each action. Google DeepMind's AndroidControl dataset gives every task both high-level and low-level human-written instructions so researchers can study how much task complexity an agent can handle [1].
Record who wrote each instruction and when. An instruction typed by the real requester before the work started, such as an email, a ticket description or a chat message, is a different training signal from one an annotator wrote after watching a recording. Require an instruction_origin field so the two are never mixed silently. For historical business data, also ask for the observed monthly frequency of each task, so you can weight training toward the work that dominates your deployment.
Observation fields: what the agent would have seen
Observations are the inputs available at each step, and the spec must name each one with its format, because a supplier cannot add a missing modality after collection. For computer-use data, vendor requirements usually cover a lossless PNG screenshot with its pixel resolution and display scale factor, the accessibility tree or DOM snapshot, the active window title or URL, and a timestamp. For tool-use agents, the observation is the tool result payload plus any user message that arrived. For agents working inside business systems, it is the record as it stood before the step, not its final version.
Ask for raw structure rather than a supplier's condensed version. Mind2Web, built from more than 2,000 tasks on 137 real websites across 31 domains, found that filtering raw HTML with a small language model significantly improved both the effectiveness and the efficiency of LLM agents [2]. Filtering is therefore a modeling choice you want to control, so request the unfiltered DOM or accessibility tree alongside any reduced view. Step-level field requirements for screen data are set out in what every computer-use step record must contain.
Action space: declare it before collection starts
The action space is the closed list of actions an agent can take, each with typed arguments, and it belongs in the spec so that every supplier and annotator records the same thing. For screen-based agents, list the primitives (click, double-click, type, key combination, scroll, drag, wait, done and fail), state whether coordinates are absolute pixels or normalized, and require the target element's identifier next to the coordinates. For historical data from business systems, the action is usually a field-level change (record, field, old value, new value, user, time), which field-level audit trails can supply.
For tool-use agents, require the tool definitions themselves, versioned, alongside every call. In the Model Context Protocol (MCP) specification revision dated 2025-11-25, each tool has a unique name and an inputSchema that must be a valid JSON Schema object, defaulting to the 2020-12 dialect [3]. The same revision separates protocol errors, returned as JSON-RPC errors, from failures inside tool execution, returned in the result with isError: true [3]. Ask suppliers to keep both error channels, since they are the failure data that MCP tool-use data and error-recovery training depend on.
Environment context: versions, configuration and the policies in force
Environment context is everything about the system that decides whether an action succeeds: application name and version, operating system, locale, screen resolution, tenant configuration such as custom fields and approval rules, and the policies the worker had to follow. OSWorld pairs each of its 369 tasks with an initial-state setup and an execution-based evaluation [4]. τ-bench combines realistic databases and APIs with domain-specific policy documents, so an agent is judged on following the rules as well as on reaching the goal [5].
Ask for the same ingredients from real operations: a start-state snapshot or reset script per episode, the configuration export, and the version of each policy or standard operating procedure in force on the episode date. Where a supplier delivers process records rather than recordings, object-centric event logs in the OCEL 2.0 format can record how object attributes changed over time and how one event touches several objects [6], which preserves the state history the work acted on. For building replayable start states from real records, see seed data and state snapshots for enterprise agent sandboxes.
Outcome labels: define success the way you will grade it
Outcome labels say how each episode ended, and they should use the terms your evaluation will use: a checkable end state, not a narrative judgment. τ-bench compares the database state at the end of a conversation with an annotated goal state, and it introduced pass^k, the probability that an agent succeeds on all k independent trials of a task, to measure reliability [5]. Business records support the same kind of check: payment posted and not reversed, case resolved and not reopened, order shipped with no credit memo.
Specify the full label set, not just "success": completed, completed with correction, escalated to a human, abandoned, failed and policy violation. Name the source of each label (a system field, a downstream event or a reviewer) and the observation window for downstream events. Require that failed and escalated episodes are delivered rather than filtered out, because they carry the error and recovery steps. Deriving these labels from business records is covered in task success labels for agent trajectories.
Coverage targets: minimum counts per cell, not one total
Coverage targets set a minimum number of episodes for each combination of application, process variant and exception type, because a large total can hide empty cells. AndroidControl, which spans 833 Android apps, found that fine-tuned agents keep improving in domain as data grows, while out-of-domain performance scales significantly more slowly, especially for high-level tasks [1]. The applications and variants your agent will meet in production therefore need their own cells; extra episodes from other apps are a weak substitute.
Hold out whole cells for evaluation. Mind2Web reports separate test settings for unseen tasks, unseen websites and unseen domains [2], and a spec can do the same by reserving entire applications or customer tenants for the test split instead of sampling episodes at random. Exception cells (missing data, duplicate records, system errors, rejected approvals, timeouts) usually need deliberate sourcing; see exception handling records, and for setting the numbers, sizing agent training data.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Application | Process variant | Exception type | Min. episodes | Of which failed or escalated | Split |
|---|---|---|---|---|---|
| ERP, accounts payable | PO invoice, three-way match | None | 400 | 40 | Train |
| ERP, accounts payable | PO invoice, three-way match | Quantity mismatch | 120 | 30 | Train |
| ERP, accounts payable | Non-PO invoice, approval routing | Approver absent | 80 | 20 | Train |
| ERP, second tenant | PO invoice | Any | 150 | 30 | Test only |
If the agent will be part of a high-risk AI system under the EU AI Act, Article 10 requires data governance covering, among other things, data collection processes and the origin of the data, examination of possible biases and identification of data gaps [7]. As of October 2026, the Act has been amended by Regulation (EU) 2026/1744, and secondary sources report that Annex III high-risk obligations now apply from 2 December 2027 [8]. A coverage table with its known gaps written down is a practical start on that record; a coverage gap analysis compares it with your deployment distribution.
Historical extract, new capture or both: say which
The spec must say whether you want records from systems that already ran the work, new recordings of people doing it, or both, because the rights, cost and fidelity differ.
| Historical extract | New capture | Both, linked | |
|---|---|---|---|
| Typical source | Event logs, audit trails, ticket histories, API logs | Recording tool on workers' computers | Recordings of tasks that also leave system records |
| Observations | Record states; rarely screens | Screens, input events, accessibility data | Both, aligned by timestamp |
| Distribution | Real frequencies and real exceptions | Only the tasks you script | Real tasks with full observation |
| Main gap | Steps taken outside the system (calls, email, spreadsheets) | Task selection bias and staged behavior | Alignment effort |
| Who must agree | The data owner, within its existing permissions | The data owner, plus worker notice or consent where monitoring law requires it | Both |
| Formats | XES (IEEE 1849-2023) [9] or OCEL 2.0 [6] event logs; system exports | Episode folders or JSONL step files with images | Both, joined on episode IDs |
OpenCUA's AgentNet Tool shows what new capture involves: people install a cross-platform recorder on their own Windows, macOS or Ubuntu computers, and it records demonstrations together with the matching computer states [10]. Historical extracts need no recorder, but they show what changed in a system, not what was on screen. The trade-offs are compared in licensed workflow records, commissioned demonstrations or synthetic trajectories, and recording at work is covered in consent and monitoring law for computer-use demonstrations.
SourceX sources operational datasets from US companies on request, including support and sales histories, finance and legal workflows and new recordings of hands-on work. It looks for businesses that hold the data a buyer describes, every release is approved by the supplying company, and a request does not guarantee a matching dataset. You can submit an agent data specification to SourceX rather than approaching companies yourself.
Privacy acceptance tests for screens, typed text and notes
Privacy acceptance criteria state, for each data surface, what must be removed or replaced and how you will test the delivered sample, because agent data carries personal data in places a text-only redaction pass never inspects. List the surfaces separately:
- Pixels: names, account numbers and addresses in screenshots, including notification pop-ups, other open windows and browser tabs. Require detection on the rendered image, not only on the DOM.
- Typed input: the text argument of every type action and the arguments of every tool call, which hold what the worker entered.
- Structure and URLs: accessibility-tree labels, DOM attributes and URL query strings that embed customer IDs or email addresses.
- Free text: ticket notes, comments and chat messages inside observations.
- Consistency: the same person maps to the same replacement token in the screenshot, the DOM, the action argument and the outcome record, or the trajectory stops making sense.
Set a measurable threshold and a test. NIST SP 800-188 recommends adopting a de-identification standard with measurable performance levels and running re-identification studies to gauge residual risk, and it cautions that traditional de-identification has inherent limits [11]. A workable clause names the number of episodes you will inspect, the residual-identifier rate that triggers rejection and who performs the inspection. Where SourceX sources the data, personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded per dataset and a sample is checked after processing; no method is perfect, so keep your own test. Screen-specific techniques are in PII redaction for screen recordings and computer-use trajectories.
Delivery format, sample and acceptance tests
The delivery section fixes the file layout, the sample you see before signing, the tests a delivery must pass and how often new data arrives. A common layout is an episode manifest plus step records in JSON Lines, which requires UTF-8 without a byte order mark and one JSON value per line [12], with screenshots stored as image files referenced by path and checksum or packed into WebDataset tar shards, where files sharing a basename form one sample [13]. Ask for dataset-level metadata in Croissant, a schema.org-based JSON-LD vocabulary for datasets, files and record structure [14]. NeurIPS 2026 sets responsible-AI metadata requirements built on it for its Evaluations and Datasets Track [15].
Request a sample of complete episodes that touches every coverage cell, including failed and escalated ones, before you commit; evaluating an agent data sample lists what to check. Acceptance tests worth writing into the spec:
- Every step validates against the delivered JSON Schema, and timestamps increase within each episode.
- Each recorded action resolves to an element present in that step's observation.
- An agreed share of episodes replays from its start state to its labeled outcome.
- Outcome labels match an independent re-review of a sample at an agreed rate.
- The privacy inspection above passes.
Tie the refresh cadence to application releases, since an interface change can break the link between older screenshots and current element identifiers. General inspection mechanics are covered in acceptance criteria for licensed training data.
Uses to name before the supplier agrees
The spec should list every use you intend, because agent data is reused in more ways than most training data and a license covers only what it names. Uses to request explicitly:
- Training: supervised fine-tuning on demonstrations, reinforcement learning, and training reward models or judges on the outcome labels.
- Evaluation: a held-out split kept private. OpenAI stopped reporting SWE-bench Verified after stating that the benchmark had become increasingly contaminated [16], and a split that never leaves your environment reduces that risk for your own tests.
- Environment construction: building a working replica of an application and its data, as WebArena does with self-hosted websites in four domains [17].
- Derived synthetic tasks: generating new tasks or trajectories from the licensed episodes, and who owns them.
- After the term: what happens to trained models, derived tasks and built environments when the license ends.
Every dataset SourceX sources goes through rights review and is delivered under a license that defines which records are included, what they can be used for, how long the license runs and how delivery happens. Draft the use list with license terms for agent data and, for evaluation-only sets, the evaluation dataset specification guide. If the agent also needs paired audio or video, the multimodal specification template covers those fields.
Illustrative specification excerpt
The sections above, condensed into one machine-readable annex for a finance-operations agent, could read as follows.
Illustrative example: invented to show structure; it does not describe an available dataset.
spec_id: AP-AGENT-01
source_mode: historical_extract_plus_new_capture
tasks:
- task_id: AP-03
goal_instruction: "Post the attached supplier invoice and route it for approval"
step_instructions: required
instruction_origin: [requester_email, annotator_written] # labeled per episode
applications: [{name: "ERP, accounts payable module", version: required}]
start_state: snapshot_or_reset_script
monthly_frequency: required
observations:
screen: {format: png, store_resolution: true, store_scale_factor: true}
structure: {accessibility_tree: raw, dom: raw_if_web}
record_state: before_each_step
actions:
space: [click, double_click, type, key_combo, scroll, drag, wait, done, fail]
coordinates: absolute_pixels
target_element_id: required
tool_calls: {definitions: versioned, schema_dialect: "JSON Schema 2020-12", keep_error_results: true}
outcomes:
labels: [completed, completed_with_correction, escalated, abandoned, failed, policy_violation]
label_source: required_per_label
downstream_window_days: 30
coverage: see_coverage_table # minimum episodes per cell, failures included
privacy:
surfaces: [pixels, typed_text, tool_arguments, urls, accessibility_labels, free_text]
pseudonyms: consistent_across_surfaces
acceptance: {inspect_episodes: 200, max_residual_identifier_rate: agreed_in_writing}
delivery:
layout: [manifest.json, steps.jsonl, images/]
metadata: croissant
sample: every_cell_including_failures
refresh: on_major_application_release
permitted_uses: [training, private_evaluation, environment_construction, derived_synthetic_tasks]
Send SourceX your agent data specification
If you have drafted these sections, you can describe the agent data you need on the SourceX buyer page, submit it as a data request or talk to SourceX about it. SourceX looks for US companies that hold the records you describe, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Specify your workflow dataset.
Sources
- Li et al., Google DeepMind, "On the Effects of Data Scale on UI Control Agents" (2024). https://arxiv.org/abs/2406.03679
- Deng, Su et al., The Ohio State University, "Mind2Web: Towards a Generalist Agent for the Web" (2023). https://arxiv.org/abs/2306.06070v1
- Model Context Protocol, "Tools (specification revision 2025-11-25)" (2025). https://modelcontextprotocol.io/specification/2025-11-25/server/tools
- Xie et al., XLANG Lab, HKU and collaborators, "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
- Yao et al., Sierra, "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- OCEL standard authors, "OCEL (Object-Centric Event Log) 2.0 Specification" (2023). https://arxiv.org/pdf/2403.01975
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- EUR-Lex, "Regulation (EU) 2026/1744 (Digital Omnibus on AI) amending Regulation (EU) 2024/1689" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
- IEEE Standards Association, "IEEE 1849-2023: IEEE Standard for eXtensible Event Stream (XES) for Achieving Interoperability in Event Logs and Event Streams" (2023). https://standards.ieee.org/ieee/1849/10907
- Wang, Yu et al., "OpenCUA: Open Foundations for Computer-Use Agents" (2025). https://arxiv.org/abs/2508.09123
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- WebDataset project, "webdataset (GitHub repository)". https://github.com/webdataset/webdataset
- Akhtar et al., MLCommons Croissant working group, "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- NeurIPS, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- Zhou, Xu et al., Carnegie Mellon University, "WebArena: A Realistic Web Environment for Building Autonomous Agents" (2024). https://arxiv.org/abs/2307.13854v4
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.