Agent, workflow and domain-reasoning data
Computer-use trajectory data: what every step record must contain
Quick answer
A computer use agent trajectory data format has to let you rebuild what the agent saw and did at every step. Each step record needs a lossless screenshot with its pixel resolution and display scale factor, the accessibility tree or DOM where one exists, one action from a declared, versioned action space with normalized coordinates and a target element, and links to the observations before and after it. Each episode adds the instruction, start state, software versions and a success label that says how success was judged.
By SourceX Editorial · Updated
No shared standard yet: how current formats store a trajectory
As of October 2026, no single interchange standard exists for computer-use trajectories; benchmarks, open datasets and agent frameworks each define their own layout. Most share a core of instruction, ordered steps, a per-step screenshot and a structured action, but differ in what an observation holds, how actions are encoded, and whether the format serves training or replay. Which records help at all is covered under training data for computer-use agents, the computer-use data glossary entry and the AI agent training data hub.
| Format | Observation per step | Action encoding | Built for |
|---|---|---|---|
| OSWorld [1] | Screenshot, accessibility tree, or both | pyautogui code plus WAIT, FAIL and DONE | Execution-based evaluation in real computer environments |
| WebArena [2] | URL, open tabs, page as screenshot, HTML DOM or accessibility tree | Element-ID actions (click, type, hover), keys, scrolling, tab and URL navigation, stop with an answer | Self-hosted websites in four domains |
| Mind2Web [3] | Page HTML from real websites | Click, Type or Select Option on a page element | Crowdsourced action sequences on 137 websites |
| AgentNet-style rows on Hugging Face [4] | Screenshot per row | Numeric action code, normalized coordinates, click type, keyboard text, OS type, episode and step indices | Flat training rows |
| OpenCUA AgentNet [5] | Demonstrations with matching computer states | State-action pairs plus reflective chain-of-thought | Training data across Windows, macOS and Ubuntu |
| Cua framework logs [6] | Screenshot files in per-turn folders | Model API requests and responses, call results | Logging agent runs |
| ATIF, as documented in NVIDIA NeMo Agent Toolkit [7] | Not GUI-specific | Ordered steps under a root Trajectory with schema version, session and trajectory IDs, agent metadata | Interchange of agent runs |
Action vocabularies do not translate one-to-one: a WebArena element-ID click carries no coordinates, a pyautogui call carries no element, and an integer action code means nothing without the supplier's codebook. Formats are also still changing: STAMP, a paper on mobile GUI agents, bundles a screenshot, action, reasoning and memory targets into each step [8], and as of October 2026 the ATIF reference lists its version as ATIF-v1.7 [7]. Fix the target layout in your agent data specification before collection starts, and judge scale per application, not only in total: OpenCUA's AgentNet holds 22.6K trajectories across more than 100 applications and 200 websites [5].
Observation fields: the screen exactly as the agent saw it
An observation is everything available immediately before an action, and a usable one keeps the original display fidelity plus the geometry needed to map pixels to coordinates.
| Field | Why it matters | Common defect |
|---|---|---|
| Screenshot: lossless PNG at physical resolution, path and SHA-256 | Small text stays legible; grounding needs exact pixels | JPEG artifacts, downscaling, frames from compressed video |
| Capture times: UTC plus a monotonic millisecond counter | Orders observations and actions without clock drift | Captured after the action began, showing hover or pressed states |
| Display geometry: physical size, OS scale factor, monitor layout | Converts logged coordinates to image pixels | Logical-pixel coordinates against a physical-pixel image |
| Application and window: process name, app version, title, bounds | Coverage, splits, replay on the right build | Missing version, so replay breaks after updates |
| Accessibility tree: node ID, role, name, value, states, bounding box | Names the target; text without OCR | Absent in canvas apps and remote sessions; stale captures |
| Browser state: URL, tabs, DOM or accessibility snapshot | Web agents act on elements and URLs | DOM pruned before delivery |
| Focus and cursor: focused node, caret, selection, pointer | Disambiguates typing and shortcuts | Not captured |
On desktops the accessibility tree comes from the operating system's accessibility interface (UI Automation on Windows, the macOS accessibility API, AT-SPI on Linux), and its quality depends on how well each application implements it. Ask the supplier to report, per application, the share of actions whose target resolves to a tree node; where it is low, the model learns from pixels. Terminal emulators, thick clients and Citrix or RDP sessions often expose only pixels; see interaction data from legacy desktop, terminal and virtual-desktop applications.
Take the DOM unpruned. Mind2Web's authors reported that filtering raw page HTML with a small language model significantly improved the agent's effectiveness and efficiency [3]; that filter is a modeling decision, and a supplier who prunes before delivery has made it for you. Element-level grounding labels are a separate product, described in GUI grounding data from enterprise applications.
Action fields: one entry from a declared action space
Every action should be one entry from a closed, versioned vocabulary, with coordinates that survive resizing and links to its target element and the observations on either side. Action spaces differ by harness, as the WebArena and OSWorld rows above show [1][2], so the vocabulary and its version belong in the record.
- Type from the published vocabulary: click, double_click, right_click, drag, type, hotkey, scroll, wait, plus terminal done and fail, with the vocabulary version stored on the episode. An explicit fail action, as in OSWorld [1], gives infeasible tasks a correct label.
- Coordinates, twice: physical pixels of the stored screenshot and values normalized to 0-1. AgentNet-style rows store normalized coordinates [4], which survive resizing only if the reference frame is stated.
- Target element: node ID, role, accessible name and bounding box from the pre-action tree, or an explicit null when no tree exists.
- Keys and text: canonical key names, held modifiers, and typed text with typed redaction placeholders rather than blanks.
- Scroll and drag: direction, amount and unit (wheel notches and pixels differ by OS and application), drag start, end and duration.
- Timing and links: start and end on the monotonic clock, and the pre- and post-action observation IDs.
Worked example: one click on a high-DPI display
A demonstrator works on a 2880 × 1800 panel at 200% scaling, so the logical (scaled) screen is 1440 × 900. The recorder saves a 2880 × 1800 PNG but logs the click at logical (1158, 206); read as image pixels, that point lands above and left of the Hold button actually pressed at physical (2316, 412).
Normalizing against the logical screen gives (0.8042, 0.2289), which maps to (2316, 412) on the stored image and about (1029, 183) on a 1280 × 800 training resize. With the scale factor and coordinate space recorded, conversion is mechanical; without them, every high-DPI click is suspect.
Raw input events or semantic actions: deliver both, aligned
Ask for both layers, the raw event stream captured from the operating system and the semantic actions derived from it, joined by an alignment table, because the conversion is lossy and every recorder does it differently.
| Layer | Example | Strength | Failure mode |
|---|---|---|---|
| Raw input events | mouse and key down/up, wheel ticks, millisecond timestamps | Lossless; actions can be re-derived | Not what a model emits; keystroke-level personal data |
| Semantic actions | click, type "INV-88213", scroll down 3 | Matches the model's action space | Merge rules misfire on double-click timing, jittery drags, autocomplete and input-method composition |
| Element-referenced actions | click node n-377 | Survives layout and resolution changes | Needs a reliable tree; unusable on remote desktops |
| Code actions | a pyautogui click at (x, y) | Executable for replay | Tied to one library and screen geometry |
The alignment table needs one row per semantic action: its ID, the first and last raw event IDs it absorbed, the merge rule applied, and the observation chosen as its keyframe. OpenCUA converts recorded demonstrations into state-action pairs [5], and every such pipeline makes keyframe and merge choices that should be documented and re-runnable. The keyframe must precede the action's first raw event, because a screenshot taken mid-press shows a pressed button or open menu the model never sees at inference time. Recordings without event logs are a different problem, covered in turning screen recordings into action-labeled trajectories.
Episode metadata: instruction origin, start state and outcome
Episode metadata decides whether a trajectory can be replayed, graded or split correctly; what matters most is the instruction's origin, a reproducible start state, and a success label that names its judge.
- Instruction and its origin. Record whether the instruction came before the demonstration (task-first) or was written afterward (hindsight). OS-Genesis, for example, builds GUI agent trajectories by reverse task synthesis, deriving tasks from interactions after the fact [9]; hindsight instructions describe what happened, not always how users phrase requests.
- Start state. OSWorld pairs each task with an initial-state setup and an execution-based evaluation [1]. For business applications that means a VM or container image, seeded records, the logged-in role and open files, each hashed; see seed data and state snapshots for agent sandboxes.
- Environment. OS build, application and browser versions, locale, time zone and display geometry.
- Outcome. A success flag, the judge type (state-check script, human reviewer or demonstrator self-report) and its version, plus partial-progress checkpoints where defined. Keep failed and abandoned episodes; task success labels for agent trajectories covers label definitions and computer-use agent evaluation tasks covers verifiable end states.
- Provenance. Capture tool and version, a pseudonymous demonstrator ID with role, collection batch, and references to the notice or consent record and license schedule; capture at operating businesses is covered in collecting computer-use demonstrations at work.
Reasoning and intent notes: a separate, attributable layer
Store per-step reasoning, intent notes and memory targets as a separate annotation layer keyed to step IDs, with a field naming who or what wrote each note, so notes can be dropped, regenerated or audited without touching the raw log. Research formats already carry such fields: STAMP's steps include reasoning and memory targets [8], and OpenCUA adds reflective chain-of-thought to its state-action pairs [5].
Sources behave differently. Demonstrator think-aloud notes are contemporaneous but sparse; annotator notes written afterward can mention information not yet visible at that step, leaking the future into the training target. Model-generated reasoning should record the model, prompt version and date, since its permitted use may depend on that model's terms.
Illustrative episode header and step record
The episode header holds everything constant across an episode, and each step line references image and tree files instead of embedding them.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"episode_id": "ep-2026-000417",
"schema_version": "1.2",
"instruction": {"text": "Put invoice INV-88213 on hold for a price mismatch on line 2",
"source": "task_first", "language": "en-US"},
"environment": {"os": "Windows 11 23H2", "locale": "en-US", "time_zone": "America/Chicago",
"apps": [{"name": "erp-desktop-client", "version": "9.4.2"}],
"display": {"width_px": 2880, "height_px": 1800, "scale": 2.0, "monitors": 1}},
"start_state": {"snapshot_ref": "vm/ap-clerk-base-0912.qcow2", "sha256": "<hash>", "account_role": "ap_clerk"},
"action_space": {"name": "desktop-semantic", "version": "2", "coordinates": "physical_px + normalized_to_screenshot"},
"frame_policy": "PNG at physical resolution before every action and 500 ms after",
"outcome": {"success": true, "judge": "state_check", "judge_ref": "checks/invoice_hold_v3.py",
"checked_fields": ["invoice.status", "invoice.hold_reason"]},
"provenance": {"capture_tool": "<recorder and version>", "demonstrator": "dem-0193",
"notice_ref": "<notice or consent record>", "license_ref": "<license schedule>"}
}
{
"episode_id": "ep-2026-000417", "step": 6,
"obs_before": {"obs_id": "o-0011", "t_utc": "2026-09-12T15:04:21.318Z", "t_mono_ms": 48213,
"screenshot": "shards/000012.tar/ep-2026-000417_o-0011.png", "sha256": "<hash>",
"window": {"app": "erp-desktop-client", "title": "Invoice INV-88213", "bounds_px": [0, 0, 2880, 1760]},
"a11y_tree": "shards/000012.tar/ep-2026-000417_o-0011.json", "focused_node": "n-204"},
"action": {"type": "click", "button": "left", "clicks": 1,
"xy_px": [2316, 412], "xy_norm": [0.8042, 0.2289],
"target": {"node_id": "n-377", "role": "button", "name": "Hold", "bbox_px": [2260, 388, 2380, 436]},
"t_start_mono_ms": 49120, "t_end_mono_ms": 49188, "raw_events": ["e-1903", "e-1904"]},
"obs_after": {"obs_id": "o-0012", "t_mono_ms": 49690},
"annotations": {"intent": {"text": "Open the hold dialog", "source": "demonstrator_think_aloud"}}
}
Normalized coordinates divide the pixel position by the screenshot size (2316 / 2880, 412 / 1800), and the point falls inside the target's bounding box, the first thing an acceptance script checks. A typing step would carry a placeholder such as [VENDOR_NAME_1], and the same placeholder must appear wherever that value shows in screenshots, tree names and URLs; see PII redaction for screen recordings and agent trajectories.
Packaging that survives replay: JSONL, image shards and a manifest
Deliver step records as JSON Lines that reference screenshots stored as separate files, shard the images for streaming, and include a manifest with checksums plus a machine-readable dataset description.
- Step records in JSONL. Lines must be UTF-8 with no byte order mark, each one a valid JSON value [10]; put one step per line and episode headers in their own file.
- Screenshots as files, not base64. Inline base64 inflates image bytes by about a third; list each PNG's path and SHA-256 in a manifest instead.
- Shards for large sets. WebDataset stores samples in POSIX tar shards and groups files sharing a basename into one sample [11], so giving a screenshot and its tree the same basename (
ep-2026-000417_o-0011.pngandep-2026-000417_o-0011.json) keeps them in one sample. - A columnar step index. In Parquet, the file footer records where each column chunk starts, so readers fetch only the columns they need [12]; filtering by application or action type never touches an image.
- A dataset description. Croissant describes datasets in schema.org-based JSON-LD, covering files and record-level structure [13], and NeurIPS 2026 set responsible-AI metadata requirements for its Evaluations and Datasets Track built on Croissant RAI fields [14]. Add a frame-sampling policy: when screenshots were taken (before every action, on a timer, on screen change) and at what resolution.
- A replay format, not a summary. Letta's July 2026 coding-agent trajectory package strips detail for token efficiency, and its authors contrast it with replay-oriented formats such as ATIF [15]; a compact log built for agents rereading past sessions does not replace a full step record.
Transfer and access choices are covered in the dataset delivery formats guide.
Acceptance checks for a trajectory delivery
Run mechanical checks on every file and replay checks on a random sample before acceptance, because geometry, timing and redaction defects otherwise surface only in training.
- Every line validates against the delivered JSON Schema, with no action type outside the published vocabulary.
- Every step carries display geometry and a declared coordinate space; normalized coordinates sit within 0-1.
- Overlay test: plot sampled pointer actions on their pre-action screenshots; each must land inside the target's bounding box.
- Target node IDs exist in the pre-action tree; unresolved targets are reported per application.
- Times run monotonically from pre- to post-action observation; step indices are contiguous.
- Re-running the merge rules on raw events reproduces the alignment table.
- Redaction placeholders match across typed text, screenshots, tree names and URLs.
- Every episode names its outcome judge; failed and infeasible episodes, and counts by application, version, OS and action type, match the specification.
- Evaluation splits hold out whole tasks, websites or applications, not random steps; Mind2Web, for instance, reports cross-task, cross-website and cross-domain test settings [3].
- Checksums match the manifest, and sampled episodes replay from their start state where one was delivered.
Sample planning is covered in evaluating an agent data sample before you buy.
When trajectories come from licensed business work
Business screen recordings and activity logs seldom carry accessibility trees, raw events and success labels together, so decide which fields are mandatory and which you will derive. See licensing screen recordings for AI training, workflow and screen activity data, a comparison of licensed, commissioned or synthetic trajectories, and license terms for agent data for replay and derived-task grants.
SourceX sources operational datasets from US companies, including new recordings of hands-on work, and manages the licensing agreements. Datasets are sourced on request, not held in stock; each goes through rights review and is delivered under a license defining which records are included and what they can be used for. Personal details such as names, emails and account numbers are removed or replaced before delivery, with the method recorded, though no de-identification method is perfect. Teams can send SourceX their step-record specification rather than approaching companies themselves.
Sourcing computer-use trajectories to your record spec?
Describe the applications, tasks, observation and action fields, outcome labels and volumes you need, and the uses you need licensed. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees, and a request does not guarantee a matching dataset. Specify your computer-use dataset with SourceX.
Sources
- Xie et al. (XLANG Lab, HKU and collaborators), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
- Zhou, Xu et al. (Carnegie Mellon University), "WebArena: A Realistic Web Environment for Building Autonomous Agents" (arXiv v4, 2024). https://arxiv.org/abs/2307.13854v4
- Deng, Su et al. (The Ohio State University), "Mind2Web: Towards a Generalist Agent for the Web" (2023). https://arxiv.org/abs/2306.06070v1
- TESS-Computer (Hugging Face dataset card), "tess-agentnet README" (accessed 2026). https://huggingface.co/datasets/TESS-Computer/tess-agentnet/blob/main/README.md
- Wang, Yu et al., "OpenCUA: Open Foundations for Computer-Use Agents" (2025). https://arxiv.org/abs/2508.09123
- Cua (framework documentation), "Trajectories" (accessed 2026). https://cua.ai/docs/cua/guide/advanced/trajectories
- NVIDIA, "NeMo Agent Toolkit 1.8: ATIF trajectory API reference" (accessed 2026). https://docs.nvidia.com/nemo/agent-toolkit/1.8/api/nat/atif/trajectory/index.html
- arXiv:2605.29324, "STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments" (2026). https://arxiv.org/pdf/2605.29324
- arXiv:2412.19723, "OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis" (2024). https://arxiv.org/pdf/2412.19723
- jsonlines.org, "JSON Lines" (accessed 2026). https://jsonlines.org/
- WebDataset project, "webdataset (GitHub repository)" (accessed 2026). https://github.com/webdataset/webdataset
- The Apache Software Foundation, "Parquet: File Format" (accessed 2026). https://parquet.apache.org/docs/file-format/
- Akhtar et al. (MLCommons), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- NeurIPS blog, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
- Letta (vendor blog), "Trajectory" (2026). https://www.letta.com/blog/trajectory/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.