Agent, workflow and domain-reasoning data
Browser action logs for web agents: DOM events, page states and outcomes
Quick answer
A browser action log dataset for web agents pairs every user action (click, type, select, scroll, navigate) with the page state the user saw just before it, a resolvable element target, the resulting URL and network calls, and a task-level outcome. Analytics clickstream records that a page was viewed or an event fired, not which element was acted on or what the page looked like, so it rarely supports imitation learning or replayable evaluation. Source instrumented browser logs, or session-replay exports with known masking, from the internal applications your agent must operate.
By SourceX Editorial · Updated
Why analytics clickstream is not web agent training data
Clickstream from tools like product analytics or web server logs answers "what happened in aggregate," while an agent needs "given this screen, which element did a competent person act on, and why." A page-view table with session_id, url, timestamp and a named custom event has no pre-action observation, no element locator and usually no typed values, so a model cannot learn a policy from it.
The gap shows up in three concrete failure modes:
- No observation. Without a DOM or accessibility snapshot taken before the action, you cannot build an (observation, action) pair, which is the basic unit in Mind2Web-style training [2].
- No grounding. An event named
submit_invoice_clickeddoes not tell the model which of four buttons on the page it was, or how to find it after a UI release. - No outcome. Funnels show drop-off, but not whether the task (approve this invoice, update this ticket) succeeded, which is what WebArena-style evaluation scores [4].
Clickstream is still useful as a sampling frame: it tells you which workflows are frequent enough to collect deeply. If what you actually need is case-level sequences across systems rather than screen-level actions, a process mining event log is the better fit.
The fields that make a browser log trainable
A trainable record is one step of a task, with the state before the action, the action itself and the observable effect after it. Public datasets converge on this shape: a community-published Mind2Web human-demonstration subset on Hugging Face ships golden action sequences, checkpoints, clicks, typing, scrolling, DOM states and screenshots [3], and WebChain aligns visual, structural and action data across 31,725 human trajectories and 318k steps on real websites [1].
Request these fields at minimum:
| Field group | What to ask for | Why it matters |
|---|---|---|
| Task and session | Task ID, natural-language goal or ticket reference, session ID, step index, actor role | Links steps into trajectories and gives a conditioning goal |
| Pre-action observation | Serialized DOM or accessibility tree, viewport screenshot, scroll position, focused element | The model's input; without it there is no training pair |
| Action | Event type (click, input, change, keydown, scroll, navigate), coordinates, typed value or value class | The supervised target |
| Element target | CSS or XPath selector plus stable attributes (id, name, aria-label, role, data-testid, visible text) and a node ID in the snapshot | Lets you re-ground actions after UI changes |
| Effect | Resulting URL, DOM diff or post-action snapshot, triggered XHR/fetch requests with method, path and status | Detects silent failures and gives tool-use signals |
| Outcome | Task success flag, final record state, validation errors, human override or abandonment | Needed for filtering, reward and evaluation |
Network calls are often the most underused field. Recording the request triggered by a click (for example POST /api/invoices/{id}/approve returning 409) gives you a ground-truth success signal and a bridge to API call logs as tool-use training data. HAR files are the familiar JSON export for HTTP transactions from browser developer tools, but exports differ by tool and can include cookies and auth headers, so agree a concrete schema and field list with the supplier rather than assuming "HAR" settles it.
DOM, accessibility tree or pixels
Ask for both structural and visual observations, because each covers the other's failure modes. The UGround authors argue that HTML and accessibility-tree inputs can be noisy and incomplete, which pushes some teams toward pixel-level grounding [6]; structural snapshots, in turn, give exact element identity and text that screenshots alone lose. For internal applications built on component libraries, shadow DOM, canvas grids and cross-origin iframes are where structural capture breaks first, so ask suppliers which of those their capture missed.
If your agent will run on legacy desktop or virtual-desktop clients rather than in a browser, browser logs are the wrong source; see interaction data from legacy desktop and terminal applications and task mining data. For pixel-first computer-use agents, see the computer-use agent training data overview.
Session-replay exports: close, but check what was masked
Session-replay tools capture DOM mutations, inputs and pointer events, so an export can approximate a browser action log, but many tools mask input values by default. Replay SDKs built on or modeled after rrweb typically expose a setting such as maskAllInputs, and its default differs by vendor and version. That protects privacy at capture time, and it also means typed values, which are often the hardest part of a form-filling policy, may simply not exist.
Before scoping a purchase from replay data, confirm four things with the supplier:
- Which SDK and version recorded the sessions, and what masking and blocking settings were active during the collection window.
- Whether masked inputs kept a value class (date, currency, free text) or length, which is often enough for training.
- Whether replays are full snapshots plus mutations or sampled, and whether network requests were captured at all.
- Whether sessions can be joined to a task outcome in a system of record (ticket closed, invoice approved).
Privacy and pseudonymization in internal web app logs
Internal web logs leak identity through places reviewers forget: URLs, page titles and DOM text, not just form fields. A path like /acme-tenant/customers/48213/edit exposes a tenant name and a record ID, and page titles often repeat customer names. Plan consistent pseudonymization of hostnames, tenant slugs, record IDs in paths and query strings, and visible text in snapshots, and keep the mapping stable within a trajectory so navigation still makes sense.
Screenshots need separate treatment from DOM text, since text-level redaction does not touch pixels. The detailed playbook is in PII redaction for screen recordings and agent trajectories. Also ask for documentation you will need later: as of October 2026, California's AB 2013 requires developers of generative AI systems made available to Californians to post documentation about their training data [9], and a machine-readable datasheet such as Croissant-RAI makes provenance and preparation steps easier to carry forward [8].
How licensed internal logs compare with public web agent datasets
Public datasets are strong for public websites and benchmarks; licensed operational logs matter when your agent must work inside line-of-business applications that no benchmark covers. Mind2Web spans over 2,000 tasks on 137 real websites across 31 domains [2], WebArena provides self-hosted sites in four domains where the best GPT-4 agent in v4 reached 14.41% task success against 78.24% for humans [4], and Go-Browse and OpenCUA show synthetic exploration and tooling-driven collection at scale [5][7].
What none of them contain is your users' real procedures: approval screens with custom validation, admin consoles, multi-tab lookups and the exceptions people work around. That long tail is why teams pair public pretraining with a smaller, deeper set of logs from systems like their own. If you plan to replay tasks, you will also need matching application state, covered in seed data and state snapshots for agent sandboxes. When the target is an internal application category, you can describe the data to SourceX, which looks for US businesses that hold it; you describe the data, not the businesses.
A request template for browser action logs
Write the request around applications, tasks and fields, not around named companies. The template below follows the structure in writing an agent data specification.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"request": "Browser action logs from a web-based accounts payable application",
"tasks": ["approve invoice", "reject with reason code", "edit vendor bank details"],
"volume_target": "trajectories per task, with abandoned and failed attempts kept",
"observation": {"dom_snapshot": "pre-action, serialized", "screenshot": "viewport PNG", "a11y_tree": "preferred"},
"action_fields": ["event_type", "selector", "stable_attrs", "node_id", "value_or_value_class", "timestamp_ms"],
"effect_fields": ["url_after", "xhr_requests[method,path,status]", "dom_diff"],
"outcome_fields": ["task_success", "final_record_status", "validation_errors", "human_override"],
"privacy": {"pseudonymize": ["hostname", "tenant_slug", "record_ids_in_url", "names_in_dom"], "screenshots": "redacted, method recorded"},
"documentation": ["capture SDK and masking settings", "collection window", "app version changes"]
}
A matching single step might look like this, with pseudonymized identifiers:
Illustrative example: invented to show structure; it does not describe an available dataset.
{"task_id": "T-0192", "step": 4, "event_type": "click",
"selector": "button[data-testid='approve-btn']", "stable_attrs": {"role": "button", "aria-label": "Approve"},
"url_before": "https://app.tenant-a.example/invoices/INV-7F3A", "url_after": "https://app.tenant-a.example/invoices/INV-7F3A",
"xhr": [{"method": "POST", "path": "/api/invoices/INV-7F3A/approve", "status": 200}],
"outcome": {"task_success": true, "final_record_status": "approved"}}
Sourcing browser action logs for your web agent
SourceX sources operational datasets, including engineering records and new recordings of hands-on work, from US companies on request; nothing is held in stock and a request does not guarantee a match. Every dataset is rights-reviewed, personal details are removed or replaced before delivery with the method recorded, and delivery happens under a license defining records, uses, term and delivery. Start with the agent training data hub, see workflow and screen activity data licensing, then describe the browser action logs you need.
Frequently asked questions
How many trajectories do we need from an internal application?
There is no published threshold for internal applications; scale benchmarks like WebChain's 31,725 trajectories [1] describe public-web coverage, not one app. Teams usually size by task count and variance: enough attempts per task to include errors, edge cases and multiple valid paths.
Can we use logs captured by an existing session-replay tool?
Often yes, if the capture settings are known and network calls or outcomes can be joined. Check masking settings first, because input values may never have been recorded.
Should failed and abandoned sessions be included?
Yes. Failures with validation errors or human overrides teach recovery and give negative examples for evaluation, and filtering to success only hides exactly the cases agents struggle with.
Sources
- CVPR 2026, "WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces" (2026). https://cvpr.thecvf.com/virtual/2026/poster/39540
- The Ohio State University (arXiv:2306.06070; NeurIPS 2023), "Mind2Web: Towards a Generalist Agent for the Web" (2023). https://arxiv.org/abs/2306.06070v1
- Hugging Face (josancamon), "mind2web-subset-human dataset card". https://huggingface.co/datasets/josancamon/mind2web-subset-human/blob/main/README.md
- Carnegie Mellon University (arXiv:2307.13854v4), "WebArena: A Realistic Web Environment for Building Autonomous Agents" (2024). https://arxiv.org/abs/2307.13854v4
- Hugging Face (CMU neulab), "Go-Browse WebArena trajectories (agent-data-collection)". https://huggingface.co/datasets/neulab/agent-data-collection/blob/main/go-browse-wa/README.md
- The Ohio State University and Orby AI (arXiv:2410.05243; ICLR 2025), "Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents" (2024). https://arxiv.org/pdf/2410.05243
- arXiv:2508.09123, "OpenCUA: Open Foundations for Computer-Use Agents" (2025). https://arxiv.org/abs/2508.09123
- MLCommons Croissant RAI task force (arXiv:2407.16883), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.