Skip to content

Evaluation and benchmarking datasets

Computer-use agent evaluation tasks: real software workflows with verifiable end states

Quick answer

Computer-use agent evaluation works best when each task starts from a scripted application state, gives the agent a natural-language goal drawn from real work, and grades the final state of the software (records, files, settings) rather than the clicks taken. Public benchmarks such as OSWorld and WebArena demonstrate the method, but enterprise teams need private tasks from the apps their users run, enough tasks per app for stable estimates, and a maintenance plan for UI drift.

By SourceX Editorial · Updated

Why public GUI agent benchmarks are not enough for enterprise apps

Public benchmarks establish the grading pattern, but they rarely cover the line-of-business software your agent will actually operate. OSWorld runs real desktop environments (Ubuntu, with Windows and macOS support) and its 369 tasks each pair an initial-state setup with an execution-based evaluation script [1]. WebArena self-hosts working sites in four domains (e-commerce, forums, collaborative software development and content management) and reported 14.41% end-to-end success for its best GPT-4 agent against 78.24% for humans [2]. Mind2Web widened coverage to over 2,000 tasks on 137 real websites, but grades against crowdsourced action sequences rather than final state [3].

None of these exercise a mid-market ERP's three-way match screen, a ticketing tool's custom fields, or a claims system's approval queue. Business-scenario benchmarks built on synthetic CRM data, such as CRMArena-Pro, show that enterprise workflows remain difficult for agents [7], yet synthetic records lack the messiness of real tenants. Public suites also leak: as of October 2026, OpenAI has stopped reporting SWE-bench Verified because contamination made score gains increasingly reflect training exposure [9]. The cluster overview at LLM evaluation datasets for buyers and private evaluation sets vs public benchmarks cover that trade-off in depth.

End-state verification beats trajectory matching for most tasks

Grade the state the agent leaves behind, not the path it took, because valid workflows have many correct click sequences. tau-bench formalized this for tool agents: after the episode, the database is compared with an annotated goal state, and reliability is measured across repeated trials with pass^k [4]. The same logic transfers to GUI agents (an editorial inference): if the invoice record shows status approved, amount 1,240.00 and GL code 6100, it matters little whether the agent used the search bar or the filtered list view.

Trajectory matching still has a role. Use it for safety constraints (the agent must not open the payroll module), for efficiency budgets (step counts, wall-clock time), and for diagnosing where failures happen. For grounding-specific errors, such as clicking the wrong icon on a dense 4K screen, a dedicated grounding benchmark like ScreenSpot-Pro isolates the perception problem from planning [6]. Treat these as secondary metrics layered on a state-based pass/fail.

State checks fail in predictable ways. Common failure modes include:

  • Under-specified checks: the checker verifies a record exists but not its field values, so a half-completed form passes.
  • Over-specified checks: the checker requires an exact timestamp or auto-generated ID, so correct runs fail.
  • Collateral damage blind spots: the target record is correct, but the agent also deleted a draft or changed a user setting; add negative assertions on untouched objects.
  • UI-only state: success lives in an unsaved modal or local cache; verify via API or database read after the session ends.

Read agent evaluation task suites and state-based grading for the general framework and agent safety tests for irreversible actions for negative assertions.

Where real task sources come from

The strongest task sources are workflows people already perform in the software, observed in licensed screen activity, ticket histories and system audit logs. Our working hypothesis is that screen activity data shows which multi-step, multi-app sequences actually occur and how often, which is what an eval distribution should mirror. Tools such as the AgentNet capture tool in OpenCUA show how human computer-use demonstrations can be recorded at scale with screenshots and actions [8]; the same recordings can seed eval goals once the end state is reconstructed.

Practical task sources by type:

  • Screen recordings and keystroke-level activity logs: reveal real navigation, workarounds and app switching. See licensed workflow and screen activity data and screen recordings for AI.
  • Ticket and request histories: give the goal wording users actually write ("reissue PO 4471 with corrected ship-to").
  • Application audit logs and database diffs: provide ground-truth end states, since they record which fields changed.
  • SOPs and runbooks: define policy constraints the agent must respect, which become negative assertions.

Recordings carry names, account numbers and customer screens, so plan redaction before tasks are derived; the guide to PII in screen recordings and computer-use trajectories covers frame-level methods. For the step-record schema used when trajectories are stored alongside tasks, see computer-use trajectory data format.

How to write a task specification with a verifiable end state

A usable task spec pins down the starting state, the goal, the application version and a programmatic checker, so any run can be reproduced and graded without a human. OSWorld's split between an initial-state config and an evaluator function is a good template [1]. The example below adapts it to an enterprise web app.

Illustrative example: invented to show structure; it does not describe an available dataset.

task_id: ap-invoice-approve-0137
app: accounts-payable-web
app_version: "2026.3.2"          # pinned build or container digest
viewport: { width: 1920, height: 1080 }
seed_snapshot: snapshots/ap_tenant_2026-09-01.sql.gz
setup:
  - login_as: role=ap_clerk
  - open_url: /invoices?status=pending
goal: >
  Approve the pending invoice from the office supplies vendor dated
  12 Aug, but only if it matches its purchase order; otherwise put it
  on hold with reason "PO mismatch".
source_provenance:
  derived_from: ticket_history + screen_session (redacted)
  redaction_method: entity replacement, sample-checked
end_state_checks:
  - query: "SELECT status, hold_reason FROM invoices WHERE invoice_no='INV-88213'"
    expect: { status: "on_hold", hold_reason: "PO mismatch" }   # PO qty differs
  - query: "SELECT count(*) FROM invoices WHERE status_changed_at > :run_start AND invoice_no <> 'INV-88213'"
    expect: 0                                                  # no collateral edits
forbidden_actions:
  - navigate: /payroll/*
  - api_call: DELETE /invoices/*
budgets: { max_steps: 40, max_minutes: 10 }
grading: pass_if_all_checks_true
trials: 5                                  # report pass^k, not best-of-k

Note the deliberate trap: the goal invites approval, but the seeded PO quantity differs, so the correct end state is a hold. Tasks with conditional outcomes separate agents that check the data from agents that pattern-match the instruction.

How many tasks per app you need for a reliable estimate

Plan for dozens to hundreds of tasks per application, because small suites produce pass rates with intervals too wide to act on [5]. A worked example using the Wilson 95% interval: an agent that passes 4 of 5 tasks has a plausible true success rate between roughly 38% and 96%. At 30 of 50 the interval narrows to roughly 46% to 72%, and at 120 of 200 to roughly 53% to 67%.

Illustrative example: invented to show structure; it does not describe an available dataset.

Tasks per appObserved pass rateApprox. 95% intervalUsable for
580%38% to 96%Smoke test only
5060%46% to 72%Spotting large regressions
20060%53% to 67%Comparing model versions

Run each task several times too. Computer-use agents are stochastic, and tau-bench showed that consistency across repeated trials (pass^k) falls well below single-trial success [4]. Label quality matters as much as count: audits of popular test sets found average label error rates of at least 3.3% [10], so have a second reviewer confirm every expected end state before a task enters the suite.

Plan maintenance for app versions and UI drift

UI drift breaks tasks silently, so treat a computer-use eval suite as maintained software rather than a static dataset. Our working hypothesis, based on how SaaS vendors ship, is that relabeled buttons, moved menus, new consent dialogs and changed default filters will invalidate a share of tasks each release, and the failure will look like agent error.

Controls that limit drift damage:

  1. Pin versions: record the app build, container digest or SaaS release tag per task, and run against frozen environments where licensing allows.
  2. Separate checker from UI: query database or API state rather than DOM text, so cosmetic changes do not break grading.
  3. Canary tasks: keep a small set of trivially solvable tasks per app; if a strong reference agent fails them, the environment changed, not the model.
  4. Seed snapshots: restore the same tenant state before every run; seed data and state snapshots for agent sandboxes covers how to license them.
  5. Retire and refresh: rotate tasks on a schedule and hold a private reserve to limit contamination, as described in contamination-resistant evaluation design.

Buyer checklist for sourcing computer-use eval tasks

Before licensing workflow data for evaluation, confirm the source can support state-based grading and that the license covers how you will use it.

Illustrative example: invented to show structure; it does not describe an available dataset.

Question to ask the data holderWhy it matters
Which applications and versions appear in the recordings or logs?Tasks must run on software you can reproduce
Are audit logs or record-level diffs available alongside screen sessions?Provides ground-truth end states
Can a de-identified seed snapshot of the app state be supplied?Enables reproducible setup
How were names, account numbers and on-screen customer data handled?Recordings are dense with personal data
Does the license permit deriving tasks, running environments and publishing scores?Eval use often differs from training use
Is the data held out from any training sets you or others license?Protects against contamination

License scope deserves particular care: license terms for agent workflow data and evaluation-only data license terms list the clauses to negotiate. If you are sourcing training data rather than eval tasks, start from training data for computer-use agents and evaluation datasets built from real business work.

How SourceX helps source workflow data for computer-use evaluation

SourceX sources operational datasets from US companies on request, including engineering records, support histories, documents and new recordings of hands-on work, and manages licensing through Find, Assess, Agree, Transact and Manage. Nothing is held in stock and a request does not guarantee a match; every release is approved by the supplying company, rights-reviewed, and delivered under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery with the method recorded and a sample checked, though no method is perfect. Teams can describe the applications and workflows they need to evaluate.

Source computer-use evaluation tasks from real business workflows

Describe the software, workflows and end-state evidence your computer-use evals need, not the businesses that might hold them. SourceX looks for US companies holding that data, serves AI teams wherever they are based, and contracts nothing until a supplier agrees. Start a buyer request.

Sources

  1. Xie et al. (arXiv; NeurIPS 2024), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
  2. Zhou, Xu et al. (arXiv), "WebArena: A Realistic Web Environment for Building Autonomous Agents" (2024). https://arxiv.org/abs/2307.13854v4
  3. Deng, Su et al. (arXiv; NeurIPS 2023), "Mind2Web: Towards a Generalist Agent for the Web" (2023). https://arxiv.org/abs/2306.06070v1
  4. Yao et al., Sierra Research (arXiv 2406.12045), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  5. mer.vin, "Your agent eval passed 4/5. That number means almost nothing". https://mer.vin/news/your-agent-eval-passed-4-5-that-number-means-almost-nothing/
  6. Li et al. (arXiv 2504.07981), "ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use" (2025). https://arxiv.org/abs/2504.07981v1
  7. Salesforce AI Research (arXiv 2505.18878), "CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions" (2025). https://arxiv.org/pdf/2505.18878
  8. Wang, Yu et al. (arXiv 2508.09123), "OpenCUA: Open Foundations for Computer-Use Agents" (2025). https://arxiv.org/abs/2508.09123
  9. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  10. Northcutt, Athalye, Mueller (arXiv; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data