Evaluation and benchmarking datasets
Agent evaluation task suites: tasks, environments and state-based grading
Quick answer
An AI agent evaluation benchmark is a suite of tasks in which each task fixes a starting environment state, the tools the agent may call, the policy it must follow, a scripted or simulated user, and an annotated goal state. A grader checks the environment's final state against that goal rather than judging the conversation, and each task runs several times so the suite measures reliability (pass^k), not a single run. Building a domain suite means sourcing realistic records, written policies and verified outcomes.
By SourceX Editorial · Updated
What every agent evaluation task has to specify
An agent task is a reproducible episode: the same starting state, tools, rules and request must produce an end state a program can check, every time it runs. τ-bench (Tool-Agent-User benchmark), published by Sierra researchers in 2024, made this structure explicit by building each domain from realistic databases and APIs, domain-specific policy documents, and instructions for user scenarios with ground-truth annotations [1]. Generalized, a task needs six things written down:
- Seed state: the records the agent can read and change at step zero, stored as a snapshot with a content hash.
- Tool surface: the functions, API endpoints or user interface the agent may use, split into read and write calls.
- Policy: the rules that decide what a correct action is, with a version and an effective date.
- Trigger: a user conversation, or for back-office agents, a queue item such as a blocked invoice or an unassigned ticket.
- Goal state: expected record changes, facts the agent must report, and calls it must never make.
- Trial protocol: number of runs, step limits, simulator settings and the suite version.
Enterprise suites add two details. The agent acts under a role with limited permissions, so the forbidden-call list should include writes outside the role. Policies also change: a task seeded from a case resolved under last year's approval limits must ship with those limits, or its goal state will contradict the policy the agent reads. Other kinds of held-out data are mapped in the buyer's guide to evaluation and benchmarking datasets.
How public agent benchmarks build environments, tasks and graders
Public agent benchmarks make three engineering choices that a private suite must also make: how real the environment is, where tasks come from, and what the grader inspects.
| Benchmark | Environment | Task source | Grader | Scale as published | Lesson for a private suite |
|---|---|---|---|---|---|
| τ-bench [1] | Databases behind tool APIs, a policy document and a simulated user | Authored scenarios in two synthetic domains | Final database vs annotated goal state; pass^k over trials | 115 retail and 50 airline tasks | State-based grading needs no user interface |
| WebArena [2] | Self-hosted, fully functional websites in four domains: e-commerce, forums, collaborative software development, content management | Authored tasks on those sites | End-to-end task success | 812 tasks | Real application stacks can be hosted and reset privately |
| OSWorld [3] | Real computer environments (the framework supports Ubuntu, Windows and macOS) with per-task setup | Web and desktop apps, file I/O, multi-app workflows | Execution-based evaluation | 369 tasks | Highest fidelity, highest infrastructure cost |
| Mind2Web [4] | Real websites | 137 sites in 31 domains | Crowdsourced action sequences as reference | Over 2,000 tasks | Recorded demonstrations suit step-level checks, not end-state checks |
| SWE-bench [5] | Repository at the pre-fix commit | Real GitHub issues and pull requests, 12 Python repositories | Repository tests | 2,294 problems | Tasks mined from real work arrive with verification attached |
| SWE-Bench Pro [7] | Repository environments | Copyleft repositories plus startup codebases under partnership agreements | Tests | Public, held-out and commercial subsets | Licensed private sources resist contamination |
Treat every score in the table as dated, because benchmarks keep changing. The τ-bench family added τ²-bench, whose telecom domain lets both the agent and the simulated user operate tools; according to the maintainers' documentation as of October 2026, a v1.0.0 release in March 2026, branded τ³-bench, adds corrected task sets, full-duplex voice evaluation and a knowledge-retrieval banking domain [8]. Scores from different versions are not comparable.
Grading the end state, the path or the words
State-based evaluation asks whether the systems ended up correct, which lets any valid sequence of actions pass; path and wording checks catch what a state diff cannot see. τ-bench compares the database at the end of the conversation with the annotated goal state [1], the right default for agents that change records. The engineering work is in making that comparison fair.
| Grader | What it verifies | Engineering trap | Mitigation |
|---|---|---|---|
| Full state diff | Every table matches the goal snapshot | Timestamps, generated IDs and audit columns differ on every run | Mask volatile fields and keep the mask list in the task |
| Assertions on invariants | Named rows and fields hold expected values | Too few assertions let collateral writes pass | List tables the task must leave unchanged |
| Multiple acceptable end states | Any policy-permitted outcome passes | Authors record only the outcome the historical case took | Enumerate alternatives from the policy, not the case |
| Output check | Facts the agent must report, such as a variance amount | Read-only tasks leave nothing to diff | Assert required facts in the final reply |
| Tool-call log rules | Forbidden calls, required ordering (verify before write) | A write later reversed leaves a clean end state | Fail on the call itself |
| Rubric or LLM judge | Tone, explanation, disclosure wording | Judge drift and bias | Calibrate against human labels; re-check when the judge model changes |
A state diff also misses side effects outside the graded tables, such as a notification sent during a cancel-and-recreate sequence, so log every tool call even when grading only the end state. Support-specific grading rules (refund limits, identity checks, required disclosures) are worked through in policy-following evaluation for customer support agents, and tasks built around irreversible actions in agent safety evaluation for irreversible actions.
pass^k versus pass@k, and which one your deployment needs
pass^k is the probability that an agent succeeds on all k independent trials of the same task; τ-bench introduced it and estimates it per task as C(c,k)/C(n,k) for c successes in n trials, averaged across tasks [1]. It is the mirror image of pass@k, the code-generation metric that credits a model when any one of k attempts succeeds.
Worked example: one task, eight trials, six successes.
| k | pass^k = C(6,k)/C(8,k) | pass@k = 1 − C(2,k)/C(8,k) |
|---|---|---|
| 1 | 0.750 | 0.750 |
| 2 | 0.536 | 0.964 |
| 4 | 0.214 | 1.000 |
| 8 | 0.000 | 1.000 |
Choose the metric from how the agent will run. If a verifier can pick a good attempt before anything ships, as when generated code must pass tests before merge, pass@k describes the useful capability. If every run touches a live record or a customer, pass^k is the one that predicts failures, because the same situation recurs and each occurrence is a fresh draw.
Public results show how steep the decay can be: a third-party leaderboard summarizing τ-bench lists a 2024 Claude 3.5 Sonnet result of 0.460 at pass^1 on the airline domain, falling to 0.225 at pass^4 [9]. Run at least as many trials per task as the largest k you report.
Tasks versus trials: why a five-task suite says little
The number of distinct tasks, not the number of trials, sets how precisely a suite estimates success on your work. Practitioners put it bluntly: an agent eval that "passed 4/5" means almost nothing [10]. The 95% Wilson score interval for 4 of 5 tasks runs from about 38% to 96%; for 40 of 50 it narrows to about 67% to 89% (our calculation, assuming tasks are independent draws from the work you care about).
- Trials do not replace tasks. Twenty runs of five tasks still sample five situations, so compute intervals by resampling tasks, not runs.
- Avoid normal-approximation intervals on small suites. A position paper at ICML 2025 argues that CLT-based intervals are too narrow for evaluations with fewer than a few hundred datapoints [11].
- Size every reported slice. If invoice holds, vendor changes and duplicate payments are reported separately, each needs its own task count; sizing an eval set for statistical power covers the arithmetic.
Environment fidelity, reset and determinism
The agent evaluation environment decides what a suite can observe, and an environment that does not reset identically turns infrastructure noise into agent failures. OpenAI's screening for SWE-bench Verified found unreliable development-environment setup in the original benchmark, alongside overly specific tests and underspecified problem descriptions, and shipped a Docker-based harness [6]. Pick the lowest fidelity that still exposes the failures you care about.
| Environment type | Catches | Reset method and determinism | Seeded from | Go deeper |
|---|---|---|---|---|
| API mock over a database | Policy errors, wrong arguments, wrong records | Restore snapshot; highly repeatable | Record exports and API schemas | Seed data and state snapshots for agent sandboxes |
| Self-hosted application replica | Navigation and form errors as well as logic errors | Container and database restore; moderate | Populated application database plus configuration and customizations | Configuration data for enterprise app replicas |
| Full desktop or VM | Multi-application, file-system and settings errors | VM snapshot restore; slowest, most timing noise | Files, accounts, installed applications | Computer-use agent evaluation tasks |
| Repository container with tests | Code changes that break behavior | Pinned image; repeatable when dependencies are mirrored | Repository at base commit plus tests | SWE task environments with tests |
Measure flakiness before scoring any agent: run the reference solution many times and repair or remove tasks it does not pass every time. The training-side use of the same environments is covered on training data for RL environments.
The simulated user is a test parameter
When a task has a conversational trigger, the user is usually an LLM following hidden scenario instructions, so its model and prompt are part of the test. The EvalScope harness for τ-bench, for example, exposes a user_model setting for the model that plays the user, defaulting to qwen-plus [12], so two teams using different simulators are not running the same evaluation. Pin the simulator model, version, temperature and scenario text, and change them only between suite versions.
Queue-triggered back-office tasks avoid this variable, so report them separately from conversational tasks. Grounding simulator instructions in real conversations is covered in user simulator scenarios for agent evaluation.
Sourcing the records behind a domain suite
A domain suite needs four kinds of real material: record populations to seed state, the policies staff applied, real triggers, and outcomes someone verified. Invented records are usually cleaner than production data, so they can omit the duplicates, stale fields and exceptions that break agents. The table maps each part to source systems in three enterprise domains; enterprise workflow datasets and enterprise workflow agents describe the underlying work histories.
| Suite part | Accounts payable agent | IT service agent | Sales operations agent |
|---|---|---|---|
| Seed state | ERP vendor master, purchase orders, goods receipts, open invoices | CMDB entries, user directory, open ticket queue | CRM accounts, opportunities, price books |
| Policy | AP policy, match tolerances, approval matrix by amount | Runbooks, change-approval policy, access rules | Discount approval rules, deal-desk guidelines |
| Trigger | Exception-queue items, vendor emails | Incident and request tickets | Rep requests for quotes or discounts |
| Verified outcome | Posted invoices, holds, approval logs, reversals | Resolution codes, change records, reopen flags | Approved quotes, booked orders, rejected requests |
Two properties matter more than volume. First, outcomes must be verified, because a closed record is not a correct one; task success labels from business records covers proxy outcomes and their traps. Second, records should never have been public: in a February 2026 post, OpenAI said it had stopped reporting SWE-bench Verified because gains increasingly reflected training-time exposure [13], while SWE-Bench Pro added a commercial set from startup codebases under partnership agreements and keeps that code private [7]. Contamination-resistant evaluation design covers the governance side.
When these records sit inside other companies, SourceX can help: it sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing agreements. Datasets are sourced on request rather than held in stock; each goes through rights review, and personal details such as names, emails and account numbers are removed or replaced before delivery, with the method recorded. Buyers describe the records a suite needs to SourceX rather than finding suppliers themselves. Ask for de-identification that keeps the joins and values your assertions depend on (de-identifying evaluation data without breaking the test) and for license terms covering environment construction and derived tasks (license terms for agent data).
Illustrative task record: an accounts-payable exception
A task record should hold everything needed to rerun and regrade the task without its author, including where the task came from and under what license.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"task_id": "ap-exc-0093",
"suite": {"name": "ap-exceptions", "version": "2026.10-r2", "task_set_sha256": "<hash>"},
"split": "held_out",
"environment": {
"type": "api_mock_over_db",
"image": "ap-sandbox:4.1.0",
"seed_snapshot": "snapshots/erp_ap_2026-06-30.parquet",
"seed_sha256": "<hash>",
"reset": "restore snapshot before every trial"
},
"role": {
"name": "ap_clerk_agent",
"write_tools": ["hold_invoice", "route_for_approval", "message_vendor"],
"forbidden_tools": ["post_payment", "edit_vendor_bank_details"]
},
"policy": {"doc": "ap_policy_v7.md", "effective_from": "2026-01-01",
"clauses": ["3-way match tolerance", "price variance above tolerance needs buyer approval"]},
"trigger": {"type": "queue_item", "text": "INV-55120 blocked: price variance on line 2"},
"user_scenario": null,
"goal_state": {
"assertions": [
{"table": "invoices", "key": {"invoice_id": "INV-55120"},
"expect": {"status": "on_hold", "hold_reason": "price_variance"}},
{"table": "approval_requests", "key": {"invoice_id": "INV-55120"},
"expect": {"approver_role": "buyer", "row_count": 1}}
],
"unchanged_tables": ["payments", "vendor_master"],
"volatile_fields_masked": ["updated_at", "approval_request_id"],
"required_outputs": ["variance amount", "reason for hold"]
},
"reference_solution": "solutions/ap-exc-0093.jsonl",
"trial_protocol": {"trials": 8, "max_steps": 40},
"provenance": {"source_record": "<supplier exception ID, stored outside the suite>",
"outcome_verified_by": "approval log and posted invoice",
"license_ref": "<agreement schedule>", "deidentification": "<method and date>"}
}
The user scenario is null because the trigger is a queue item; the unchanged-tables list catches collateral writes such as an early payment; and the reference solution makes the solvability check below possible. Supplier IDs stay outside the suite, so the task can circulate internally. Finance-specific ground truth is covered in accounting agent evaluation sets.
Acceptance tests before you trust a suite
Accept a suite, in-house or vendor-built, only after its tasks prove solvable, not trivially passable and resettable, because a broken task scores as an agent failure.
- Solvable: the reference solution reaches the goal state on every trial in a freshly reset environment.
- Not trivially passable: an agent that makes no tool calls fails every task requiring a write, and a script that calls every write tool fails the forbidden-call and unchanged-table checks.
- Resettable: the seed hash after reset matches the recorded hash before each trial.
- Graded on the right fields: each masked field is justified, and each assertion maps to a policy clause.
- Adjudicated: two domain reviewers agree on each goal state, resolving disagreements against the policy text.
- Anchored to people: domain staff attempt a sample of tasks. WebArena reported 78.24% human success against 14.41% for its best agent [2], and OSWorld over 72.36% against 12.24% [3]; those 2024 model scores are outdated, but the method is not. See human baseline data for agent evaluation.
- Held out: tasks are excluded from training and prompt-tuning pipelines, and access is logged; see keeping a private eval set private.
- Versioned: suite version, task-set hash, environment image, tool schemas and simulator settings are stored with every score.
Building an agent evaluation suite from real business records?
Describe the seed records, policies, triggers and verified outcomes your suite needs, and the uses you need licensed, such as building environments and deriving tasks. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees, and a request does not guarantee a matching dataset. Describe your agent evaluation data needs.
Sources
- Yao, Shinn, Razavi, Narasimhan (Sierra), "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- Zhou, Xu et al. (Carnegie Mellon University), "WebArena: A Realistic Web Environment for Building Autonomous Agents" (arXiv v4, 2024). https://arxiv.org/abs/2307.13854v4
- Xie et al. (XLANG Lab, HKU and collaborators), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
- Deng, Su et al. (The Ohio State University), "Mind2Web: Towards a Generalist Agent for the Web" (2023). https://arxiv.org/abs/2306.06070v1
- Jimenez et al. (Princeton, UChicago), "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
- OpenAI, "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
- Scale AI researchers, "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
- Sierra Research (repository documentation, indexed by Algolia DocSearch), "τ-Bench: sierra-research/tau2-bench" (accessed 2026). https://docsearch.algolia.com/mcp/docs/repo/sierra-research/tau2-bench
- Steel.dev leaderboard (third-party aggregator), "Tool use benchmark - Public: tau-bench" (accessed 2026). https://leaderboard.steel.dev/registry/benchmarks/tau-bench
- mer.vin (practitioner blog), "Your agent eval passed 4/5: that number means almost nothing" (accessed 2026). https://mer.vin/news/your-agent-eval-passed-4-5-that-number-means-almost-nothing/
- ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
- EvalScope documentation, "tau bench" (accessed 2026). https://evalscope.readthedocs.io/en/latest/benchmarks/tau_bench.html
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.