Skip to content

Evaluation and benchmarking datasets

Customer support AI agent evaluation: policy-following tests from real cases

Quick answer

Customer support AI agent evaluation should test two things: whether the agent leaves the account in the state your written policy requires, and whether it does so on every run. Build each task from a real resolved case: the policy version in force, an account or order snapshot, the customer's goal and the end state a correct resolution produces. Grade the final state and actions, run every task several times, and report pass^k (the share of tasks solved in all k runs) next to single-run success.

By SourceX Editorial · Updated

For support training data and general eval sets, see customer support AI training data and eval sets and business data for customer support AI; the evaluation datasets hub maps the cluster.

What τ-bench established, and why teams look for a real-data alternative

τ-bench (tau-bench) is a customer service agent benchmark with a reusable task structure, but its domain data are synthetic, its tasks are public and its policies are not yours. Use it to compare base models, and tasks built from your own policies and cases to decide whether an agent is ready for customers.

The 2024 τ-bench paper builds each domain from three parts: a database with programmatic APIs the agent can call, a domain policy document, and simulated-user scenario instructions, each paired with an annotated goal state as ground truth [1]. Grading compares the database at the end of the conversation with that goal state, so a courteous transcript that issues the wrong refund still fails [1]. The original release has 165 tasks in two synthetic domains, 115 retail and 50 airline, and introduced pass^k to measure consistency across repeated trials [1].

The family keeps changing. τ²-bench added a dual-control telecom domain in which both the agent and the simulated user operate tools; as of October 2026, the maintainers' documentation describes version 1.0.0 (March 2026), branded τ³-bench, as adding corrected task sets, full-duplex voice evaluation and a knowledge-retrieval banking domain, with the original repository kept as a historical release [2], so scores from different versions are not comparable. Public task sets also age: in February 2026 OpenAI stopped reporting the coding benchmark SWE-bench Verified, saying score gains increasingly reflected training-time exposure to its tasks [3].

τ-bench familyEval set built from your support records
PolicyWritten for the benchmark domainYour policy documents, macros and approval limits, by version
Account stateSynthetic databasePoint-in-time snapshots of real accounts, orders or subscriptions
Customer behaviorSimulated from scenario instructionsSimulator instructions grounded in real transcripts
Case mixChosen by the benchmark authorsYour ticket distribution, with rare and costly cases oversampled
Ground truthAnnotated goal stateVerified outcomes, re-adjudicated against the policy version in force
ExposurePublicPrivate, if kept out of training and prompt tuning

Anatomy of a policy-following support task

A gradable task links six records: the policy excerpt and its version, the initial account state, the tools the agent may call, the customer scenario, the acceptable end states, and the actions that must never happen. Without all six, grading falls back to reading transcripts, which is slow and inconsistent.

Support agent test scenarios need more than a goal. Real customers hold facts the agent must ask for, change their request mid-conversation and push back on a denial, so simulator instructions should say which facts to volunteer and which to withhold (user simulator scenarios grounded in real conversations). The expected outcome is often a set: if policy allows either a refund to the original payment method or store credit, both should pass.

Illustrative example: invented to show structure; it does not describe an available dataset.

task_id: subs-billing-0147
source_case: case_7f3a9c          # pseudonymized; same surrogate key in ticket, account and billing tables
channel: chat
opened_at: 2026-04-20T15:12:00Z   # 18 days after renewal
policy:
  doc_id: billing-refund-policy
  version: 2026-03-01             # version in force when the case was opened
  clauses:
    - "4.2 annual plans: prorated refund if requested within 30 days of renewal"
    - "4.5 refunds above 200.00 require tier-2 approval"
initial_state:                    # snapshot rebuilt as of the first customer message
  account: {plan: annual_pro, renewed_on: 2026-04-02, seats: 12}
  invoices: [{id: inv_88, amount: 1188.00, status: paid, method: card_on_file}]
  credits_last_12_months: 1
tools: [get_account, list_invoices, issue_refund, apply_credit, escalate_to_tier2, send_export_link]
user_scenario:
  goal: "cancel and get money back for the unused months"
  volunteers: [renewal date, seat count]
  withholds_until_asked: [wants data export before cancellation]
  behavior: ["says the renewal notice never arrived", "asks twice for a full refund"]
expected:
  acceptable_end_states:
    - {refund_issued: none, escalation: tier2, reason_code: refund_over_limit, export_link_sent: true}
  forbidden_actions: ["issue_refund with amount > 200.00", "apply_credit before escalation"]
  required_statements: [explains prorated basis, sets expectation for tier-2 follow-up]
ground_truth:
  origin: historical_resolution
  verified_by: [qa_scorecard_pass, not_reopened_30_days, no_chargeback]
  adjudicated_against: billing-refund-policy@2026-03-01
slices: [annual_plan, over_approval_limit, escalation_required]
trials: 8

Here the correct behavior is to escalate, not refund, because the prorated amount exceeds the agent's approval limit. A transcript-similarity grader would reward an agent that sounds helpful and issues the refund; an end-state grader with a forbidden-action list fails it. For other domains, see agent evaluation task suites.

Where each piece of a task lives in a support operation

The parts of a task sit in different systems; the hardest to recover is the account state when the customer wrote in. Ask a supplier which of these records exist, how far back they go, and whether they join on a shared case or account key.

RecordWhat it supplies to a taskCommon defect
Policy documents and help-center articles with revision historyPolicy text for each versionEdits overwrite history
Macros (saved replies) with change logsRules staff actually applied; required wordingMacros drift from the written policy
Ticket threads: public replies, internal notes, tags, status changes, timestampsCustomer scenario, pushback, final dispositionNotes carry customer data and staff names; inconsistent tags
Order, billing or subscription recordsInitial and expected end stateOnly current state kept; an event log is needed to rebuild state at case open
Action and approval logs (refunds, credits, cancellations, overrides)Tool calls, amounts, approversActions taken in another system with no ticket link
QA scorecardsPer-case policy-compliance judgmentsSampled cases only; rubric changes over time
Reopens, chargebacks, escalations, satisfaction surveysEvidence that an outcome heldArrive weeks after the ticket closes

The same linked records train policy-following agents (policy-following service agent data); for retail, see e-commerce order-support conversations with order state. Record types are described on customer support ticket datasets, licensed customer support transcripts and QA scorecards.

Which historical resolutions can serve as ground truth

A closed ticket records what a human agent did, not what the policy required. Use a case as ground truth only when its outcome was checked, it matches the policy version in force, and any exception was approved by someone with authority to grant it.

Workable filters are a passing QA score where one exists, no reopen within a set window, no chargeback or reversal, and an action that matches the policy clause for that version. Goodwill credits granted outside policy belong in a separate slice whose expected agent behavior is usually escalation, not the credit. When a refund window or approval limit changes, retire or relabel every task keyed to the old version. Outcome-labeled evaluation data covers outcome signals across domains.

Audit the goal states before acceptance. Northcutt et al. estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets and showed that such errors can change model rankings [4]. Have two reviewers re-adjudicate a random sample against the policy text and measure agreement with Krippendorff's alpha, where 1 means perfect agreement and 0 means agreement no better than chance [5]. Low agreement usually means the policy text is ambiguous; see gold-label audits and adjudication.

Measuring reliability with pass^k and repeated runs

Single-run success overstates how an agent will treat customers, because the same situation recurs across many conversations and each one is a fresh draw. Run every task n times, compute pass^k for each k up to n, and report the curve with the trial count and simulator settings.

The τ-bench paper defines pass^k as the chance that all k independent trials of a task succeed, estimated for a task with c successes in n trials as C(c,k)/C(n,k) and averaged across tasks; pass^1 is the ordinary success rate [1].

Worked example with invented counts (not results from any agent): four tasks, eight runs each.

TaskSuccesses out of 8pass^1pass^4 = C(c,4)/C(8,4)
A81.00070/70 = 1.000
B70.87535/70 = 0.500
C60.75015/70 = 0.214
D40.5001/70 = 0.014
Mean0.7810.432

A dashboard reading 78% success hides that fewer than half of these situations are handled correctly four times in a row. Task B's 87.5% success rate still means one customer in eight in that situation gets the wrong outcome, and only half of its four-run sequences are fully correct.

The simulated customer is itself a model: as of October 2026, the EvalScope harness for τ-bench exposes a user_model parameter, defaulting to qwen-plus, that sets which model plays the user [6]. Fix and report the simulator model, prompt and temperature, and change them only between eval versions. Runs of one task are not independent samples of the task population, so compute intervals by resampling tasks, not runs. A position paper at ICML 2025 argues that CLT-based intervals are too narrow for evals with fewer than a few hundred items [7]; if your task count is in that range, see sizing an eval set for statistical power.

Grading actions and wording, not only the end state

End-state comparison catches wrong refunds and missed cancellations, but support policy also governs sequence and wording. To evaluate support agent policy compliance fully, add checks for forbidden calls, ordering and required disclosures, each tied to a record showing what correct looks like.

Policy requirementHow to grade itRecords needed
Correct account changeFinal state matches one of the acceptable end statesBefore and after snapshots; action log
No action above the agent's authorityTool-call log has no forbidden call or argument, even if later reversedApproval limits by policy version
Verify identity before changing an accountOrdering assertion over tool callsVerification steps defined in policy
Required disclosure (fees, timelines, cancellation terms)Rubric item scored by humans, or by an LLM judge checked against human labelsMacros and QA scorecard items
Escalate when policy says soHandoff call with the correct reason codeEscalation records with reasons
Refuse when policy says noState unchanged and the reason explainedDenied-request cases with verified outcomes

Tasks built around irreversible actions, such as account closure or payouts, deserve their own slice and more trials; see agent safety evaluation for irreversible actions.

Rights and privacy questions for support records

Support records carry customers' personal data and the supplier's promises to those customers, so rights review starts with what those customers were told. Check the privacy notice in force when the records were created, sector rules, recording consent, and whether de-identification preserves the joins your tasks need.

  • Privacy promises. A 2024 FTC staff blog post warned that adopting more permissive data practices, such as AI training or third-party sharing, through a surreptitious, retroactive change to terms or a privacy policy may be unfair or deceptive [8]. Ask which notice versions covered the tickets in scope; see whether customer consent is needed to license support tickets.
  • California deidentification. Under Cal. Civ. Code §1798.140, information counts as deidentified only if, among other conditions, the business contractually binds recipients to the definition's requirements, which include not attempting re-identification [9]. Expect that clause in a license from a CCPA-covered business.
  • Financial-services support. Tickets from a financial institution can contain nonpublic personal information, and 12 CFR 1016.11 (Regulation P) limits a recipient's reuse and redisclosure of it, even if the recipient is not a financial institution [10].
  • Call recordings. California Penal Code §632 prohibits recording a confidential communication without the consent of all parties [11], so voice tasks need the consent record (voice agent evaluation sets).
  • Redaction that keeps tasks gradable. Customers can paste card numbers into free text, and the Presidio project cautions that its ML-based detection cannot guarantee it finds all sensitive information [12]. Use consistent surrogates across ticket, account and billing tables and one date offset, or a 30-day refund rule stops being testable (de-identifying evaluation data without breaking the test).
  • Evaluation use. Settle whether tasks may go to hosted model APIs, whether scores may be published and whether you may generate task variants (evaluation-only data license terms).

SourceX sources operational datasets from US companies, including support histories and business documents, and manages the licensing process. It sources on request rather than from stock, so a request does not guarantee a match. Each dataset goes through rights review (does the business own or may it share the records, and are required consents in place) and is delivered under a license defining the included records, allowed uses, term and delivery.

Names, emails, phone numbers and account numbers are removed or replaced before delivery and a sample is checked, though no de-identification method is perfect; diligence materials on source, rights, preparation and allowed use are prepared per dataset for your reviewers. If your eval needs linked policies, cases and account states, describe the support records your evaluation needs.

Request checklist and red flags

A request for support-agent eval data should name the policies, the state data and the outcome evidence, not only a ticket count. Treat each red flag as a reason to ask more before acceptance.

Specify:

  • Channels and domains: chat, email or voice; billing, orders, subscriptions or account security
  • Policy documents and macros with revision dates covering the whole case window
  • Account or order state at case open and after resolution, plus action and approval logs, joined by a shared case key
  • Outcome evidence: QA scores and reopen, chargeback and escalation flags
  • Minimum cases per slice, including exceptions, denials and escalations
  • De-identification method, with consistent surrogates and date offsets across tables
  • License scope: evaluation use, hosted-API exposure, publication of scores, derived tasks

Red flags:

  • Transcripts with no account state, gradable only by reading
  • "Resolved" status as the only ground truth
  • Policies paraphrased from a public benchmark domain and presented as real-world
  • Single-run pass rates with no trial count, simulator model or temperature
  • Tasks also licensed as training data to the developers of the models you will test (keeping a private eval set private)

Building a support-agent eval from real policies and cases?

Describe the support data your evaluation needs: domains and channels, policy versions, state and action records, outcome evidence and license scope. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Specify your support-agent eval data.

Sources

  1. Yao, Shinn, Razavi, Narasimhan (Sierra; arXiv:2406.12045), "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  2. Sierra Research tau2-bench repository (indexed by Algolia DocSearch), "τ²-bench / τ³-bench repository documentation" (2026). https://docsearch.algolia.com/mcp/docs/repo/sierra-research/tau2-bench
  3. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  4. Northcutt, Athalye, Mueller (arXiv:2103.14749; NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  5. Klaus Krippendorff (University of Pennsylvania, Annenberg School for Communication), "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  6. EvalScope documentation, "tau bench (benchmark documentation)". https://evalscope.readthedocs.io/en/latest/benchmarks/tau_bench.html
  7. ICML 2025 (poster), "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  8. Federal Trade Commission, Office of Technology (Tech@FTC staff blog), "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  9. California Legislature (California Legislative Information), "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  10. Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
  11. California Legislature (California Legislative Information), "California Penal Code section 632 (eavesdropping on or recording confidential communications)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632
  12. Microsoft presidio project (indexed on pkg.go.dev), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data