Skip to content

Evaluation and benchmarking datasets

User simulator scenarios for agent evaluation: grounding simulated users in real conversations

Quick answer

A user simulator for agent evaluation is a language model that plays the customer from a written scenario: a goal, information it holds back, a persona and constraints. Benchmarks such as τ-bench score agents against these simulated users, so the scenarios and the simulator model both shape the score [1][2]. To make results predictive, write scenarios from the distribution of real conversations, script the difficult behaviors real customers show, and check simulated dialogues against held-out real transcripts before you trust a leaderboard.

By SourceX Editorial · Updated

Why the simulated user decides your agent score

The simulated user is half of every test episode, so a weak simulator produces scores that describe the simulator rather than the agent. In τ-bench, an LM-simulated user receives scenario instructions and talks to the agent, which must call domain APIs while following a policy document; grading compares the final database state to the expected one [1]. If the simulated user volunteers the order ID in turn one, never misremembers a date and accepts the first offer, the agent never has to ask, verify or refuse.

The harness setting matters as much as the scenario text. EvalScope's τ-bench integration exposes a user_model parameter, and its documentation notes that results vary with the simulator chosen [2]. Pin the simulator model, version, temperature and system prompt in every reported run, or two "τ-bench retail" numbers are not comparable.

Reliability needs repeated trials. τ-bench introduced pass^k, which asks whether an agent succeeds on all k independent attempts at the same task, because a simulated user that phrases things differently each run exposes agents that only work on the lucky phrasing [1]. Repeated trials only help if the simulator varies the way customers vary, which is a data question.

What a user simulator scenario contains

A usable scenario separates what the simulated user wants from what it reveals, how it speaks and what it will not accept. This structure descends from agenda-based simulators in task-completion dialogue research, which drove the simulated user from a user goal with constraints and requests [3], and τ-bench's instruction-driven users follow a similar pattern of goal plus details to reveal [1].

FieldPurposeFailure if missing
goalThe outcome the user will accept, stated as an end state (refund issued, flight moved)Simulator drifts, agrees to anything
known_infoFacts the user can state when asked (order number, travel date)Simulator invents IDs the database lacks
hidden_infoFacts revealed only if the agent asks the right questionAgent never tested on elicitation
personaRegister, patience, literacy, channel habitsEvery user sounds like the model's default voice
constraintsThings the user refuses (no store credit, no callback)Agent passes by steering to an easy outcome
behaviorsScripted difficulties: change of mind, partial answers, wrong assumptionsOnly the happy path is tested
stop_conditionWhen the user ends the chat (goal met, escalation, frustration limit)Endless loops or premature exits
expected_end_stateThe database or ticket state the grader checksGrading falls back to subjective judging

Keep goal and expected_end_state aligned but separate. The goal is written in the customer's words; the end state is written in your system's fields, such as a refund_status change on a specific order row, so a state-based grader can check it without an LLM judge. For grader design, see the agent evaluation task suites guide.

Grounding scenarios in real conversation distributions

Real transcripts give you the scenario mix, the opening utterances and the friction that synthetic brainstorming misses. Our working hypothesis, which you should test on your own data, is that scenarios written by sampling from real conversation logs produce more predictive scores than scenarios written from a product team's mental model of users.

Work from the transcripts outward:

  1. Sample intents by frequency and by cost. Pull contact reasons from your ticketing or chat platform (for example a disposition code or tag field on each conversation) and sample in proportion to volume, then oversample the expensive tail such as chargebacks, account takeovers and regulatory complaints.
  2. Extract the goal and the end state from the resolution. The agent's actions in the record, such as a refund issued or an address changed, tell you what the customer actually needed, which is often different from the first message.
  3. Mine hidden information from clarifying turns. Every place a human agent asked "which order?" or "when did you buy it?" marks a fact the customer did not lead with. Those become hidden_info entries.
  4. Lift persona traits from language, not demographics. Register, message length, typo rate, use of screenshots and patience before escalation are observable in logs; age or ethnicity guesses are not, and they invite stereotyped role-play.
  5. Keep policy context with each scenario. Datasets such as ABCD were built on the premise that customer-service dialogue involves the agent's guidelines and actions, not only slot values [4]. Store the policy version that applied when the real conversation happened.

The same logic applies on the query side of retrieval; the synthetic vs real queries guide covers it for retrievers.

Scripting the behaviors real customers show

Real users rarely state complete, correct requests in one turn, and scenarios must force the simulator to behave the same way. A 2026 preprint, PersonaForge, argues that much agent data assumes complete single-turn queries while real users reveal needs across turns [6], and CRMArena-Pro evaluates business agents in both single-turn and multi-turn settings because the two differ [5].

Behaviors worth encoding as explicit behaviors entries, based on what support transcripts typically contain (a hypothesis to check against your logs):

  • Ambiguity: "my last order" when there are two orders in the same week.
  • Change of mind: asks for a refund, then switches to an exchange after hearing the restocking fee.
  • Partial information: knows the city but not the booking reference, or gives the last four digits of the wrong card.
  • False premises: insists a promotion applies that expired, testing whether the agent holds policy.
  • Multi-intent: bundles a billing dispute with an address change in one message.
  • Social pressure: threatens to cancel or asks for a supervisor, testing escalation rules covered in the customer support policy evaluation guide.

Write each behavior as a trigger and a response, for example "if the agent quotes a fee, switch goal to exchange," rather than a vague trait like "indecisive." Conditional triggers are easier to audit than adjectives, and they make failures reproducible across pass^k trials.

User simulator bias and how to measure it

A user simulator is biased when its dialogues differ systematically from real customers in ways that change agent scores. Common patterns include over-cooperation (answering questions the agent never asked), politeness collapse into one voice, leaking hidden_info early, accepting policy-violating offers the persona should reject, and self-preference effects when the simulator and agent come from the same model family. Treat these as hypotheses to measure, since they vary by simulator model [2].

Validate the simulator against held-out real transcripts before using it for model selection:

CheckHow to computeWarning sign
Turn-count distributionCompare turns per resolved conversation, simulated vs real, per intentSimulated chats much shorter
First-message completenessShare of openers that contain the key identifierSimulator front-loads every fact
Clarification rateAgent questions per conversation needed to reach the goalNear zero in simulation
Goal-switch rateShare of conversations where the requested outcome changesNever happens in simulation
Escalation and abandonmentShare ending in human handoff or user exitSimulated users never give up
Score stabilityAgent success with two different simulator modelsRanking of agents flips
Human discrimination testReviewers label mixed real and simulated transcriptsReviewers spot simulations easily

If agent rankings flip when you swap the simulator, report both numbers and treat the benchmark as unresolved for that decision. Keep the real transcripts used for validation out of any training mix and out of prompt examples, for the reasons covered in contamination-resistant evaluation design.

Illustrative scenario record

A scenario file should be machine-readable, versioned and traceable to the real conversation pattern it was drawn from, without carrying that conversation's personal data.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "scenario_id": "retail-return-0137",
  "source_pattern": {"intent_tag": "return_after_window", "sample_stratum": "tail_high_cost", "policy_version": "returns-2026-03"},
  "goal": "Get a refund to the original card for a jacket bought 34 days ago",
  "known_info": {"email_on_account": "<redacted-pseudonym>", "item": "rain jacket, size M"},
  "hidden_info": {"order_id": "revealed only if asked for order number", "gift_receipt": "true, revealed only if asked how it was purchased"},
  "persona": {"register": "terse, lowercase", "patience_turns": 6, "channel": "web chat"},
  "constraints": ["refuses store credit twice", "will not call a phone line"],
  "behaviors": [
    {"trigger": "agent cites 30-day window", "response": "claims the item arrived late"},
    {"trigger": "agent offers exchange", "response": "accepts only if same color is in stock"}
  ],
  "stop_condition": "goal met, or patience_turns exceeded, or human handoff",
  "expected_end_state": {"returns.status": "approved_exception", "refund.method": "store_credit"},
  "simulator": {"model": "<pinned model id>", "temperature": 0.7}
}

Note that the expected end state differs from the stated goal: under this invented policy the correct agent behavior is an exception for late delivery, paid as store credit despite the user's push. Scenarios where the right answer disappoints the user are the ones that separate policy-following agents from agreeable ones.

Sourcing conversations for simulator grounding

Grounding needs real conversations with outcomes, policies and enough metadata to stratify, licensed for evaluation use and stripped of personal details. Specify these fields in any request:

  • Transcript turns with speaker role, timestamps and channel (chat, email, voice transcript); for spoken interactions see voice agent evaluation sets.
  • Contact reason or disposition tags and resolution codes, ideally with the back-office action taken.
  • Policy or macro version in force at the time.
  • Outcome signals: reopen within 7 days, CSAT where collected, escalation flag.
  • De-identification method for names, emails, phone numbers and account numbers. Automated PII tools such as Presidio state they cannot guarantee finding everything [7], so ask for the method and a reviewed sample.
  • Provenance and consent basis. FTC staff have warned that quietly expanding how customer data is used for AI may be unfair or deceptive [8], so ask how the supplier's customer terms cover the release.
  • Documentation as a dataset card with license, language and size metadata [9].

Your own production logs are the first candidate, but they only describe your current users. If you are launching a new vertical or channel, you need conversations from businesses that already serve it. SourceX sources operational datasets, including support and sales histories, from US companies on request; it does not hold them in stock, and a request does not guarantee a match. Describe the conversations you need on the SourceX buyer page, and see the owner pages for licensed chat logs, customer support transcripts and evaluation sets built from real business work.

Keeping simulator scenarios private and fresh

Scenario files are evaluation data and lose value once they appear in training corpora or public repos. Store them with the same access controls as a private test set, rotate a share of scenarios each release, and keep a frozen subset for trend lines. The tradeoffs between held-out and public sets are covered in private evaluation sets vs public benchmarks, and the wider map of options is on the LLM evaluation datasets hub.

Request real conversations to ground your user simulator

SourceX looks for US businesses that hold the conversation data you describe, reviews ownership and consents, and delivers under a license defining records, uses, term and delivery once the supplying company approves. Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Describe your scenario needs at sourcex.si/buyers.

Sources

  1. Yao et al., Sierra (arXiv:2406.12045), "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  2. EvalScope (ModelScope), "tau bench (EvalScope benchmark documentation)". https://evalscope.readthedocs.io/en/latest/benchmarks/tau_bench.html
  3. Microsoft Research (Li et al.), "A User Simulator for Task-Completion Dialogues" (2016). https://www.microsoft.com/en-us/research/publication/user-simulator-task-completion-dialogues/
  4. Chen et al., NAACL 2021 (arXiv:2104.00783), "Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems" (2021). https://arxiv.org/abs/2104.00783v1
  5. Salesforce AI Research (arXiv:2505.18878), "CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions" (2025). https://arxiv.org/pdf/2505.18878
  6. arXiv (Hanglong Lv et al.), "PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems" (2026). https://arxiv.org/abs/2608.28378
  7. Microsoft (microsoft/presidio), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  8. Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  9. Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data