Agent, workflow and domain-reasoning data
User simulator training data: what real requester behavior agents need
Quick answer
A credible user simulator is calibrated on the requester side of real conversations: what people wanted, what they disclosed and when, how they corrected or abandoned the exchange, and how often each pattern occurs. Buy or license de-identified session logs with intent labels, turn-level disclosure, outcome codes and segment metadata, then hold out a slice of real sessions to test whether simulated users reproduce the real distributions before you trust any RL reward or evaluation score.
By SourceX Editorial · Updated
This page covers simulator building. For scenario design on the evaluation side, see user simulator scenarios for agent evaluation; for agent-side demonstrations and workflow records, start at the agent training data hub.
Why agent teams need real requester data, not just agent transcripts
Simulators exist because live users cannot be put in every RL rollout or regression run, and the simulated user is only as realistic as the data it was fitted to. The problem is old: Microsoft's task-completion simulator was built from example dialogues precisely because collecting fresh human dialogues for every policy iteration was too expensive [3]. Modern benchmarks such as tau-bench put an LLM-simulated customer in the loop with tool APIs and domain policies, so the simulated user's behavior directly shapes the measured pass rate [4].
The failure mode is a simulator that is too cooperative. It states its full goal in turn one, never contradicts itself, and accepts the first plausible answer. The PersonaForge authors argue that most agent training data assumes complete single-turn queries, while 75.9% of the 16K real sessions they analyzed were multi-turn [1]. An agent tuned against an over-helpful simulator learns to skip clarification, and its offline scores overstate production performance.
What a simulated user has to model
A simulated user needs a hidden goal, a disclosure policy, a patience budget and a style, each estimated from real sessions. Balog and Zhai frame user simulation as a bottleneck for both evaluation and training data and argue it should be parameterized by user type, such as novice versus expert [2]. In practice, the parameters below are the ones agent teams most often under-specify.
- Goal and sub-goals. The underlying task (refund, plan change, access request, purchase order status), including goals that change mid-conversation.
- Information disclosure over turns. Which facts the requester volunteers, which only appear when asked, and which are never provided (order number missing, wrong account referenced).
- Ambiguity and error. Vague first messages, wrong product names, mistaken dates and self-corrections.
- Patience and escalation. Turns or minutes before the requester asks for a human, repeats themselves, or abandons.
- Policy pressure. Requests for exceptions the agent must refuse; the ABCD corpus is useful here because agent actions are bound by company policies rather than just slot filling [5].
- Style and channel. Message length, typos, multi-question messages, email versus chat versus voice transcript conventions.
Which records calibrate those parameters
Each simulator parameter maps to a specific field in operational systems, which is why requester-side records from ticketing and messaging tools are the core input. Knowledge-grounded simulator research shows that tying simulated users to external facts increases diversity and realism [6], and real account and order context plays that role in enterprise settings.
| Simulator parameter | Source system and field | What to compute |
|---|---|---|
| Intent mix | Ticket category, disposition code, macro used (Zendesk, Salesforce Service Cloud, ServiceNow) | Share of sessions per intent, by channel and segment |
| Disclosure policy | Message-level text with turn index; first turn where order or account ID appears | Probability a slot is volunteered vs elicited, by turn |
| Patience budget | Reopen flag, escalation timestamp, transfer events, CSAT | Turns and elapsed time to escalation or abandonment |
| Corrections | Edited fields, "actually" or "sorry, I meant" turns, reversed actions | Rate and position of self-corrections |
| Policy pressure | Exception requests, supervisor overrides, denial reasons | Frequency of out-of-policy asks and the requester's reaction to refusal |
| Outcome | Resolution code, refund amount bucket, follow-up within 7 days | Ground-truth success label for reward or scoring |
Distributions matter more than any single transcript. A simulator that reproduces individual conversations but gets the intent mix or escalation rate wrong will weight your evaluation toward easy cases. Ask suppliers for aggregate statistics alongside the raw sessions so you can check representativeness before signing.
An illustrative requester-session record for simulator fitting
The most useful delivery format is one record per session with requester turns separated from agent turns and labels attached at session level. The schema below shows the minimum structure an evaluation lead should request; field names are generic and should be mapped to the supplier's export.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"session_id": "s_8f21c",
"channel": "chat",
"segment": {"customer_tenure_bucket": "1-3y", "plan_tier": "business", "region": "US-West"},
"intent_primary": "billing.dispute_charge",
"intent_shift": [{"turn": 5, "to": "account.cancel"}],
"turns": [
{"idx": 1, "role": "requester", "text": "I got charged twice this month??", "slots_disclosed": []},
{"idx": 2, "role": "agent", "text": "[AGENT_TURN]"},
{"idx": 3, "role": "requester", "text": "it's the account under [EMAIL_1]", "slots_disclosed": ["account_ref"]},
{"idx": 5, "role": "requester", "text": "honestly just cancel it", "slots_disclosed": []}
],
"escalated": true,
"escalation_turn": 7,
"resolution_code": "refund_partial",
"reopened_within_7d": false,
"deid": {"method": "entity replacement with typed placeholders", "sample_reviewed": true}
}
Typed placeholders such as [EMAIL_1] preserve the fact that a requester disclosed an identifier at a given turn, which the disclosure policy needs, without carrying the identifier itself.
How to validate a simulator against held-out real sessions
Validation means comparing simulated sessions with held-out real sessions on distributions, not reading a few transcripts and judging them plausible. Split the licensed data by time, not at random, so the held-out slice reflects drift in products and policies. Keep the held-out slice out of prompt construction, persona extraction and few-shot examples.
Simulator validation checklist
- Intent distribution: compare simulated and real intent shares; flag any intent off by more than your tolerance.
- Turn-count distribution: compare histograms of requester turns per session, including the long tail.
- First-turn completeness: share of sessions where the goal and key slots appear in turn one.
- Escalation and abandonment: rate and turn position, by segment.
- Correction rate: frequency of self-corrections and contradictory statements.
- Agent ranking stability: run two or three agent versions against both real replays and the simulator and check whether the ranking agrees.
- Leakage: confirm no held-out session text appears verbatim in simulator prompts or persona cards.
- Memorization: probe the simulator for verbatim requester text, since language models can emit training sequences including contact details [7].
If rankings disagree between real replays and the simulator, the simulator is not yet fit for RL reward shaping, whatever its transcripts look like. Document the comparison in a data card covering sources, collection methods and intended use [10].
Privacy and notice issues specific to requester text
Requester messages are personal data written by customers and employees, so de-identification and notice checks come before modeling. Free text carries names, emails, phone numbers, account numbers, addresses and health or financial details that structured redaction misses. Under the CCPA, information counts as deidentified only if the business also takes reasonable measures against re-association, publicly commits not to re-identify, and contractually binds recipients [8].
Notice matters as much as redaction. The FTC has warned that quietly changing terms to permit AI training on customer data may be unfair or deceptive [9], so ask which privacy notice was in force when each session was collected; see matching records to the notice in force at collection. Voice transcripts raise further questions where recordings may involve biometric identifiers.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
How simulator data differs from other agent data purchases
Simulator data describes the people asking, while most agent data describes the people or systems doing the work. Agent-side demonstrations, exception handling records and cross-system workflow records teach the policy what to do; requester-side sessions teach the environment how to behave. Many buyers need both from the same support operation, which is why the customer support agents use case and RL environments capability pages sit next to this one.
For scoring, pair simulator outputs with real outcome labels; task success labels from business records and golden evaluation datasets built from business records cover that side. The broader evaluation sets use case explains how held-out business work becomes a test set.
How SourceX fits a requester-data request
SourceX sources operational datasets, including support and sales histories, from US companies and manages the licensing process. Data is sourced on request rather than held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Buyers describe the data they need, such as session-level chat logs with disposition codes and escalation events, and SourceX looks for US businesses that hold it; you can describe your requester-data need on the buyers page.
Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. SourceX does not train models and does not source scraped web content.
Get requester-side session data for your user simulator
SourceX sources support and sales histories from US companies on request, with rights review and de-identification before delivery under a license. Nothing is contracted until a supplier agrees. Describe the requester data your simulator needs.
Frequently asked questions
Can I build a user simulator from synthetic personas alone?
You can start with synthetic personas, but without real sessions you cannot measure whether the simulator's intent mix, disclosure behavior and escalation rate match production. Real logs are what you calibrate and validate against [1].
How many real sessions do I need?
There is no fixed number. Size the sample so every intent you plan to evaluate has enough sessions for its turn-count and escalation distributions to be stable, and reserve a time-based held-out slice.
Should the agent turns be included?
Yes, as context. The simulator conditions on what the agent said, so requester turns without agent turns lose the cause of each reaction; label roles clearly so agent text is not used as requester behavior.
Sources
- arXiv, "PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems" (2026). https://arxiv.org/abs/2608.28378
- arXiv, "The Indispensable Role of User Simulation in the Pursuit of AGI (Balog and Zhai)" (2025). https://arxiv.org/pdf/2509.19456v1
- Microsoft Research, "A User Simulator for Task-Completion Dialogues". https://www.microsoft.com/en-us/research/publication/user-simulator-task-completion-dialogues/
- arXiv (Sierra Research), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- arXiv (Chen et al., NAACL 2021), "Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems" (2021). https://arxiv.org/abs/2104.00783v1
- arXiv, "Knowledge-augmented user simulators for training language model assistants (arXiv 2401.16454v1)" (2024). https://arxiv.org/html/2401.16454v1
- USENIX Security 2021 (Carlini et al.), "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
- California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
- Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.