Skip to content

Agent, workflow and domain-reasoning data

User simulator training data: what real requester behavior agents need

Quick answer

A credible user simulator is calibrated on the requester side of real conversations: what people wanted, what they disclosed and when, how they corrected or abandoned the exchange, and how often each pattern occurs. Buy or license de-identified session logs with intent labels, turn-level disclosure, outcome codes and segment metadata, then hold out a slice of real sessions to test whether simulated users reproduce the real distributions before you trust any RL reward or evaluation score.

By SourceX Editorial · Updated

This page covers simulator building. For scenario design on the evaluation side, see user simulator scenarios for agent evaluation; for agent-side demonstrations and workflow records, start at the agent training data hub.

Why agent teams need real requester data, not just agent transcripts

Simulators exist because live users cannot be put in every RL rollout or regression run, and the simulated user is only as realistic as the data it was fitted to. The problem is old: Microsoft's task-completion simulator was built from example dialogues precisely because collecting fresh human dialogues for every policy iteration was too expensive [3]. Modern benchmarks such as tau-bench put an LLM-simulated customer in the loop with tool APIs and domain policies, so the simulated user's behavior directly shapes the measured pass rate [4].

The failure mode is a simulator that is too cooperative. It states its full goal in turn one, never contradicts itself, and accepts the first plausible answer. The PersonaForge authors argue that most agent training data assumes complete single-turn queries, while 75.9% of the 16K real sessions they analyzed were multi-turn [1]. An agent tuned against an over-helpful simulator learns to skip clarification, and its offline scores overstate production performance.

What a simulated user has to model

A simulated user needs a hidden goal, a disclosure policy, a patience budget and a style, each estimated from real sessions. Balog and Zhai frame user simulation as a bottleneck for both evaluation and training data and argue it should be parameterized by user type, such as novice versus expert [2]. In practice, the parameters below are the ones agent teams most often under-specify.

  • Goal and sub-goals. The underlying task (refund, plan change, access request, purchase order status), including goals that change mid-conversation.
  • Information disclosure over turns. Which facts the requester volunteers, which only appear when asked, and which are never provided (order number missing, wrong account referenced).
  • Ambiguity and error. Vague first messages, wrong product names, mistaken dates and self-corrections.
  • Patience and escalation. Turns or minutes before the requester asks for a human, repeats themselves, or abandons.
  • Policy pressure. Requests for exceptions the agent must refuse; the ABCD corpus is useful here because agent actions are bound by company policies rather than just slot filling [5].
  • Style and channel. Message length, typos, multi-question messages, email versus chat versus voice transcript conventions.

Which records calibrate those parameters

Each simulator parameter maps to a specific field in operational systems, which is why requester-side records from ticketing and messaging tools are the core input. Knowledge-grounded simulator research shows that tying simulated users to external facts increases diversity and realism [6], and real account and order context plays that role in enterprise settings.

Simulator parameterSource system and fieldWhat to compute
Intent mixTicket category, disposition code, macro used (Zendesk, Salesforce Service Cloud, ServiceNow)Share of sessions per intent, by channel and segment
Disclosure policyMessage-level text with turn index; first turn where order or account ID appearsProbability a slot is volunteered vs elicited, by turn
Patience budgetReopen flag, escalation timestamp, transfer events, CSATTurns and elapsed time to escalation or abandonment
CorrectionsEdited fields, "actually" or "sorry, I meant" turns, reversed actionsRate and position of self-corrections
Policy pressureException requests, supervisor overrides, denial reasonsFrequency of out-of-policy asks and the requester's reaction to refusal
OutcomeResolution code, refund amount bucket, follow-up within 7 daysGround-truth success label for reward or scoring

Distributions matter more than any single transcript. A simulator that reproduces individual conversations but gets the intent mix or escalation rate wrong will weight your evaluation toward easy cases. Ask suppliers for aggregate statistics alongside the raw sessions so you can check representativeness before signing.

An illustrative requester-session record for simulator fitting

The most useful delivery format is one record per session with requester turns separated from agent turns and labels attached at session level. The schema below shows the minimum structure an evaluation lead should request; field names are generic and should be mapped to the supplier's export.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "session_id": "s_8f21c",
  "channel": "chat",
  "segment": {"customer_tenure_bucket": "1-3y", "plan_tier": "business", "region": "US-West"},
  "intent_primary": "billing.dispute_charge",
  "intent_shift": [{"turn": 5, "to": "account.cancel"}],
  "turns": [
    {"idx": 1, "role": "requester", "text": "I got charged twice this month??", "slots_disclosed": []},
    {"idx": 2, "role": "agent", "text": "[AGENT_TURN]"},
    {"idx": 3, "role": "requester", "text": "it's the account under [EMAIL_1]", "slots_disclosed": ["account_ref"]},
    {"idx": 5, "role": "requester", "text": "honestly just cancel it", "slots_disclosed": []}
  ],
  "escalated": true,
  "escalation_turn": 7,
  "resolution_code": "refund_partial",
  "reopened_within_7d": false,
  "deid": {"method": "entity replacement with typed placeholders", "sample_reviewed": true}
}

Typed placeholders such as [EMAIL_1] preserve the fact that a requester disclosed an identifier at a given turn, which the disclosure policy needs, without carrying the identifier itself.

How to validate a simulator against held-out real sessions

Validation means comparing simulated sessions with held-out real sessions on distributions, not reading a few transcripts and judging them plausible. Split the licensed data by time, not at random, so the held-out slice reflects drift in products and policies. Keep the held-out slice out of prompt construction, persona extraction and few-shot examples.

Simulator validation checklist

  1. Intent distribution: compare simulated and real intent shares; flag any intent off by more than your tolerance.
  2. Turn-count distribution: compare histograms of requester turns per session, including the long tail.
  3. First-turn completeness: share of sessions where the goal and key slots appear in turn one.
  4. Escalation and abandonment: rate and turn position, by segment.
  5. Correction rate: frequency of self-corrections and contradictory statements.
  6. Agent ranking stability: run two or three agent versions against both real replays and the simulator and check whether the ranking agrees.
  7. Leakage: confirm no held-out session text appears verbatim in simulator prompts or persona cards.
  8. Memorization: probe the simulator for verbatim requester text, since language models can emit training sequences including contact details [7].

If rankings disagree between real replays and the simulator, the simulator is not yet fit for RL reward shaping, whatever its transcripts look like. Document the comparison in a data card covering sources, collection methods and intended use [10].

Privacy and notice issues specific to requester text

Requester messages are personal data written by customers and employees, so de-identification and notice checks come before modeling. Free text carries names, emails, phone numbers, account numbers, addresses and health or financial details that structured redaction misses. Under the CCPA, information counts as deidentified only if the business also takes reasonable measures against re-association, publicly commits not to re-identify, and contractually binds recipients [8].

Notice matters as much as redaction. The FTC has warned that quietly changing terms to permit AI training on customer data may be unfair or deceptive [9], so ask which privacy notice was in force when each session was collected; see matching records to the notice in force at collection. Voice transcripts raise further questions where recordings may involve biometric identifiers.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

How simulator data differs from other agent data purchases

Simulator data describes the people asking, while most agent data describes the people or systems doing the work. Agent-side demonstrations, exception handling records and cross-system workflow records teach the policy what to do; requester-side sessions teach the environment how to behave. Many buyers need both from the same support operation, which is why the customer support agents use case and RL environments capability pages sit next to this one.

For scoring, pair simulator outputs with real outcome labels; task success labels from business records and golden evaluation datasets built from business records cover that side. The broader evaluation sets use case explains how held-out business work becomes a test set.

How SourceX fits a requester-data request

SourceX sources operational datasets, including support and sales histories, from US companies and manages the licensing process. Data is sourced on request rather than held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Buyers describe the data they need, such as session-level chat logs with disposition codes and escalation events, and SourceX looks for US businesses that hold it; you can describe your requester-data need on the buyers page.

Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. SourceX does not train models and does not source scraped web content.

Get requester-side session data for your user simulator

SourceX sources support and sales histories from US companies on request, with rights review and de-identification before delivery under a license. Nothing is contracted until a supplier agrees. Describe the requester data your simulator needs.

Frequently asked questions

Can I build a user simulator from synthetic personas alone?

You can start with synthetic personas, but without real sessions you cannot measure whether the simulator's intent mix, disclosure behavior and escalation rate match production. Real logs are what you calibrate and validate against [1].

How many real sessions do I need?

There is no fixed number. Size the sample so every intent you plan to evaluate has enough sessions for its turn-count and escalation distributions to be stable, and reserve a time-based held-out slice.

Should the agent turns be included?

Yes, as context. The simulator conditions on what the agent said, so requester turns without agent turns lose the cause of each reaction; label roles clearly so agent text is not used as requester behavior.

Sources

  1. arXiv, "PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems" (2026). https://arxiv.org/abs/2608.28378
  2. arXiv, "The Indispensable Role of User Simulation in the Pursuit of AGI (Balog and Zhai)" (2025). https://arxiv.org/pdf/2509.19456v1
  3. Microsoft Research, "A User Simulator for Task-Completion Dialogues". https://www.microsoft.com/en-us/research/publication/user-simulator-task-completion-dialogues/
  4. arXiv (Sierra Research), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  5. arXiv (Chen et al., NAACL 2021), "Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems" (2021). https://arxiv.org/abs/2104.00783v1
  6. arXiv, "Knowledge-augmented user simulators for training language model assistants (arXiv 2401.16454v1)" (2024). https://arxiv.org/html/2401.16454v1
  7. USENIX Security 2021 (Carlini et al.), "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
  8. California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  9. Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  10. Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data