Skip to content

Agent, workflow and domain-reasoning data

Agent-to-human handoff data: when to escalate and what context to pass

Quick answer

Agent-to-human handoff data is the record of every point where an automated agent, or a tier 1 person, passed a case to someone more capable: the trigger, the context passed, what the receiver had to re-ask, and how the case ended. To train escalation policies and handoff summarizers, buyers need these four parts linked per case, plus labeled missed escalations. Public dialogue corpora rarely contain them, so most usable data comes from licensed support, claims and service-desk histories.

By SourceX Editorial · Updated

Why public "handoff" datasets rarely fit escalation work

Most datasets labeled "handoff" or "handover" describe physical object transfer between robots and people, not conversational escalation, so a keyword search mostly returns the wrong modality. Research corpora for service agents, such as the Action-Based Conversations Dataset (ABCD), model agents acting under company guidelines [3], and τ-bench tests tool-using agents against policy documents and simulated users [2]. Neither is built around the moment an agent gives up and a human takes over, nor the human's work afterward; where a benchmark offers a transfer action, it scores the task outcome, not the quality of the handoff. If your query is "agent human handoff dataset", rephrase internal requests as "chatbot escalation to human agent records" or "AI agent escalation training data" to avoid sourcing the robotics kind.

The richest source is operational: contact-center platforms, ITSM queues and claims systems that already log transfers. Platforms instrument this event explicitly: Open.cx, for example, logs each AI-to-human handoff with a reason, sentiment and the articles or tools referenced [1]. Those logs are the skeleton; the value is in joining them to what happened next.

What a usable escalation record contains

A usable record links five things per case: pre-handoff state, the trigger, the handoff payload, the receiver's actions, and the outcome. Missing any one of them limits what you can train. Without pre-handoff state you cannot learn when; without receiver actions you cannot grade the summary; without outcomes you cannot tell a good escalation from a needless one.

  • Pre-handoff state. Full transcript or ticket thread up to the transfer, timestamps, channel, intent or category, tools called and their results, and any confidence or retrieval scores the agent logged.
  • Trigger. The reason code (customer request, low confidence, policy boundary, missing tool or permission, risk signal such as legal threat or vulnerability), who or what fired it, and whether it was rule-based or model-decided.
  • Handoff payload. The summary, notes or metadata passed to the receiver, including queue, skill and priority routing fields. On bot platforms this is often a free-form metadata object defined by the implementing team; in ticketing systems it is the internal note and field changes at reassignment.
  • Receiver actions. The first messages and actions after transfer, especially questions that repeat what the customer already said, plus any re-transfers (tier 2 to tier 3, or back to the bot).
  • Outcome. Resolution code, reopen within N days, CSAT or sentiment change, refund or exception granted, and handle time after transfer.

Escalation triggers: which signals to label and how

Label triggers with a closed taxonomy so a model can learn a policy, not just imitate one team's habits. Free-text reasons from agents ("cust upset") are useful as raw evidence but need mapping. A practical starting taxonomy has five families, each with its own failure mode.

Trigger familyTypical evidence in recordsCommon labeling failure
Explicit request"agent", "human", "representative" utterances; IVR zero-outCounting repeated requests in one session as separate events
Low confidenceRetrieval score, intent confidence, repeated fallback intentsThresholds changed mid-period without a version field
Policy boundaryRefund above limit, account change needing verificationPolicy documents not versioned with the case date
Missing capabilityTool error, no API for the action, knowledge gapLogged as "other"; Open.cx-style reason fields help [1]
Risk signalLegal threat, safety, vulnerability, fraud indicatorsUnder-labeled because agents escalate silently via phone

Ask for the policy documents in force at the case date. τ-bench shows how much agent behavior depends on the written domain policy [2]; an escalation label is only interpretable against the rule that applied at that time.

Handoff summaries: grading quality by what the receiver had to re-ask

A handoff summary is good if the receiver did not need to ask for anything the sender already knew. That gives you a measurable label without subjective rating: diff the facts in the pre-handoff transcript against the receiver's first questions. Every re-asked fact (order number, error message, steps already tried) is a summary defect.

Request records where both the summary and the receiver's first five turns are present. Then build pairs: the transcript as input, the summary as a candidate, and re-asks as negative signal. Pairs where a senior agent rewrote the note are rare but valuable reference summaries. Keep summary style consistent across sources, or the model learns team formatting instead of content selection.

Missed escalations and needless transfers: the rare, decisive labels

The most important labels are the cases where the agent should have escalated and did not, and the cases where it escalated when it could have resolved. Both are scarce in logs because nothing explicitly marks them. You find them through outcomes, not through the transfer event.

Proxy signals for missed escalations: the customer re-contacts on another channel within 48 to 72 hours, a complaint or chargeback follows a bot-only session, or QA reviewers flagged the session. Proxy signals for needless transfers: the human resolved with a knowledge article the bot had retrieved, or handle time after transfer was under a minute. Ask suppliers whether QA scorecards exist; reviewer flags are the cleanest ground truth. See long-tail and edge-case coverage for how to quantify how many rare cases you actually need.

Specification template for escalation and handoff records

Write the request around fields and joins, not around a vendor name, so suppliers can check whether their systems hold it. The template below shows one record as a single JSON Lines row: UTF-8, one JSON object per line [6].

Illustrative example: invented to show structure; it does not describe an available dataset.

{"case_id": "c_000184", "channel": "chat", "domain": "billing",
 "pre_handoff": {"turns": 9, "bot_intents": ["refund_status", "fallback", "fallback"],
   "tool_calls": [{"name": "get_invoice", "status": "ok"}, {"name": "issue_refund", "status": "denied_policy"}],
   "max_retrieval_score": 0.41},
 "trigger": {"family": "policy_boundary", "code": "refund_over_limit", "fired_by": "rule", "policy_version": "2025-11"},
 "handoff": {"from": "bot", "to": "tier2_billing", "summary": "Customer disputes duplicate charge; refund over bot limit.",
   "routing": {"queue": "billing_t2", "priority": "normal"}},
 "receiver": {"first_turns": 5, "re_asked_fields": ["invoice_number"], "retransfer": false},
 "outcome": {"resolution": "refund_issued", "reopened_7d": false, "csat": 4, "post_handoff_minutes": 11},
 "labels": {"escalation_correct": true, "summary_defects": 1, "qa_flag": null},
 "pii_method": "names, emails, account numbers replaced with tokens"}

Checklist to send with the request:

  1. Systems: which platform logged the transfer (CCaaS, ticketing, ITSM, claims), and can sessions be joined to tickets by ID.
  2. Coverage: share of escalations by trigger family, channels, languages, and months covered; whether policy changes fall inside the window.
  3. Payload: whether summaries, internal notes and routing fields at reassignment are preserved, not overwritten.
  4. Outcomes: resolution codes, reopen windows, CSAT, QA scorecards.
  5. Negatives: bot-only sessions with outcomes, so you can mine missed escalations.
  6. Privacy: de-identification method for transcripts and free text.
  7. Documentation: a data card covering sources, collection, annotation and intended use [5].

For how these records sit beside other agent data, see ticket histories reconstructed as agent trajectories and policy-following service agent data.

Privacy, rights and domain constraints on escalation transcripts

Escalation transcripts are among the most identifier-dense records a company holds, because escalations concentrate on account, payment and complaint issues. Names, emails, phone numbers and account numbers appear in free text and in summaries, so de-identification must cover notes and metadata, not just structured fields. Health-plan and provider escalations add stricter rules: health records need HIPAA de-identification by Safe Harbor, which removes 18 listed identifiers, or Expert Determination [4].

Ask how voice escalations were transcribed and whether audio is in scope, since recordings carry voice as an identifier. Confirm the supplier holds rights to use customer conversations for this purpose, and that the license covers the derived artifacts you plan (summaries, labels, evaluation sets). Our license terms guide for agent data and the privacy hub cover these points in depth.

Where escalation records come from and how SourceX sources them

Escalation records sit in ordinary operations: support desks, service centers, insurance claims and IT help desks. SourceX's workflow pages show what these look like in customer support escalation and insurance service escalation. SourceX sources operational datasets, including support and sales histories, from US companies on request; it does not hold escalation data in stock, and a request does not guarantee a match.

Buyers describe the data they need, not the companies, and SourceX looks for US businesses that hold it. Each dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phones and account numbers are removed or replaced before delivery with the method recorded and a sample checked, and no method is perfect. Nothing is contracted until the supplier agrees, and every release is approved by the supplying company. You can describe your escalation data need to SourceX using the checklist above.

Source escalation and handoff records for your agents

SourceX sources operational records such as support histories from US companies on request and manages the license that defines records, uses, term and delivery. Delivery runs through private, access-controlled workflows after an executed agreement and supplier approval. Start in the agent data hub or submit your escalation data request.

Frequently asked questions

Can synthetic conversations replace real escalation records?

Synthetic dialogues can cover trigger families you already understand, but they cannot tell you how often real customers hit a policy boundary or what receivers actually re-ask. Use real records to set the distribution and labels, then augment. The trade-offs are covered in licensed, commissioned or synthetic trajectories.

Are tier 1 to tier 2 transfers useful if there is no bot involved?

Yes. Human-to-human transfers carry the same trigger, payload and outcome structure, and the reasons experienced tier 1 staff escalate are close to what you want an agent to learn. Mark the sender type so you can weight or separate them.

How should escalation data be split for evaluation?

Split by time and by policy version, not randomly, so the test set reflects policies the model has not seen applied. Hold out a set of QA-flagged missed escalations as a dedicated evaluation slice; see task success labels.

Sources

  1. Open.cx, "Handoff analytics (API reference)". https://docs.open.cx/api-reference/handoff-analytics
  2. Sierra Research, "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  3. Chen et al., "Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems" (2021). https://arxiv.org/abs/2104.00783v1
  4. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  5. Pushkarna, Zaldivar, Kjartansson, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  6. jsonlines.org, "JSON Lines". https://jsonlines.org/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data