Industry-specific operational data
After-call work notes and disposition codes for call summarization AI
Quick answer
Call summarization training data is a set of contact-center transcripts paired with what the agent actually wrote and selected after the call: the wrap-up note, the disposition code (with its code-set version), and the CRM fields changed. Off-the-shelf corpora rarely include these pairs; they usually sell audio and transcripts alone. Buyers should source the pairing directly from operating contact centers, verify client authorization and payment-card redaction, and treat human notes as noisy references that need QA-reviewed subsets for evaluation.
By SourceX Editorial · Updated
Why transcript-only corpora do not train a summarizer
Transcript-only corpora lack the target side of the task, so they cannot supervise note generation or auto-disposition. As of October 2026, commercial call-center datasets are typically marketed as audio with time-stamped transcripts, with samples available on request [1]. That is the right input for ASR, which our call-center audio dataset page covers, but a summarization model needs the human output tied to each interaction.
Public research data does not close the gap. Academic call-center corpora are often non-commercial; CC BY-NC 4.0 is reported for at least one, so confirm the license on each corpus before any product use [2].
The commercial pull is clear: after-call work is paid handle time on every interaction, and summarizers that draft notes and codes for agents to confirm target it directly. That makes the operational pairing, not more audio, the scarce asset.
What a usable transcript-to-note record contains
A usable record links one interaction ID to the transcript, the agent's free-text note, the selected disposition, and the downstream system changes. Without a stable join key across the ACD or CCaaS platform (Genesys, NICE CXone, Five9, Amazon Connect), the WFM system and the CRM (Salesforce Service Cloud, Zendesk, Dynamics 365), the pairs cannot be rebuilt reliably.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value | Why it matters |
|---|---|---|
| interaction_id | INT-7F3A21 | Join key across telephony, QA and CRM |
| channel / queue | voice / billing_tier2 | Controls domain mix and routing context |
| transcript_turns | speaker-labeled turns with start/end ms | Input side; speaker labels drive attribution in notes |
| asr_engine_version | asr-2025.11 | Lets you control for transcript error drift |
| call_reason_ivr | "billing dispute" | Weak label to compare against the agent's code |
| acw_note | "Cust disputes 3/12 late fee. Waived per policy 4.2. Cb if not reflected." | Summarization target (terse, abbreviated) |
| disposition_code + code_set_version | BIL-WAIVE / v14 | Classification target; version prevents label drift |
| secondary_dispositions | RETENTION-OFFER-DECLINED | Multi-label wrap-up |
| crm_fields_changed | case.status: open→resolved; credit: 15.00 | Structured action target for CRM auto-fill |
| follow_up_created | task: callback, due +2d | Tests whether summaries capture commitments |
| aht_s / acw_s | 412 / 96 | Separates rushed notes from careful ones |
| qa_reviewed / qa_note_score | true / 4 of 5 | Flags gold-quality references for eval |
| redaction_method | PCI pause-resume + NER mask, version recorded | Evidence for privacy and PCI review |
Where human notes and codes mislead models
Agent notes and disposition codes are noisy labels, so a model trained naively on them learns the shortcuts agents take under handle-time pressure. Expect three recurring failure modes, and test for each in any sample before licensing.
- Omission under ACW pressure. Short acw_s values correlate with notes that drop commitments ("will call back Friday") or amounts. Filter or down-weight by acw_s, and score summaries against QA-reviewed notes rather than raw notes. Our QA scorecards page covers the evaluation side.
- Default and catch-all codes. Codes such as "General Inquiry" or the first item in a drop-down are overused. Compare the code distribution against IVR call reason and transcript content, and ask for the code-set history so you can map retired codes.
- Templated text. Macros and canned notes ("Cust called re: acct, issue resolved") inflate apparent consistency and teach boilerplate. Detect and handle them using the approach in templates and canned replies in business records.
Also check the transcript side: ASR errors in names, amounts and dates propagate into generated notes, so request ASR engine version and, where possible, a human-corrected subset.
Rights, recording consent and payment-card redaction
Rights review for this data turns on who owns the interaction and what the callers agreed to, not only on who stores the files. In BPO settings, the outsourcer typically processes calls on behalf of its clients, so reuse usually depends on the client's contract; see client data held by service providers. Agent-written notes raise separate employee-authorship and notice questions covered in employee-authored records in training data.
Recording consent varies by state: California Penal Code 632 requires consent of all parties to record a confidential communication [3]. Ask for the disclosure script and the states covered, and confirm that the consent language does not restrict secondary use.
Payment calls need specific evidence. Card numbers spoken on calls and captured in recordings, transcripts or notes bring that data into PCI DSS scope, so the cleanest position is that it never reaches the licensed corpus. Require proof of pause-and-resume recording, DTMF masking or transcript redaction, plus a spot-check of notes, because agents sometimes type card fragments into free text. Calls into health plans or providers carry PHI and need HIPAA de-identification via Safe Harbor or Expert Determination [4]; see HIPAA-compliant AI training data.
Supplier qualification checklist
Use this checklist before signing, and ask the supplier to answer it in a dataset documentation sheet modeled on Data Cards, which cover sources, collection and annotation methods, and intended use [5].
Illustrative example: invented to show structure; it does not describe an available dataset.
- Interaction ID joins transcript, note, disposition and CRM changes for at least a test sample.
- Disposition code set provided with version history, definitions and retired-code mappings.
- Queue, line-of-business and channel mix stated, including chat versus voice share.
- acw_s and aht_s included so rushed notes can be filtered.
- QA-reviewed subset identified, with the scoring rubric.
- Macro or template usage flagged per note.
- Redaction method named and versioned for transcripts and notes, with a sample audit result.
- PCI handling evidence for payment queues (pause-resume, DTMF masking or redaction).
- Written client authorization where the supplier is a BPO.
- Recording-consent disclosure and states covered.
- Allowed uses stated: SFT, classifier training, evaluation, or all three.
Designing evaluation splits for generated notes
Evaluation for call summarization should use held-out, QA-reviewed notes split by time and by client or queue, not random splits. Random splits leak macros, recurring customers and agent writing habits across train and test, which inflates ROUGE and BERTScore.
Score generated notes on factual consistency with the transcript (amounts, dates, commitments), coverage of follow-ups recorded in follow_up_created, and agreement with crm_fields_changed. For auto-disposition, report per-code precision and recall and a confusion matrix across adjacent codes, since overall accuracy hides catch-all code inflation. Hold back a slice from a later code-set version to test robustness to taxonomy changes.
Adjacent industry guides use the same note-plus-outcome pattern: debt collection conversation data, FNOL intake conversations and provider-to-payer calls. For the full map of operational records, start at the industry-specific operational data hub or the AI data buyer hub. If you need the audio side specified too, use the contact-center audio requirements template.
How SourceX handles requests for this data
SourceX sources operational datasets from US companies on request, including support and sales histories, and manages licensing and ongoing purchases; it does not hold data in stock, and a request does not guarantee a match. You describe the data you need, not the businesses, and every release is approved by the supplying company after rights review of ownership and consents. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Contact-center buyers can also review buyers in BPO and contact centers, customer support buyers and customer support transcripts, or describe your requirement to SourceX.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Request transcript-to-note training data
Describe the transcripts, wrap-up notes, disposition codes and CRM fields you need, and SourceX will look for US businesses that hold them. Nothing is contracted until a supplier agrees, and each dataset is delivered under a license that defines records, uses, term and delivery. Start a buyer request.
Sources
- Unidata, "Call Center Audio Dataset". https://unidata.pro/datasets/call-center-audio
- arXiv, "Call-center conversation corpus paper (arXiv 2507.02958)" (2025). https://web3.arxiv.org/pdf/2507.02958
- California Legislative Information, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN§ionNum=632
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- arXiv (Pushkarna, Zaldivar, Kjartansson), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.