Evaluation and benchmarking datasets
Summarization evaluation sets: human-written summaries paired with real sources
Quick answer
A summarization evaluation dataset for business use should pair real source material (meeting transcripts, support calls, case files) with summaries that people wrote to do their jobs: minutes, agent wrap-up notes, closing memos. Those natural pairs show what your users actually need preserved. Treat each human summary as a candidate reference, not ground truth: screen it, convert it into required and forbidden claims, and score model outputs for coverage and faithfulness at the claim level, with consent and PII handling settled before any transcript moves.
By SourceX Editorial · Updated
Why natural pairs beat commissioned benchmark summaries
Natural pairs encode the decisions, owners and dates that a downstream reader relied on, which commissioned summaries rarely capture. Public benchmarks were built for research: Open4Business, for example, is an open-access dataset for summarizing business documents [1]. TWEETSUMM covers customer-service dialog summarization, but from chat conversations rather than phone calls [2]. Both are useful for method comparison, yet neither reflects your call taxonomy, your CRM fields or the house style your reviewers expect.
A natural pair also carries an implicit relevance judgment. When a project manager's minutes omit a ten-minute digression and record a single action item, that omission is a signal about what matters. Annotators hired for a benchmark see the transcript cold and tend to summarize evenly, which inflates scores for models that produce balanced but unhelpful prose.
The catch is that natural summaries were never written to be complete or literal. They abbreviate, assume context, and sometimes include facts the writer knew but the source never said. The rest of this guide is about turning that raw material into a defensible reference set.
Where source-summary pairs already exist in business records
Most operating companies already produce summary pairs as a side effect of normal workflows; the work is finding the join key between source and summary. The table below lists common pair types and the failure modes to check for before using them.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Pair type | Source artifact | Human summary | Typical join key | Main risk as a reference |
|---|---|---|---|---|
| Meeting to minutes | Zoom or Teams transcript (VTT), audio | Minutes doc, action-item list | Calendar event ID, meeting date | Minutes include decisions made after the call |
| Support call to wrap-up note | Call recording plus ASR transcript | Agent disposition and notes in Salesforce Service Cloud or Zendesk | Ticket or interaction ID | Notes are terse, templated, written under handle-time pressure |
| Sales call to CRM note | Gong or similar call recording | Opportunity next-step field, call summary | Activity ID linked to opportunity | Rep optimism; notes record intent, not what the buyer said |
| Case file to closing memo | Claims, legal or HR case documents | Closing or disposition memo | Case number | Memo draws on documents outside the delivered file |
| Incident to postmortem | Slack channel export, pager timeline | Postmortem summary section | Incident ID | Root cause written after investigation, not from the channel |
| Thread to handoff note | Email or ticket thread | Shift handoff or escalation summary | Thread ID, escalation timestamp | Handoff covers several threads at once |
Ask suppliers which system holds each side and whether the join key survives export. A wrap-up note without a reliable interaction ID is useless, because you cannot prove which transcript it summarizes. For meeting and call sources specifically, see the owner pages for licensed meeting transcripts and licensed sales call recordings.
Screening human summaries before they become references
Reference quality varies widely, so every human summary should pass a screening step before it is allowed to score a model. Market practice treats expert-validated references as the standard for golden sets, which implies a validation pass rather than blind trust in whatever the business wrote [4]. ISO/IEC 5259-4 frames this as a data quality process for evaluation data, including how labels are produced and checked [7].
A practical screen has four checks. First, grounding: every statement in the summary must trace to the source; flag anything that does not as "external knowledge" and either strip it or attach the supporting document. Second, completeness against task: if your feature promises action items, a summary with no action items is a bad reference unless the meeting genuinely had none.
Third, style and templating: disposition codes and boilerplate ("Customer called re: billing, resolved") carry little signal and should be kept for a separate classification eval. Fourth, reviewer agreement: have a second domain reviewer mark which source spans the summary covers, and drop or repair pairs where the two disagree on key content.
Expect attrition. Natural pairs need no commissioned writing, but a meaningful share will fail grounding, especially closing memos and postmortems written after later investigation. Budget for repair rather than assuming every pair is usable.
Scoring with claims instead of n-gram overlap
Claim-level scoring separates coverage (did the model include what matters) from faithfulness (did it say anything unsupported), which overlap metrics conflate. One practical method expresses ground truth as required and forbidden claims checked against each output, which works even when no single reference answer exists [3]. A human-written minutes document is a good seed for required claims; the source transcript is the only arbiter of faithfulness.
Factual-consistency research gives you a vocabulary for the forbidden side. Useful error categories include wrong entities (the wrong customer or product), wrong predicates ("refunded" instead of "credited"), wrong circumstances (dates, amounts, locations) and unsupported additions. Sentence-level natural language inference checkers, which compare each summary sentence against source passages, are a common automated first pass. Calibrate any checker on your own domain pairs before trusting it, because legal and support language confuses general-purpose models.
If you use an LLM judge to apply the claim checks, validate it against human labels first; see the guide to LLM-as-a-judge calibration sets. For answers grounded in retrieved passages rather than a single source, the related page on faithfulness evaluation sets with grounded and unsupported labels covers labeling at the answer level.
An illustrative evaluation record
Each eval item should keep the source, the original human summary, the derived claims and the provenance in one record, so reviewers can audit any score. The schema below is a starting point for a call-summary eval.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"item_id": "callsum-000417",
"source": {
"type": "support_call_transcript",
"format": "vtt",
"duration_sec": 742,
"asr_engine": "vendor_asr_v3",
"speaker_labels": ["agent", "customer"],
"pii_method": "names, phones, account numbers replaced with typed placeholders"
},
"human_summary": {
"origin": "agent_wrap_up_note",
"written_within_min": 4,
"text": "Cust [PERSON_1] disputed late fee on [ACCOUNT_1]. Waived one-time. Sent autopay enrollment link."
},
"screen": {
"grounded": true,
"external_knowledge_spans": [],
"second_reviewer_agreement": "full"
},
"required_claims": [
"Customer disputed a late fee",
"Agent waived the fee as a one-time exception",
"Agent sent an autopay enrollment link"
],
"forbidden_claims": [
"Fee was refunded to a card",
"Customer agreed to enroll in autopay during the call"
],
"strata": {"intent": "billing_dispute", "length_bucket": "10-15min", "resolution": "resolved"},
"provenance": {"system": "contact_center_platform", "export_date": "2026-09-30", "license_ref": "LIC-REF"}
}
Note the second forbidden claim: the transcript shows the link was sent, not that the customer enrolled. That gap between action and outcome is one of the most common hallucinations in call summaries, and it belongs in your forbidden set deliberately.
Consent, PII and recording law for transcript sources
Transcripts are the riskiest half of each pair, because they capture third parties who never signed anything with you. Recording law varies by state, and several states, California among them, require consent from all parties to record certain conversations, so ask how each recording was noticed and whether the consent covers secondary use for model evaluation. Confirm with counsel how those consents apply to your use.
De-identification must cover both sides of the pair. A wrap-up note can contain the same account number as the transcript, and a summary that keeps a name the transcript scrubbed defeats the exercise. Placeholder types should be consistent across source and summary so that claims like "[PERSON_1] disputed the fee" still score correctly.
Health-related sources are a separate tier. HIPAA de-identification relies on either Expert Determination or Safe Harbor, the latter requiring removal of 18 listed identifiers [5]. Clinical visit notes paired with transcripts are attractive summarization references, but they are not usable without one of those methods documented.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Building and protecting the set
A summarization eval set is only informative if it is stratified, held out and documented. Stratify by meeting or call type, length bucket (long transcripts stress context handling differently) and outcome, then cap any single team or agent so one writer's style does not dominate. For long inputs, the guide to long-context evaluation with real document sets covers length allocation.
Keep the references out of training. If the same company's minutes feed a fine-tuning run, split by source system or time window, and follow the practices in contamination-resistant evaluation design. Training pairs themselves are covered separately in summarization fine-tuning data from business records.
Document the set with a Data Card that records upstream sources, how summaries were written and screened, intended use and known gaps [6]. For voice products, pair this eval with outcome-based tests from voice agent evaluation sets. The wider map of options sits in the evaluation datasets hub, and rights questions are covered in the data provenance guide.
What to ask a supplier of summary pairs
A short request template keeps conversations concrete and makes non-matches obvious early.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Pair type and systems: for example, "Teams transcripts (VTT) joined to minutes in SharePoint by calendar event ID."
- Who wrote the summaries, in what role, and how long after the source event.
- Whether summaries reference documents outside the delivered source.
- Volume by stratum (call intent, meeting type, length bucket), not just a total.
- Recording notice and consent basis for each source channel.
- De-identification method, applied to both source and summary, and how it was sample-checked.
- Allowed uses: evaluation only, or evaluation plus training, and any limits on publishing scores.
- Documentation available: source description, collection method, known quality issues.
SourceX sources operational datasets, including support and sales histories and documents, from US companies on request; categories are not inventory and a request does not guarantee a match. You can describe the summary pairs you need to SourceX using the template above. For broader context on eval sets from business work, see AI evaluation datasets built from real business work.
Request summarization evaluation pairs from real business work
SourceX looks for US businesses that hold the source-summary pairs you describe, rights-reviews each dataset for ownership and consents, and delivers under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Describe your summarization evaluation needs.
Sources
- arXiv, "Open4Business(O4B): An Open Access Dataset for Summarizing Business Documents" (2020). https://arxiv.org/pdf/2011.07636
- arXiv (Feigenblat et al., Findings of EMNLP 2021), "TWEETSUMM - A Dialog Summarization Dataset for Customer Service" (2021). https://arxiv.org/abs/2111.11894v1
- OneUptime, "How to Build Ground Truth for RAG Evaluation When No Reference Answers Exist" (2026). https://oneuptime.com/blog/post/2026-08-31-build-rag-ground-truth-without-reference-answers/markdown
- Sigma AI, "Golden datasets: Evaluating fine-tuned large language models". https://sigma.ai/?p=26358
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- arXiv (Pushkarna, Zaldivar, Kjartansson, Google Research; FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Artificial intelligence - Data quality for analytics and machine learning - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.