Skip to content

Evaluation and benchmarking datasets

Summarization evaluation sets: human-written summaries paired with real sources

Quick answer

A summarization evaluation dataset for business use should pair real source material (meeting transcripts, support calls, case files) with summaries that people wrote to do their jobs: minutes, agent wrap-up notes, closing memos. Those natural pairs show what your users actually need preserved. Treat each human summary as a candidate reference, not ground truth: screen it, convert it into required and forbidden claims, and score model outputs for coverage and faithfulness at the claim level, with consent and PII handling settled before any transcript moves.

By SourceX Editorial · Updated

Why natural pairs beat commissioned benchmark summaries

Natural pairs encode the decisions, owners and dates that a downstream reader relied on, which commissioned summaries rarely capture. Public benchmarks were built for research: Open4Business, for example, is an open-access dataset for summarizing business documents [1]. TWEETSUMM covers customer-service dialog summarization, but from chat conversations rather than phone calls [2]. Both are useful for method comparison, yet neither reflects your call taxonomy, your CRM fields or the house style your reviewers expect.

A natural pair also carries an implicit relevance judgment. When a project manager's minutes omit a ten-minute digression and record a single action item, that omission is a signal about what matters. Annotators hired for a benchmark see the transcript cold and tend to summarize evenly, which inflates scores for models that produce balanced but unhelpful prose.

The catch is that natural summaries were never written to be complete or literal. They abbreviate, assume context, and sometimes include facts the writer knew but the source never said. The rest of this guide is about turning that raw material into a defensible reference set.

Where source-summary pairs already exist in business records

Most operating companies already produce summary pairs as a side effect of normal workflows; the work is finding the join key between source and summary. The table below lists common pair types and the failure modes to check for before using them.

Illustrative example: invented to show structure; it does not describe an available dataset.

Pair typeSource artifactHuman summaryTypical join keyMain risk as a reference
Meeting to minutesZoom or Teams transcript (VTT), audioMinutes doc, action-item listCalendar event ID, meeting dateMinutes include decisions made after the call
Support call to wrap-up noteCall recording plus ASR transcriptAgent disposition and notes in Salesforce Service Cloud or ZendeskTicket or interaction IDNotes are terse, templated, written under handle-time pressure
Sales call to CRM noteGong or similar call recordingOpportunity next-step field, call summaryActivity ID linked to opportunityRep optimism; notes record intent, not what the buyer said
Case file to closing memoClaims, legal or HR case documentsClosing or disposition memoCase numberMemo draws on documents outside the delivered file
Incident to postmortemSlack channel export, pager timelinePostmortem summary sectionIncident IDRoot cause written after investigation, not from the channel
Thread to handoff noteEmail or ticket threadShift handoff or escalation summaryThread ID, escalation timestampHandoff covers several threads at once

Ask suppliers which system holds each side and whether the join key survives export. A wrap-up note without a reliable interaction ID is useless, because you cannot prove which transcript it summarizes. For meeting and call sources specifically, see the owner pages for licensed meeting transcripts and licensed sales call recordings.

Screening human summaries before they become references

Reference quality varies widely, so every human summary should pass a screening step before it is allowed to score a model. Market practice treats expert-validated references as the standard for golden sets, which implies a validation pass rather than blind trust in whatever the business wrote [4]. ISO/IEC 5259-4 frames this as a data quality process for evaluation data, including how labels are produced and checked [7].

A practical screen has four checks. First, grounding: every statement in the summary must trace to the source; flag anything that does not as "external knowledge" and either strip it or attach the supporting document. Second, completeness against task: if your feature promises action items, a summary with no action items is a bad reference unless the meeting genuinely had none.

Third, style and templating: disposition codes and boilerplate ("Customer called re: billing, resolved") carry little signal and should be kept for a separate classification eval. Fourth, reviewer agreement: have a second domain reviewer mark which source spans the summary covers, and drop or repair pairs where the two disagree on key content.

Expect attrition. Natural pairs need no commissioned writing, but a meaningful share will fail grounding, especially closing memos and postmortems written after later investigation. Budget for repair rather than assuming every pair is usable.

Scoring with claims instead of n-gram overlap

Claim-level scoring separates coverage (did the model include what matters) from faithfulness (did it say anything unsupported), which overlap metrics conflate. One practical method expresses ground truth as required and forbidden claims checked against each output, which works even when no single reference answer exists [3]. A human-written minutes document is a good seed for required claims; the source transcript is the only arbiter of faithfulness.

Factual-consistency research gives you a vocabulary for the forbidden side. Useful error categories include wrong entities (the wrong customer or product), wrong predicates ("refunded" instead of "credited"), wrong circumstances (dates, amounts, locations) and unsupported additions. Sentence-level natural language inference checkers, which compare each summary sentence against source passages, are a common automated first pass. Calibrate any checker on your own domain pairs before trusting it, because legal and support language confuses general-purpose models.

If you use an LLM judge to apply the claim checks, validate it against human labels first; see the guide to LLM-as-a-judge calibration sets. For answers grounded in retrieved passages rather than a single source, the related page on faithfulness evaluation sets with grounded and unsupported labels covers labeling at the answer level.

An illustrative evaluation record

Each eval item should keep the source, the original human summary, the derived claims and the provenance in one record, so reviewers can audit any score. The schema below is a starting point for a call-summary eval.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "item_id": "callsum-000417",
  "source": {
    "type": "support_call_transcript",
    "format": "vtt",
    "duration_sec": 742,
    "asr_engine": "vendor_asr_v3",
    "speaker_labels": ["agent", "customer"],
    "pii_method": "names, phones, account numbers replaced with typed placeholders"
  },
  "human_summary": {
    "origin": "agent_wrap_up_note",
    "written_within_min": 4,
    "text": "Cust [PERSON_1] disputed late fee on [ACCOUNT_1]. Waived one-time. Sent autopay enrollment link."
  },
  "screen": {
    "grounded": true,
    "external_knowledge_spans": [],
    "second_reviewer_agreement": "full"
  },
  "required_claims": [
    "Customer disputed a late fee",
    "Agent waived the fee as a one-time exception",
    "Agent sent an autopay enrollment link"
  ],
  "forbidden_claims": [
    "Fee was refunded to a card",
    "Customer agreed to enroll in autopay during the call"
  ],
  "strata": {"intent": "billing_dispute", "length_bucket": "10-15min", "resolution": "resolved"},
  "provenance": {"system": "contact_center_platform", "export_date": "2026-09-30", "license_ref": "LIC-REF"}
}

Note the second forbidden claim: the transcript shows the link was sent, not that the customer enrolled. That gap between action and outcome is one of the most common hallucinations in call summaries, and it belongs in your forbidden set deliberately.

Transcripts are the riskiest half of each pair, because they capture third parties who never signed anything with you. Recording law varies by state, and several states, California among them, require consent from all parties to record certain conversations, so ask how each recording was noticed and whether the consent covers secondary use for model evaluation. Confirm with counsel how those consents apply to your use.

De-identification must cover both sides of the pair. A wrap-up note can contain the same account number as the transcript, and a summary that keeps a name the transcript scrubbed defeats the exercise. Placeholder types should be consistent across source and summary so that claims like "[PERSON_1] disputed the fee" still score correctly.

Health-related sources are a separate tier. HIPAA de-identification relies on either Expert Determination or Safe Harbor, the latter requiring removal of 18 listed identifiers [5]. Clinical visit notes paired with transcripts are attractive summarization references, but they are not usable without one of those methods documented.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Building and protecting the set

A summarization eval set is only informative if it is stratified, held out and documented. Stratify by meeting or call type, length bucket (long transcripts stress context handling differently) and outcome, then cap any single team or agent so one writer's style does not dominate. For long inputs, the guide to long-context evaluation with real document sets covers length allocation.

Keep the references out of training. If the same company's minutes feed a fine-tuning run, split by source system or time window, and follow the practices in contamination-resistant evaluation design. Training pairs themselves are covered separately in summarization fine-tuning data from business records.

Document the set with a Data Card that records upstream sources, how summaries were written and screened, intended use and known gaps [6]. For voice products, pair this eval with outcome-based tests from voice agent evaluation sets. The wider map of options sits in the evaluation datasets hub, and rights questions are covered in the data provenance guide.

What to ask a supplier of summary pairs

A short request template keeps conversations concrete and makes non-matches obvious early.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Pair type and systems: for example, "Teams transcripts (VTT) joined to minutes in SharePoint by calendar event ID."
  • Who wrote the summaries, in what role, and how long after the source event.
  • Whether summaries reference documents outside the delivered source.
  • Volume by stratum (call intent, meeting type, length bucket), not just a total.
  • Recording notice and consent basis for each source channel.
  • De-identification method, applied to both source and summary, and how it was sample-checked.
  • Allowed uses: evaluation only, or evaluation plus training, and any limits on publishing scores.
  • Documentation available: source description, collection method, known quality issues.

SourceX sources operational datasets, including support and sales histories and documents, from US companies on request; categories are not inventory and a request does not guarantee a match. You can describe the summary pairs you need to SourceX using the template above. For broader context on eval sets from business work, see AI evaluation datasets built from real business work.

Request summarization evaluation pairs from real business work

SourceX looks for US businesses that hold the source-summary pairs you describe, rights-reviews each dataset for ownership and consents, and delivers under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Describe your summarization evaluation needs.

Sources

  1. arXiv, "Open4Business(O4B): An Open Access Dataset for Summarizing Business Documents" (2020). https://arxiv.org/pdf/2011.07636
  2. arXiv (Feigenblat et al., Findings of EMNLP 2021), "TWEETSUMM - A Dialog Summarization Dataset for Customer Service" (2021). https://arxiv.org/abs/2111.11894v1
  3. OneUptime, "How to Build Ground Truth for RAG Evaluation When No Reference Answers Exist" (2026). https://oneuptime.com/blog/post/2026-08-31-build-rag-ground-truth-without-reference-answers/markdown
  4. Sigma AI, "Golden datasets: Evaluating fine-tuned large language models". https://sigma.ai/?p=26358
  5. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  6. arXiv (Pushkarna, Zaldivar, Kjartansson, Google Research; FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  7. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Artificial intelligence - Data quality for analytics and machine learning - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data