Skip to content

Fine-tuning and post-training data

Summarization fine-tuning data from business records

Quick answer

The best summarization fine-tuning data from business records is a naturally occurring pair: a source artifact (call transcript, meeting recording, case thread) linked to the summary a professional wrote for a real reader, such as after-call notes, minutes or a closing summary. These pairs carry domain vocabulary and audience conventions that commissioned summaries often miss. Buy them only with join keys intact, faithfulness filtering, audience tags and consistent de-identification across both sides of each pair.

By SourceX Editorial · Updated

Which business records already contain source-summary pairs

Most operational systems already store a summary next to its source, because a person had to write one to close the work. The job is to find systems where the link between the two is explicit in the data, not reconstructed by timestamp guesswork.

Source artifactNaturally paired summaryTypical system and join fieldCommon quality issue
Support call recording or transcriptAfter-call work (ACW) notes, disposition codeContact-center platform plus CRM case; join on call or interaction IDNotes written in shorthand under handle-time pressure
Sales call transcriptOpportunity notes, next steps, stage changeConversation-intelligence tool plus CRM opportunity; join on activity IDRep notes omit objections the transcript shows
Internal meeting recordingMinutes, action-item list, decision logVideo platform plus wiki or project tracker; join on meeting IDMinutes written days later, partly from memory
Support ticket threadResolution summary or knowledge-base articleHelp desk; join on ticket IDClosing summary copies the last message
Incident timeline and chatPostmortem summaryIncident tool plus engineering wikiSummary includes facts learned after the incident
Claim or case fileAdjuster or case-worker closing noteCase-management system; join on case numberNote references documents not in the export

Public sets show the gap. TWEETSUMM provides customer-service dialog summaries, but from chat conversations rather than phone calls [2]. Open4Business offers about 17,458 business articles with reference summaries, and its authors note that licensing is a barrier for this kind of data [1]. If you need spoken calls or meetings in a specific vertical, owner pages on meeting transcripts, contact-center call recordings and sales call transcripts with outcomes describe the underlying record types.

Why natural summaries need filtering before SFT

Natural summaries are uneven, so the dataset you train on should be a filtered subset, not the full export. A rushed agent note like "cust called re bill, fixed" teaches the model to drop content, while a minutes document written after the fact may contain claims the recording does not support.

Curation tends to matter more than volume. LIMA fine-tuned a 65B model on 1,000 carefully selected examples and reported competitive results, which argues for spending budget on selection [3]. For summarization, the practical filters are:

  • Length ratio: drop pairs where the summary is under roughly 2% or over 50% of source tokens; tune bounds per record type.
  • Coverage: check that key entities, amounts, dates and commitments in the source appear in the summary.
  • Copy rate: flag summaries that are near-verbatim spans of the source, which teach extraction rather than abstraction.
  • Reviewer signal: prefer summaries that a supervisor approved, edited or that passed QA scoring, where that field exists.
  • Temporal leakage: remove pairs whose summary was edited after later events, since it describes facts absent from the source.

How to check faithfulness in purchased pairs

Faithfulness means every claim in the summary is supported by the source, and it is the property most likely to fail in natural data. Overlap metrics such as ROUGE reward shared wording, not support, so they cannot reliably catch it. The risk sits in the training targets: if human references contain unsupported claims, your model learns to hallucinate in the house style.

A workable buyer check uses three passes on a sample. First, split each summary into atomic claims and run an entailment model (NLI) against source chunks, flagging claims below a threshold. Second, have a domain reviewer grade 100 to 200 flagged and unflagged pairs as supported, unsupported, or supported by context outside the source (for example, the customer's account history). Third, decide whether "outside-source" claims are acceptable for your feature; for a call-summary product that sees CRM context at inference time, they may be.

Ask the supplier for the unsupported-claim rate on a reviewed sample, the reviewer guidelines, and whether filtered-out pairs can be delivered separately. Rejected pairs with reviewer edits are useful later: an original note and its corrected version form a preference pair, and methods such as DPO train directly on preference data and have been reported to match PPO-based RLHF on summarization [4]. For broader acceptance tests, see how to evaluate a fine-tuning dataset before you buy it.

Tagging summaries by audience and purpose

Summaries written for different readers are different tasks, so tag them rather than mixing them silently. An agent's internal ACW note, a customer-facing recap email and an executive escalation brief can all summarize the same call, yet they differ in register, length, redaction and what counts as important.

Useful tags include audience (internal_agent, customer, manager, legal), purpose (handoff, record, decision, follow-up), format (free text, bulleted, templated fields) and author_role. With these tags you can condition the model at training time ("Write a customer recap") and balance the mix so one high-volume note type does not dominate. If templated fields dominate, the task is closer to extraction; see structured-output fine-tuning data. If the target is a long expert document rather than a short summary, report-generation fine-tuning data is the closer pattern.

De-identifying both sides of each pair consistently

Transcripts and their summaries carry the same personal data, so de-identification must be applied to both with the same mapping. If the transcript replaces "Dana Ortiz" with [PERSON_1] but the summary keeps the name or uses [NAME], you teach the model to invent or leak identifiers, and you leave re-identification risk in the target text.

Ask for one consistent pseudonymization key per record pair, entity types preserved as typed placeholders, and handling for spoken forms such as spelled-out account numbers and email addresses read aloud. Audio itself is identifying; if you license recordings rather than transcripts, voice is a separate decision. Health-related calls or case notes that contain PHI need HIPAA de-identification through Safe Harbor, which removes 18 listed identifiers, or Expert Determination [5]. Outside HIPAA, the ICO frames identifiability as a risk assessment and recommends the motivated intruder test as a starting point, a useful lens for small-business call notes where context alone can identify a caller [6]. Healthcare specifics are covered in healthcare LLM fine-tuning data.

A record schema for source-summary pairs

A purchasable pair should arrive as a record that preserves the join, the provenance and the filtering signals, not as a flat prompt-completion file. You can render it into chat format later; see chat fine-tuning data format.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "pair_id": "p-000184",
  "source": {
    "type": "call_transcript",
    "interaction_id": "int-55120",
    "language": "en-US",
    "speaker_turns": 42,
    "tokens": 3810,
    "asr_engine_version": "recorded by supplier",
    "text_uri": "delivery://pairs/p-000184/source.jsonl"
  },
  "summary": {
    "type": "after_call_notes",
    "audience": "internal_agent",
    "purpose": "handoff",
    "author_role": "tier2_agent",
    "written_offset_minutes": 6,
    "edited_after_close": false,
    "reviewer_approved": true,
    "text": "[PERSON_1] reported duplicate charge on invoice [INVOICE_1]. Refund issued; confirmation sent to [EMAIL_1]."
  },
  "quality": {
    "length_ratio": 0.031,
    "nli_unsupported_claims": 0,
    "copy_rate": 0.12
  },
  "deid": {"method": "typed placeholders, shared key per pair", "sample_checked": true}
}

Document the collection with a datasheet-style card covering upstream systems, annotation and review methods, and intended use [7]. Teams starting from a broader record pool can compare this with turning business records into instruction-response pairs.

What to put in a summarization data request

A good request describes the source-summary relationship precisely so a supplier can tell whether its systems hold it. Use this checklist when writing one.

  • Source type, language, modality (audio, transcript, documents) and typical length range.
  • Summary type, audience and whether reviewer approval or edits are recorded.
  • Required join key and how linkage is verified.
  • Minimum filtering already applied, and whether unfiltered pairs are wanted for your own pipeline.
  • De-identification method for both sides and the placeholder scheme.
  • Intended use: SFT, preference training, or a held-out evaluation split. Keep evaluation pairs separate from training; summarization evaluation sets with expert references covers that side.
  • License scope matching the use, for example a fine-tuning-only data license.

The wider buying process for post-training data sits in the fine-tuning and post-training data guide. When you are ready to describe a pair type to a sourcing partner, SourceX's buyer intake takes a description of the data rather than named companies.

Sourcing summarization pairs with SourceX

SourceX sources operational datasets, including support and sales histories and documents, from US companies on request; it does not hold them in stock, and a request does not guarantee a match. Each dataset is rights-reviewed, personal details are removed or replaced before delivery with the method recorded, and nothing is contracted until the supplying company agrees. Describe the source-summary pairs you need at SourceX for buyers.

Frequently asked questions

Are natural summaries better than commissioned ones?

They are better at matching real domain conventions and audience, and they are cheaper per pair when the system already holds them. Commissioned summaries are more consistent and can follow your own guidelines. Many teams combine both: natural pairs for the bulk, plus a smaller commissioned or reviewer-corrected set for the style they want.

Can I use summaries edited by reviewers as preference data?

Yes, if the original and edited versions are both retained with timestamps. The pair (original note, corrected note) for the same source is a natural chosen-rejected example for preference methods [4].

How many pairs do I need?

There is no fixed number; it depends on base model, task breadth and quality. LIMA's general-purpose results with 1,000 curated examples suggest quality and diversity matter more than raw volume [3], so plan a pilot filter on a sample before sizing a purchase.

Sources

  1. arXiv, "Open4Business (O4B): An Open Access Dataset for Summarizing Business Documents" (2020). https://arxiv.org/pdf/2011.07636
  2. arXiv (Findings of EMNLP 2021), "TWEETSUMM: A Dialog Summarization Dataset for Customer Service" (2021). https://arxiv.org/abs/2111.11894v1
  3. arXiv (NeurIPS 2023), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  4. arXiv (NeurIPS 2023), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  5. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  6. Information Commissioner's Office (ICO), "How do we ensure anonymisation is effective?" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/
  7. arXiv (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data