Skip to content

Fine-tuning and post-training data

Multi-turn conversation data for chat fine-tuning

Quick answer

The most useful multi-turn conversation dataset for fine-tuning is a set of complete, real human-to-human service conversations with clear roles, outcome labels and documented rights. Map the human agent to the assistant role only where that agent's behavior is what you want the model to learn. Keep whole conversations rather than single turns, filter by resolution, tag canned macros, and replace names and account numbers with surrogates that stay consistent across turns.

By SourceX Editorial · Updated

Why real multi-turn data is scarce

Real multi-turn data is scarce because most open instruction data is single-turn, and most public multi-turn sets are synthetic or carry non-commercial licenses. The ConsistentChat authors note that existing instruction data largely neglects cross-turn coherence, and they report 20-30% chat-consistency gains from about 15,000 synthetic multi-turn conversations [1]. The TurnWise paper (March 2026) makes the same observation about open post-training data and reports a 12% gain on its multi-turn evaluation after adding roughly 10,000 multi-turn conversations [2].

Public sets that do exist tend to be small. A typical Hugging Face example is a small synthetic set in chat-messages format [3], which works for format testing but not as a production training mix. Licensing is the other constraint. The Data Provenance Initiative audit found license omission above 70% and license error rates above 50% on popular hosting sites [10], so the license line on a dataset card is not enough evidence on its own. For the rights side of dialogue corpora specifically, see human-to-human dialogue corpora with commercial training rights.

The richest supply sits inside companies: support desks (Zendesk, Salesforce Service Cloud, Intercom, Freshdesk), live-chat tools, sales chat and internal help desks. That data was not collected for training, so it needs construction work before it becomes SFT data.

What a usable source conversation contains

A usable source conversation contains the full ordered thread, speaker roles, timestamps, channel and the outcome, not just the message text. Ask suppliers for these fields before anything else, because they cannot be reconstructed later.

  • Thread identity: a stable conversation ID plus parent ticket ID, so chat sessions that continue by email or in a second session can be merged or deliberately split.
  • Ordered messages with author type: customer, human agent, bot, supervisor, system event. Bot and system messages must be separable from human agent messages.
  • Timestamps per message: needed to detect transfers, long holds and split sessions.
  • Agent metadata: a pseudonymous agent ID, team or tier, and tenure band if available. This lets you filter to experienced agents or a specific queue.
  • Macro and template markers: whether a message was inserted from a canned response library, and which one.
  • Outcome fields: resolution status, reopen flag, escalation or transfer, CSAT or thumbs rating, and the final disposition code.
  • Back-office context: the actions taken (refund issued, address changed, password reset) and the policy or knowledge article referenced. Datasets such as ABCD showed that service dialogues become more instructive when agent actions and guidelines are recorded alongside the text [6].

If you also need structured actions for an agent model, the policy-following service agent data guide covers how to pair conversations with policies and system actions. For licensing raw logs as an asset, the owner pages for chat logs and customer support transcripts cover the commercial side.

Mapping speakers to training roles

Map a human speaker to the assistant role only when that speaker's behavior is the policy you want the model to imitate. In most support data that means the human agent becomes assistant, the customer becomes user, and routing notes or policy context become system content. Chat training formats in torchtune and Vertex AI both expect role-tagged message lists of this kind [4][5].

Several cases break the naive mapping:

  • Bot-then-human handoffs: the first turns often come from an existing rules-based or LLM bot. Training on them teaches your model the old bot. Either drop those turns, keep them as context with loss masked, or truncate the conversation to start at the human handoff.
  • Multiple agents in one thread: a tier-1 agent who transfers to tier-2 produces two assistant voices with different permissions. Keep only the agent whose scope matches your deployment, or add a system note describing the transfer.
  • Internal notes: private agent notes and whisper messages are never customer-visible and must be removed from the dialogue, though they can feed a separate reasoning or summarization task.
  • Agent behavior you do not want: agents promise refunds they cannot authorize, paste policy verbatim or apologize six times. These are reasons to filter or mask, not to train.

When loss masking is supported, train on assistant turns only and mask user and system tokens; chat fine-tuning data format explains the message and masking conventions in more detail.

Keep whole conversations, segment them deliberately

Keep complete conversations as the training unit, because splitting them into single prompt-response pairs removes clarification, repair and memory of earlier turns. A model trained only on isolated turns does not learn to ask for an order number, correct a misunderstanding three messages later or recall a constraint stated at the start, which is exactly the cross-turn coherence gap the research above describes [1][2].

Segmentation is still needed, for three reasons. Long sessions can exceed your context window. Some threads contain two unrelated issues. Others span channels or days. Practical rules are to split on a new issue or a long gap (for example a new ticket ID or a gap of many hours), drop trailing pleasantries after resolution, and never cut in the middle of a clarification exchange. Record the segmentation rule so evaluation data is built the same way.

Filtering by outcome and building preference pairs

Filter supervised examples by outcome and keep the failures in a separate set, because failed conversations are valuable preference and evaluation data. LIMA showed that a small number of carefully curated examples can produce a strong chat model, while noting that curation is labor-intensive [7]. In service data, curation usually means keeping conversations that were resolved on first contact, not reopened, not escalated and rated well by the customer.

Failures should not be discarded. A reopened ticket with a poor rating, paired with a resolved conversation on the same intent, gives you material for preference training. InstructGPT trained on human rankings of outputs after SFT [9], and DPO fits a policy directly to preferred and rejected responses without a separate reward model [8]. Real service data rarely contains two responses to the identical context, so most teams either have reviewers write the better response at the failure point or generate a candidate and have reviewers rank it.

Treat CSAT carefully. It reflects the outcome (refund granted or not) as much as the agent's conduct, and survey response rates are often low and skewed toward customers with strong opinions. Combine it with reopen and escalation signals rather than using it alone.

Macros, canned text and de-identification

Canned macros should be removed or tagged, because a model trained on them learns to paste boilerplate instead of answering. In many support desks a large share of agent messages start from templates. Tag macro-inserted text with its template ID, keep lightly edited macros only where the edit carries the answer, and cap how often any single template appears in the training mix.

De-identification must use consistent surrogates across turns. If the customer says "I'm Dana" in turn 2 and the agent says "Thanks, Dana" in turn 5, both must become the same surrogate name, or the conversation stops making sense. The same holds for order numbers, account numbers, addresses, emails and phone numbers. Use typed placeholders or realistic fakes that preserve format (a 10-digit account surrogate stays 10 digits) so the model still learns to ask for and confirm identifiers. Free-text identifiers inside customer messages are the usual failure mode, so sample and review.

Where the supplier is a HIPAA covered entity or business associate and conversations contain protected health information, the de-identification standard in 45 CFR 164.514 applies, including the Safe Harbor identifier categories that must be removed [11]. Payment card numbers pasted into chat are another common leak, and should be detected and replaced before data leaves the supplier.

Conversion checklist and example record

Use the checklist below as the acceptance spec you send to a supplier or internal data team, then validate a sample against it before full delivery. More general acceptance testing is covered in how to evaluate a fine-tuning dataset before you buy it.

Illustrative example: invented to show structure; it does not describe an available dataset.

StepRuleCheck on a sample
Thread assemblyMerge messages by conversation ID, order by timestampNo orphan or out-of-order messages
Role mappingHuman agent to assistant, customer to user, bot turns masked or droppedZero bot or system text in assistant turns
Internal notesRemove private notes and whispersZero note markers in dialogue text
SegmentationSplit on new issue or long gap; keep clarification exchanges intactNo conversation ends on an unanswered question
MacrosTag template ID; cap per-template frequencyTemplate share per intent within agreed limit
Outcome filterSFT set: resolved, not reopened, not escalatedDisposition fields present on every record
Failure setReopened, escalated or low-rated conversations held separatelyKept out of SFT split
De-identificationTyped, consistent surrogates for names, accounts, addresses, cardsManual review of a random sample finds no residual identifiers
SplitsHold out by customer and agent, not by messageNo customer or agent appears in both train and eval

Illustrative example: invented to show structure; it does not describe an available dataset.

{"conversation_id": "c_000184", "channel": "web_chat", "intent": "billing.duplicate_charge",
 "outcome": {"resolved": true, "reopened": false, "escalated": false, "csat": 5},
 "messages": [
  {"role": "system", "content": "You are a billing support agent. Refunds over $100 require supervisor approval."},
  {"role": "user", "content": "I was charged twice for my subscription this month."},
  {"role": "assistant", "content": "Sorry about that. Can you confirm the last four digits of the card on file, [CARD_LAST4_1]?", "macro_id": null},
  {"role": "user", "content": "Yes, that's the one. Account [ACCOUNT_ID_1]."},
  {"role": "assistant", "content": "I see two charges on the 3rd. I've refunded the duplicate; it should appear in 3-5 business days.", "macro_id": "refund_confirm_v2", "action": "refund.issue"}
 ]}

Mixing conversation data into a fine-tuning run

Conversation data should be one component of a training mix, sized against general instruction data so the model does not narrow into a support persona. Balance intents so that password resets do not dominate, cap the share from any single agent, and keep a held-out set split by customer and agent to avoid leakage. The data mixtures for fine-tuning guide covers ratios and forgetting checks, and how to source supervised fine-tuning data covers the wider SFT supply picture.

Synthetic generation is useful for filling rare intents or rewriting failures into better responses, but it should be anchored on real conversations so the distribution of customer phrasing, typos and topic shifts stays realistic. Label synthetic and real records separately in the manifest. The fine-tuning and post-training data guide places conversation data alongside reasoning, function-calling and preference data.

Rights and diligence questions for conversation data

Conversation data needs rights diligence on two parties: the company that holds the logs and the customers and agents who appear in them. Ask the supplier which privacy notice and terms covered the customers at collection time, whether recordings or chats were disclosed as monitored, whether any customer opted out or requested deletion, and whether agent messages are covered by employment or contractor agreements. Ask how contractual restrictions with the supplier's own enterprise customers (common in B2B support) were checked.

Get the permitted uses in writing: SFT, preference training, evaluation, and whether derived synthetic data may be kept. The AI training data licensing guide covers license terms, and workplace email and chat datasets covers internal communications, which raise different consent questions from customer chats. If you would rather describe the conversations you need than search for holders yourself, SourceX works with AI data buyers to find US companies that hold them, and every release is approved by the supplying company.

Sourcing real multi-turn conversation data with SourceX

SourceX sources operational datasets, including support and sales histories, from US companies on request, and manages the licensing and ongoing purchases; a request does not guarantee a match. Every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with personal details removed or replaced before delivery. Describe the conversations, roles and outcome fields you need at SourceX for AI data buyers.

Frequently asked questions

Can I use ShareGPT-style logs of chatbot conversations instead?

Shared chatbot logs capture user-to-model behavior, so training on them mostly teaches your model to imitate another model, and their licenses are often non-commercial. They do not show how a skilled human resolves a real case with real systems behind them.

How many conversations do I need?

There is no fixed number. Published results show gains from roughly 10,000 to 15,000 multi-turn conversations [1][2], and LIMA reached strong chat quality with 1,000 curated examples [7]. For a domain assistant, quality, intent coverage and outcome filtering usually matter more than raw volume.

Should the system prompt include company policy?

Yes, where agent behavior depended on policy. Including the relevant policy excerpt in the system turn lets the model learn to apply the rule, rather than memorizing outcomes that only make sense under a policy it never saw.

Sources

  1. arXiv, "ConsistentChat: Building Skeleton-Guided Consistent Multi-Turn Dialogues for Large Language Models from Scratch" (2025). https://arxiv.org/pdf/2506.03558
  2. Graf et al. (arXiv), "TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities" (2026). (2026). https://arxiv.org/abs/2603.16759
  3. Hugging Face (DataCreatorAI), "Multi-Turn-Conversational-SFT dataset card". https://huggingface.co/datasets/DataCreatorAI/Multi-Turn-Conversational-SFT/blob/main/README.md
  4. PyTorch, "Chat Datasets (torchtune 0.4 documentation)". https://pytorch.org/torchtune/0.4/_sources/basics/chat_datasets.rst.txt
  5. Google Cloud, "Prepare supervised fine-tuning data for Gemini models". https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini-supervised-tuning-prepare
  6. arXiv (Chen et al., NAACL 2021), "Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems" (2021). https://arxiv.org/abs/2104.00783v1
  7. arXiv (Zhou et al.), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  8. arXiv (Rafailov et al.), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  9. arXiv (Ouyang et al., OpenAI), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  10. arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  11. eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data