Skip to content

Fine-tuning and post-training data

Chat fine-tuning data format: messages, roles and loss masking

Quick answer

Ask suppliers for one conversation per JSONL line, stored as a template-neutral messages array with explicit system, user, assistant and tool roles, plus a per-record metadata object. Do not accept data pre-rendered into one model's chat template. Apply the template at training time, mark which assistant turns should carry loss, and validate role order, empty turns, encoding and duplicates on arrival. That record shape loads into Hugging Face TRL, hosted tuning APIs and open-weight toolkits with only a thin adapter [1][2].

By SourceX Editorial · Updated

Why the messages array is the right delivery shape

The role-tagged messages array is the shape most trainers and tuning services consume, so it is the lowest-rework contract to put in a data specification. Hugging Face Transformers builds chat inputs from a list of messages with role and content keys, and TRL's SFT trainer accepts that "conversational" shape directly (check the current docs for your installed versions). Google's Gemini tuning format uses the same idea with different field names: a systemInstruction plus a contents array of role and parts entries [1]. Practitioner threads show the same system/user/assistant list as the answer to "what format do I upload" [5].

The older alternative is the Alpaca-style instruction record (instruction, input, output), which open-weight toolkits such as Unsloth still document and flatten into one prompt before training [2]. It works for single-turn tasks but cannot express multiple turns, tool calls or a system prompt separately from the user request. If your supplier's source is support tickets, chat logs or agent sessions, insist on messages; converting Alpaca rows into messages is trivial, while flattening messages into Alpaca loses turns, tool calls and the system prompt.

For the wider SFT buying process, see how to source supervised fine-tuning data and the fine-tuning buyer's guide. The term itself is defined in the glossary entry on supervised fine-tuning.

Roles, turn order and the system prompt

Each role should have one job, and the supplier should state the allowed order in the spec rather than leave it to inference. A workable rule set:

  • system: zero or one message, always first. It holds the operating instructions the deployed model will see (persona, policy, output format), not per-conversation facts.
  • user: the human or upstream application input. In business data this is the customer, the employee or the ticket body.
  • assistant: the target behavior. Every conversation must end on an assistant turn unless the record is explicitly an evaluation prompt.
  • tool: the result of a tool call, linked to the assistant message that requested it by a call identifier.

The system prompt deserves a deliberate decision. If you will deploy with a fixed system prompt, either have the supplier include that exact text in every record or leave the field empty and inject it in your pipeline; mixing both produces a model that behaves differently with and without it. If the source data had no system prompt, do not let the supplier invent one per record, because synthetic instructions become part of what the model learns.

Gemini places system text in a separate systemInstruction field and names the assistant side model [1]; most open-weight templates use assistant. That difference is exactly why you want role names normalized in the delivery and remapped by a short adapter.

Chat templates: deliver neutral records, render at training time

A chat template is the tokenizer-level recipe that turns a messages list into the exact string a model family expects, including special tokens, so it belongs to the model, not to the dataset. Llama, Qwen, Mistral and Gemma families each use different delimiters. A file pre-rendered with <|im_start|> markers is useless for a different base model and hard to audit.

Render with apply_chat_template at training time. The Transformers docs recommend add_generation_prompt=False for training, because the trailing "start of assistant reply" tokens are an inference aid, and warn that tokenizing a rendered string a second time can duplicate special tokens unless add_special_tokens=False is set. Templates also accept extra arguments such as tools and documents, which is how tool schemas and retrieval context reach the model.

Common failure modes when records arrive pre-templated:

  • Double beginning-of-sequence tokens after re-tokenization.
  • A system prompt silently dropped because the target template has no system slot.
  • Whitespace or newline differences between the supplier's template and the tokenizer's official one, which shifts the loss mask by a token.

Loss masking: which tokens the model learns from

Loss masking decides which tokens contribute to the gradient, and for chat SFT the default should be assistant turns only. TRL distinguishes two cases. For prompt-completion data, completion_only_loss defaults to scoring only the completion. For conversational data, assistant_only_loss=True scores assistant messages and ignores system and user text, and it requires a chat template that marks assistant spans with {% generation %} and {% endgeneration %}; recent TRL releases ship patched templates for a few model families, and for others you must check the template yourself.

The data contract matters more than the trainer flag. Ask the supplier to mark turns that should not be learned, such as an agent reply later corrected by a supervisor, a templated legal footer, or a greeting that appears thousands of times. A simple per-message boolean (for example "train": false) is template-neutral; your adapter can translate it to whichever per-message flag or mask your trainer or tuning service expects. Without it, the only safe option is to drop the whole conversation.

Multi-turn data raises one more choice: train on every assistant turn, or only the final one. Training on all turns uses more signal but teaches early, weaker replies too. Put the decision in the spec, and see multi-turn conversation data for chat fine-tuning for how to source conversations where every turn is worth learning.

Tool messages and structured outputs

Tool use needs three things in the record: the tool definitions available, the assistant's structured call, and the tool's returned result. Together AI's function-calling format, for example, puts a tools list on the line, lets assistant turns carry tool calls in place of plain text, and returns outputs as tool messages [4]. Require a stable call id linking each tool result to its call, JSON-valid arguments, and explicit no-call examples. The full treatment is on the function-calling fine-tuning data format page.

A record spec to send your supplier

The fastest way to avoid rework is to send the supplier a filled-in example and a field list before any data moves. The record below shows the shape; adapt field names to your pipeline.

Illustrative example: invented to show structure; it does not describe an available dataset.

{"id": "conv-000184",
 "messages": [
  {"role": "system", "content": "You are a billing support assistant. Answer in under 120 words."},
  {"role": "user", "content": "I was charged twice for invoice [INVOICE_ID]."},
  {"role": "assistant", "content": "Let me check that invoice.", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "get_invoice", "arguments": "{\"invoice_id\": \"[INVOICE_ID]\"}"}}], "train": true},
  {"role": "tool", "tool_call_id": "call_1", "content": "{\"status\": \"paid\", \"charges\": 2}"},
  {"role": "assistant", "content": "I can see two charges. I have opened a refund for the duplicate; it posts in 5-10 business days.", "train": true}
 ],
 "tools": [{"type": "function", "function": {"name": "get_invoice", "parameters": {"type": "object", "properties": {"invoice_id": {"type": "string"}}, "required": ["invoice_id"]}}}],
 "meta": {"source_system": "helpdesk", "created_date": "2025-11-03", "domain": "billing", "author_role": "tier-2 agent", "language": "en-US", "deid_status": "pseudonymized; method recorded", "rights_tag": "license-scope-A", "split": "train"}}

(Pretty-printed here; in the file each record is a single line.)

Metadata fields worth requiring on every record:

FieldWhy you need it
source_systemLets you rebalance mixtures and trace defects back to an origin
created_dateSupports time-based splits and recency filters
domain / author_roleCoverage analysis; separates expert from novice replies
deid_statusShows how personal details were removed or replaced
rights_tagTies each record to the license scope it was delivered under
splitPrevents train/eval leakage when the supplier creates holdouts

Document the same fields at dataset level in a dataset card, which Hugging Face recommends for every dataset to promote responsible use and inform users of potential biases [7]. See dataset cards for licensed enterprise data and the provenance guide for what the card should record.

Validation checks to run on delivery

Run automated checks before any record reaches a training job, because format defects fail silently as bad loss rather than as errors. The file-level baseline comes from the JSON Lines convention: UTF-8, no byte order mark, one valid JSON value per line and no blank lines [3].

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckRuleTypical defect it catches
ParseEvery line parses as a JSON objectEmbedded raw newlines, BOM, trailing commas
Role orderOptional system first, then user/assistant alternation; tool only after an assistant callTwo user turns merged badly, orphan tool results
Terminal turnLast message is assistantConversations cut before the reply
Empty contentNo empty or whitespace-only content unless a tool call is presentExport placeholders, stripped attachments
Token lengthRendered length under your max sequence lengthSilent truncation of the assistant reply
EncodingNo mojibake, consistent Unicode normalizationDouble-encoded UTF-8 from CSV exports
PlaceholdersRedaction tokens follow one documented patternMixed [NAME], <PERSON> and *** markers
DuplicatesExact and near-duplicate conversations removedMacro replies repeated thousands of times
Mask coverageEvery record has at least one trained assistant turnRecords that contribute no gradient

Truncation deserves special care: if you cut from the right, you remove the assistant reply that carries all the loss. Truncate from the left or drop the record. On volume, the LIMA results suggest that a small set of carefully curated examples can go a long way, so validation that removes noisy records is rarely a loss [6]. Google's Gemini tuning guide makes the same point that quality matters more than quantity [1]. For a pre-purchase sample review, use how to evaluate a fine-tuning dataset before you buy it.

Container, sharding and transcript sources

Keep the SFT record contract separate from the file container decision. JSONL is the common default for chat data; Parquet suits very large corpora, and either can carry the same record. File-type preferences are covered in what file formats AI buyers accept, and the upstream raw form (before conversion into training records) is covered in a delivery schema for chat and conversation transcripts. Ask for raw transcripts and converted SFT records together when you can, so you can rebuild records if your masking or template choices change.

If you are sourcing chat data from real business operations, SourceX sources operational datasets such as support and sales histories from US companies on request, with personal details removed or replaced before delivery and the method recorded. You can describe the records and format you need; a request does not guarantee a match.

Sourcing chat SFT data in the right format

SourceX sources operational datasets, including support and sales histories, from US companies on request and manages the licensing process. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Describe the chat fine-tuning data you need.

Frequently asked questions

Should the system prompt be in every training record?

Only if the deployed model will always see that same prompt. Otherwise leave it out of the records and inject it in your pipeline, so the model does not learn to depend on text it will not receive.

Is "train on completions only" the same as assistant-only loss?

They are related but apply to different dataset types. In TRL, completion-only loss applies to prompt-completion records, while assistant-only loss applies to conversational records and depends on template markers.

Can I convert ShareGPT-style or Alpaca data into messages?

Yes. Alpaca rows map to one user and one assistant message, ShareGPT from/value turns map to roles, and toolkits such as Unsloth document conversions to chat templates [2]. Check that converted records keep role labels and do not merge turns.

Sources

  1. Google Cloud, "Prepare supervised fine-tuning data for Gemini models". https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini-supervised-tuning-prepare
  2. Unsloth, "Datasets Guide". https://docs.unsloth.ai/basics/datasets-guide
  3. jsonlines.org, "JSON Lines". https://jsonlines.org/
  4. Together AI, "Fine-tuning for function calling". https://docs.together.ai/docs/fine-tuning-function-calling.md
  5. OpenAI Developer Community, "What is the correct format for dataset content for fine tuning the models (solved)". https://community.openai.com/t/what-is-the-correct-format-for-dataset-content-for-fine-tuning-the-models-solved/691954
  6. Zhou et al., "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  7. Hugging Face, "Create a dataset card (Datasets library documentation)". https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card
  8. Hugging Face (TRL documentation), "Chat Templates" (2026). https://huggingface.co/docs/trl/chat_templates

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data