Skip to content

Fine-tuning and post-training data

Building distillation datasets: teacher sampling, filtering and student fit

Quick answer

An LLM distillation dataset is a set of prompts paired with outputs from a stronger teacher model, filtered for correctness and quality, and used to fine-tune a smaller student. Its quality depends on three choices: which prompts you sample (coverage of the student's real task), how you sample the teacher (temperature, number of candidates, reasoning traces), and how you filter (verifiers, judges, length and duplicate rules). The teacher's license terms decide whether you may train on its outputs at all.

By SourceX Editorial · Updated

What a distillation dataset contains

A distillation dataset is a prompt pool plus teacher responses plus the metadata you need to filter and audit them. Classic knowledge distillation trained a small model to match a large model's temperature-softened probability distribution, the "soft targets". With modern LLMs, most teams use sequence-level distillation instead: the student is trained with ordinary supervised fine-tuning (SFT) on text the teacher generated, because teacher logits are often unavailable through an API.

Recent open reasoning models have been distilled this way, with small students fine-tuned on long teacher reasoning traces rather than trained with reinforcement learning from scratch [8]. Public releases show the pattern at scale: OmniThought packages roughly 2M reasoning traces from DeepSeek-R1 and QwQ-32B teachers, each annotated for verbosity and difficulty [1]. Distillation is not limited to chat, either; Rank-DistiLLM distills LLM ranking judgments into cross-encoder rerankers [3], which our page on reranker training data and distillation labels covers separately.

Two terms collide in search. "Dataset distillation" means condensing a training set into a few synthetic examples; this page is about knowledge distillation from a teacher model into a student. For the broader category context, see the fine-tuning and post-training data hub.

Why the prompt pool matters as much as the teacher

The prompt pool sets the ceiling on what the student can learn, because the teacher only demonstrates behavior on prompts you send it. A frontier teacher sampled on a narrow or synthetic prompt set produces a student that is fluent on that distribution and brittle elsewhere. Common failure modes are easy to spot once you look for them:

  • Template collapse. Prompts generated by an LLM from a handful of seed tasks repeat phrasing and structure, so the student overfits to the templates.
  • Missing long tail. Synthetic prompts underrepresent messy real inputs: pasted logs, half-finished emails, mixed-language tickets, scanned-form text with OCR errors.
  • Difficulty skew. Pools drawn from public benchmarks such as GSM8K and NuminaMath are heavy on clean math and light on ambiguous business requests [2].
  • Contamination. Prompts copied from evaluation sets inflate scores; deduplicate against every benchmark you report.

A working hypothesis many teams act on is that real user requests fill gaps synthetic prompts miss, especially for a student deployed on one workflow. Our guide to real-world prompt sets for post-training covers how to source and de-identify them. If your deployment task looks like support tickets, sales conversations or contract review, prompts drawn from those operational records are a closer match than anything a generator invents.

Teacher sampling settings that shape the data

Sample several candidates per prompt at moderate temperature, then select, rather than taking one greedy answer. Multiple samples give you material for rejection sampling (keep only candidates that pass a check), for difficulty estimates (pass rate across samples), and for preference pairs if you later move to DPO.

Record these parameters for every generation, because you will need them to reproduce or audit the set:

  • teacher model ID and exact version or snapshot date
  • system prompt and any few-shot exemplars
  • temperature, top_p, max_tokens and stop sequences
  • number of samples per prompt (k) and the seed, if the API exposes one
  • whether reasoning traces were requested, returned and kept
  • the timestamp and the terms version in force when you generated

Reasoning traces deserve a separate decision. Long chains of thought can transfer reasoning behavior to small students [1], but they also inflate sequence length, training cost and inference latency. OmniThought's verbosity and difficulty annotations are meant to help you select traces whose length and difficulty suit the student's capacity [1]. Our page on reasoning trace datasets, human-written versus model-generated goes deeper on this trade-off.

Filtering teacher outputs: verifiers, judges and rejection statistics

Filter every teacher output before it reaches training, and keep the rejection statistics as a first-class artifact. Use the strongest available check for each task type:

Task typePrimary filterSecondary filterTypical failure it catches
Math, numeric answersExact-match or symbolic check against reference answerFormat check on final-answer tagRight reasoning, wrong final value
CodeUnit tests in a sandboxLint and timeoutPlausible code that fails edge cases
Structured extraction (JSON)Schema validation plus field-level comparisonNull-rate checkHallucinated fields, wrong types
Open-ended writing, support repliesLLM judge with a written rubricHuman spot-check sampleConfident, generic or off-policy answers
Ranking and relevanceAgreement across teacher samplesComparison with human labels on a subsetPosition bias, unstable orderings

Verified filtering is common in public sets: one distillation set built from GSM8K, PRM12K and NuminaMath prompts dropped incorrect teacher answers, leaving about 92k samples [2]. Where no verifier exists, an LLM judge is the usual fallback; AlpaGasus scored teacher-generated instruction data with a judge model and trained on a smaller, higher-scored subset [5]. LIMA's 1,000 curated pairs make the same point from the human side: curation can matter more than volume [4]. See filtering instruction-tuning data for quality and RLVR datasets with prompts, reference answers and verifiers for verifier design.

Rejection statistics tell you where the teacher is weak. If 60% of samples fail on one prompt category, either the teacher cannot do that task, the reference answers are wrong, or the prompts are ambiguous. Silently dropping those prompts biases the student away from exactly the hard cases you care about.

Fitting the dataset to the student

Small students need more targeted data, not just less of it. A 1B to 8B student has limited capacity, so every sample spent on tasks it will never see in production is capacity not spent on the target workflow. Match the prompt mix to the deployment distribution, then add a smaller general slice to limit regressions on instruction following.

Practical fit checks before training:

  1. Length budget. Drop or truncate teacher responses longer than the student's context and inference latency budget allow.
  2. Format alignment. Convert teacher outputs to the student's chat template and tokenizer; check special tokens and tool-call syntax.
  3. Capability gap. Very long teacher traces on very hard problems may be unlearnable for a small student; difficulty annotations help you select a curriculum [1].
  4. Held-out eval. Build a test set from real task inputs, not from the distillation pool, and freeze it before you start generating.

For sizing, our page on how much data you need to fine-tune an LLM gives the ranges that published work supports.

Teacher terms and provenance: check before you generate

Read the teacher model's license or API terms before generating a single sample, because some explicitly reach distilled students. Gemma's terms, for example, define "Model Derivatives" to include models trained to behave like Gemma by transferring patterns from its outputs, which covers distillation and training on synthetic data, so the student falls under Gemma's terms, including its use restrictions [6]. Commercial API terms vary on whether outputs may be used to develop competing models; our page on training on other models' outputs and checking provider output terms walks through the clauses to look for.

The prompts carry rights too. An audit of more than 1,800 text datasets found license information omitted for over 70% and errors in over 50% on popular hosting sites [7], so "public" prompt sets need their own license check. Keep a provenance record per batch: prompt source and license, teacher and terms version, sampling settings, filter version and rejection counts. Our guide to provenance records for synthetic training data has a fuller schema, and the synthetic data glossary entry defines the terms.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Illustrative distillation record and request template

A distillation record should let you trace any training row back to its prompt source, teacher call and filter decision.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "dist-000412",
  "prompt": {
    "text": "Customer says invoice INV-[REDACTED] was charged twice after plan downgrade. Draft a reply and list refund steps.",
    "source": "licensed_support_tickets",
    "source_license_id": "LIC-0091",
    "deidentification": "names, emails, account numbers replaced with tokens",
    "task_category": "billing_dispute",
    "difficulty_est": 0.42
  },
  "teacher": {
    "model_id": "teacher-model-x",
    "terms_version": "2026-06-01",
    "temperature": 0.7,
    "top_p": 0.95,
    "k_samples": 4,
    "reasoning_trace_kept": false
  },
  "selected_response": "...",
  "filter": {
    "method": "llm_judge_rubric_v3",
    "score": 8.5,
    "threshold": 7.0,
    "samples_rejected": 2,
    "rejection_reasons": ["policy_violation", "missing_refund_step"]
  },
  "split": "train"
}

When you ask a data partner for licensed prompts, describe the data rather than the company. A useful request states: the deployment task; the record type (for example, inbound support tickets with resolution notes); fields needed; target volume and category mix; date range; languages; de-identification expectations; and intended uses, including that prompts will be sent to a named third-party teacher model. That last point matters, because sending prompts to an external API is a use the license has to allow. Our due diligence checklist for purchased synthetic fine-tuning data covers what to ask if you buy finished teacher outputs instead.

Where licensed operational prompts fit

Licensed operational records are most useful as the prompt side of a distillation set when the student serves a specific business workflow. SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents, and finance and legal workflows; it does not hold stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phones and account numbers are removed or replaced before delivery, and the method is recorded. You can describe the prompts you need on the SourceX buyer page.

Request licensed prompts for distillation

If your student needs prompts that look like real operational work, describe the records, fields and allowed uses you need. SourceX looks for US businesses that hold that data, and every release is approved by the supplying company and delivered under a license defining records, uses, term and delivery. Start a request at sourcex.si/buyers.

Sources

  1. arXiv, "EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models" (2025). https://arxiv.org/pdf/2505.20888
  2. Hugging Face (Zigeng), "dParallel_LLaDA_Distill_Data dataset card". https://huggingface.co/datasets/Zigeng/dParallel_LLaDA_Distill_Data/blob/main/README.md
  3. Webis Group, "Rank-DistiLLM (ECIR 2025)" (2025). https://www.webis.de/data/ecir25-rank-distillm.html
  4. arXiv (Zhou et al., Meta AI), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  5. arXiv, "AlpaGasus: Training A Better Alpaca with Fewer Data" (2023). https://arxiv.org/pdf/2307.08701v1
  6. Google AI for Developers, "Gemma Terms of Use (archived version)" (2024). https://ai.google.dev/gemma/terms-archive/terms_04_01_24
  7. arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  8. arXiv (DeepSeek-AI), "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning" (2025). https://arxiv.org/abs/2501.12948

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data