Fine-tuning and post-training data
Mid-training and annealing data: high-quality domain data late in training
Quick answer
Mid-training data is the smaller, higher-quality mix a lab switches to after bulk pre-training and before post-training, usually while the learning rate decays. Recent surveys describe it as high-quality general text blended with math, code, QA, instruction-style and domain data that has been selected, decontaminated and carefully mixed [1][2]. For a buyer, the qualifying traits are editorial quality, real expertise, low boilerplate, clean provenance and zero overlap with evaluation sets, not raw token count.
By SourceX Editorial · Updated
What mid-training and annealing mean in a training run
Mid-training is a distinct stage that spends a minority of compute on curated data to strengthen target skills without erasing general ability. One 2025 survey frames the stage as a step up in data quality beyond noisy web-scale corpora, aimed at math, coding, reasoning and long-context handling [2]. Another treats mid-training data as a hybrid pipeline of collection, synthesis, selection, decontamination and mixing, rather than a single dataset [1].
"Annealing" is the narrower term for the final phase of the learning-rate schedule. As the rate decays toward its floor, the model consolidates whatever it sees most recently, so labs reserve that window for their best data. The mid-training surveys describe this phase as an annealing-style step that sharpens data quality beyond noisy web-scale corpora [2].
The two terms overlap in practice. Some teams call the whole post-pretraining, pre-SFT block "mid-training"; others use "annealing" for the same thing. What matters for sourcing is that both phases are small relative to pre-training, so each document carries more weight. See the pretraining glossary entry and the post-training glossary entry for the stages on either side.
How mid-training differs from continued pre-training and SFT
Mid-training sits between two better-known stages, and the data requirements differ from both. Continued pre-training on a single domain, covered in our guide to sourcing domain corpora for continued pre-training, shifts a model's distribution with large volumes of domain text. Supervised fine-tuning, covered in how to source SFT data, teaches response format with prompt-response pairs.
Mid-training borrows from both. The mix is still mostly raw documents trained with next-token loss, but it is far more selective than a pre-training crawl and often includes instruction-shaped and reasoning-shaped text [1].
| Dimension | Bulk pre-training | Mid-training / annealing | SFT |
|---|---|---|---|
| Typical unit | Web page, book, repo file | Edited document, worked solution, technical record | Prompt-response pair |
| Selection pressure | Heuristic and classifier filters | Strict quality and domain targeting | Per-example review |
| Share of compute | Most | Small minority | Very small |
| Main risk | Noise and duplication | Contamination, overfitting to style | Format collapse |
| Licensing scrutiny per token | Lower | Higher | Highest |
The last row is the commercial point. Because annealing data is upweighted and seen late, each token plausibly has more influence on model behavior than crawl data, and memorized training data is known to be extractable from deployed models [6].
Which data labs put in an annealing mix
Published mixes combine three ingredients: top-decile general text, targeted domain data, and structured reasoning content. The mid-training surveys list high-quality general text, math and code, QA, instruction data and synthetic reasoning text as the main components [1][2]. Exact proportions vary by lab and are rarely disclosed in full.
Domain text that tends to qualify includes:
- Edited professional documents: final versions of technical manuals, engineering change notices, regulatory filings and contracts that went through review.
- Expert problem-solution records: resolved support escalations with root-cause notes, incident postmortems, worked calculations in finance or engineering.
- Long, coherent documents: specifications and reports that train long-context dependency, which surveys list as a mid-training target [2].
- Curated reasoning text: human-written derivations and step-by-step explanations; see reasoning trace datasets for how human and model-generated traces compare.
Openly licensed corpora such as the Common Pile, an 8TB collection of public-domain and openly licensed text built for pretraining, supply general high-quality text [7]. Their gap is proprietary operational knowledge: how real companies document, diagnose and decide.
Quality signals that qualify a domain corpus for annealing
A document qualifies for annealing when it shows evidence of editing, genuine expertise, high information density and low template residue. These are the signals model-based quality classifiers learn to approximate, and benchmarks such as DataComp-LM exist precisely to measure how such filtering choices change downstream performance [3]. Our guide to quality filtering for pretraining-scale text covers the heuristics in detail.
For licensed operational text, check these signals directly:
- Edit lineage. Does the record have a draft and a final, an approver field or a revision count? Final, reviewed versions beat drafts.
- Author expertise. Is the author role recorded (engineer, actuary, attorney), and is the writing from that role rather than a templated system?
- Boilerplate ratio. Strip signatures, disclaimers, ticket headers and auto-generated status lines; measure what share of tokens remains.
- Factual density. Count specific entities, quantities, part numbers or citations per thousand tokens.
- Near-duplicate rate. Run MinHash or similar deduplication at the document and paragraph level; templated records often collapse heavily.
- Language and encoding hygiene. Confirm UTF-8, no OCR garbage, consistent tables, no truncated documents.
Small, very clean corpora can matter more here than in bulk pre-training. The LIMA result, though about alignment rather than annealing, showed that 1,000 curated examples could carry outsized weight, and that curation was the expensive part [4]. Plan budget for selection, not just volume.
Contamination and memorization controls for late-stage data
Annealing data must be decontaminated against every evaluation set you report, because late-stage exposure inflates scores most directly. OpenAI stopped reporting SWE-bench Verified after concluding that score gains increasingly reflected training-time exposure to the benchmark [5]. Domain corpora built from public repositories, forums or standard exams are the usual leak path.
Practical controls:
- Run n-gram overlap (for example, 13-gram) and embedding-similarity checks between the candidate corpus and your eval suites before mixing.
- Hold out a random slice of the licensed corpus as an in-domain eval that never enters training.
- Record the decontamination method and thresholds in the dataset card so results stay auditable.
Memorization is the second risk. Research has shown that gigabytes of training data can be extracted from production models by querying alone [6]. Upweighted, repeated, late-seen documents are a plausible place for verbatim recall to concentrate, so personal details and confidential identifiers must be removed before the data reaches the training cluster, not filtered at inference.
Testing a candidate source with a small annealing run
The cheapest way to decide whether a domain source belongs in the mix is to anneal a mid-size checkpoint on it and compare against a control. Open model reports describe this kind of small annealing ablation for screening candidate sources, and the same pattern works for a buyer evaluating a licensed sample. Treat it as an internal experiment design, not a published standard.
Illustrative example: invented to show structure; it does not describe an available dataset.
annealing_probe:
base_checkpoint: internal-7b-step-450k # mid-schedule checkpoint, LR not yet decayed
schedule: linear decay to 0 over 20B tokens
arms:
control: { general_hq: 1.00 }
candidate: { general_hq: 0.85, licensed_sample: 0.15 }
licensed_sample:
source_type: resolved engineering incident reports
tokens: 3B # repeated to fill share; note epoch count
fields_kept: [title, symptom, root_cause, corrective_action, final_text]
fields_dropped: [reporter_name, email, asset_serial, customer_account]
pii_method: named-entity replacement, sample-checked
decontam: 13-gram overlap vs internal eval suite, 0 hits after filter
evals:
in_domain: held-out 5% of licensed sample (perplexity, task QA)
general: standard reasoning and knowledge suites
regression_guard: general score drop <= 0.5 pt
decision: include if in-domain gain is material and regression_guard passes
Keep the probe honest: match token counts across arms, track how many epochs the candidate is repeated, and report general-capability regressions alongside in-domain gains. For broader pre-purchase checks, see how to evaluate a fine-tuning dataset before you buy it.
Licensing and provenance questions for mid-training data
Because annealing data has outsized influence and memorization risk, its provenance needs more scrutiny per token than crawl data. Ask every supplier the same questions:
- Ownership and consent. Who created the documents, and did the holder have the right to license them for model training?
- Allowed uses. Does the license name pre-training, mid-training and derived model weights, or only "fine-tuning"? Ambiguous scope is the most common gap.
- Personal and regulated data. How were identifiers removed, and was a sample checked? Health records need HIPAA de-identification by Safe Harbor or Expert Determination [8].
- Third-party content. Do the documents embed vendor manuals, standards text or customer material the holder cannot sublicense?
- Refresh. Can the same source supply a later slice for the next run, under the same terms?
Licensing documentation is what lets data quality survive procurement and legal review. The foundation-model pre-training team guide covers the wider sourcing picture.
Where licensed operational data fits in a mid-training mix
Licensed operational records fill the gap open corpora leave: how experts actually write about real problems inside companies. SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents, and finance and legal workflows; categories are not held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, delivered under a license that defines records, uses, term and delivery, and has personal details removed or replaced with the method recorded and a sample checked. Teams can describe the data they need on the SourceX buyers page.
SourceX does not source scraped web content and does not train models. Every release is approved by the supplying company, and nothing is contracted until that supplier agrees. For the wider category map, start at the fine-tuning and post-training data hub or the AI data hub.
Source licensed domain text for your mid-training mix
If your annealing probes show that operational domain text moves the metrics you care about, describe the records, fields and allowed uses you need. SourceX looks for US businesses that hold that data, reviews rights and data handling, and manages licensing from first purchase through ongoing ones. Describe your mid-training data request.
Sources
- arXiv, "A Survey on LLM Mid-Training" (2025). https://arxiv.org/pdf/2510.23081
- arXiv, "Mid-Training of Large Language Models: A Survey" (2025). https://arxiv.org/html/2510.06826v1
- arXiv, "DataComp-LM: In search of the next generation of training sets for language models" (2024). https://arxiv.org/pdf/2406.11794
- arXiv (Meta AI and collaborators), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- arXiv, "Scalable Extraction of Training Data from (Production) Language Models" (2023). https://arxiv.org/abs/2311.17035v1
- NeurIPS 2025 Datasets and Benchmarks Track, "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://papers.neurips.cc/paper_files/paper/2025/hash/52acc050138d6f40dad6f12f91a4ce22-Abstract-Datasets_and_Benchmarks_Track.html
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.