Fine-tuning and post-training data
Reasoning trace datasets for fine-tuning: expert-written, model-generated or natural rationales
Quick answer
A reasoning dataset for fine-tuning pairs each problem with a step-by-step trace and a final answer that can be checked. Buyers get traces three ways: domain experts write them; a stronger model generates them and only traces with verified answers are kept; or the rationales professionals already record beside real decisions are licensed and paired with their inputs. Model-generated traces scale cheaply but can carry restrictive provider terms and unchecked steps. Expert traces cost expert hours. Natural rationales are short and need their inputs restored.
By SourceX Editorial · Updated
The chain-of-thought glossary entry defines the concept; this page covers sourcing traces for supervised fine-tuning within the fine-tuning and post-training data guide. See also building distillation datasets and RLVR datasets.
Expert-written, model-generated and natural rationales compared
The three sources differ most in what you can verify: expert traces step by step, model-generated traces usually only at the final answer, and natural rationales against what happened after the decision. Post-training datasets are commonly grouped as human-labeled, distilled or synthetic, and developers use one approach or a hybrid [1]; distilling teacher models' chain-of-thought traces is among the dominant construction strategies, while data built entirely from scratch has become rare [2].
| Expert-written traces | Model-generated, answer-filtered traces | Natural business rationales | |
|---|---|---|---|
| Who writes the steps | Credentialed writers following a style guide | A teacher model, sampled one or more times per prompt | The decision-maker, at the time, for colleagues or auditors |
| Typical shape | A clean derivation of the path that worked | Long, with exploration, self-checks and restarts; length varies by teacher and settings | A few sentences that point to documents rather than restate them |
| What can be checked | Final answer against a key; a second expert reviews the steps | Final answer against a reference, unit tests or a rule checker; steps rarely checked | The decision against later outcomes, such as a claim paid or a loan that performed |
| Main risk | Coverage gaps; undisclosed chatbot drafting | Wrong steps behind right answers; teacher terms; prompt pools that overlap benchmarks | Post-hoc justification; missing inputs; personal and client data in free text |
| Main cost driver | Expert hours per trace plus review | Prompt sourcing, sampling compute and verification | Finding rationale fields, rebuilding inputs, de-identification and rights review |
| Best fit | Domains with no reliable teacher; reference traces for evaluation | Math, code and other auto-checkable tasks; distilling into smaller students | Judgment in underwriting, claims, credit and code review, where no checker exists |
Label hybrids per record: an expert-edited model trace, a model-expanded rationale, or a model trace kept only when it matches the recorded human decision. For fully expert-authored data, see expert worked solutions and expert demonstration data.
What open reasoning datasets tell a buyer
Most public reasoning sets are distilled from a teacher model, and their documentation shows that curation and verification decide value more than row counts.
- Distillation can transfer strong results. The AM team reports that AM-Distill-Qwen-32B, trained only with supervised fine-tuning on its 1.4 million distilled responses, outperformed DeepSeek-R1-Distill-Qwen-32B on AIME2024, MATH-500, GPQA-Diamond and LiveCodeBench [3].
- A filtered slice can match the whole. MMFineReason distilled chain-of-thought rationales for 1.8 million multimodal samples (5.1 billion tokens) from Qwen3-VL-235B-A22B-Thinking and filtered them for quality and difficulty. Its authors report that a 7% difficulty-filtered subset of 123,000 samples performed comparably to the full set [4].
- Size says nothing about correctness. The ThinkChain-20M card lists 22.2 million rows and 35.8 billion tokens and states that its traces and answers were not individually verified for accuracy [5].
- Useful sets carry selection metadata. The EasyDistill authors released OmniThought, about 2 million reasoning processes generated and validated by DeepSeek-R1 and QwQ-32B, with each entry scored for reasoning verbosity and cognitive difficulty [6].
A 2026 primer on post-training reasoning data reports that data quality and construction often matter more than the training algorithm, and that a large corpus can still cover a target poorly when source mixture, leakage or lineage is uncontrolled [7]. Pay for verification and difficulty selection, then run an ablation on your own evaluation set (evaluating a fine-tuning dataset before you buy it).
Verifying answers, steps and lineage before training
Check every trace's final answer automatically, review a sample of its steps, and trace every prompt to its upstream source. SFT teaches the student to imitate the whole trace, so wrong steps are learned too.
A common filter keeps only traces whose final answer matches a reference. One public distillation set had a model answer prompts from GSM8K, PRM12K and part of NuminaMath, then removed incorrect answers, leaving about 92,000 samples [8]. That is necessary but not sufficient: a correct answer can follow two errors that cancel out, a lucky guess, or a step that contradicts an earlier one. Step-level checking is the job of process supervision data and step-level labels for process reward models.
| Check | How to run it | What it catches |
|---|---|---|
| Final answer | Normalize and compare with a reference; run unit tests for code; apply the business rule or schema | Wrong answers; traces cut off before an answer |
| Trace-answer agreement | Parse the trace's own conclusion and compare it with the answer field | Traces that argue for one result and output another |
| Step review | Experts score a stated sample of steps per task family against a rubric | Right answers reached through wrong reasoning |
| Lineage and overlap | Map every prompt to its upstream dataset, split and license; run n-gram and embedding similarity against your held-out sets | Benchmark items reused as training prompts; hidden duplicates; inflated scores |
| Degeneration | Flag repeated loops, language switching and traces that hit the generation limit | Text the student will learn to reproduce |
Lineage matters because reasoning prompt pools are heavily reused. A 2026 study of RLVR (reinforcement learning from verifiable rewards) datasets found that most are variants of a small set of upstream sources and that many carry contamination risks; its tracing framework attributed over 99.7% of 1.45 million instances to 20 atomic sources [9]. Decontaminate against your evaluation sets before training (decontaminating a licensed training set against public benchmarks).
Treat any written rationale as evidence of a plausible path, not proof of how an answer was produced. Research on rationales generated by a model trained on human think-aloud explanations cautions that they do not necessarily reveal the true decision-making process [10].
Provider terms that travel with model-generated traces
A trace generated by a commercial or open-weight model may carry that model's usage terms, which differ by provider, model and version, so record the terms in force on the generation date.
As of October 2026, Anthropic's help center, for example, states that its terms do not allow using outputs to train models that compete with Anthropic's, and names non-competing uses such as sentiment analysis, summarization and information extraction tools [11]. One public dataset card states that because its data was generated with BLOOMZ models under the BigScience RAIL License v1.0, that license would apply to classifiers fine-tuned on the dataset [12]. Lemley and Henderson argue that model output generally lacks the human authorship copyright requires, so such restrictions rest on contract rather than copyright, and their enforceability is unsettled [13]. A July 2026 Mayer Brown article on AI acquisitions advises asking which model generated synthetic data and what that model was trained on, and warns that generator restrictions could make a synthetic dataset unusable for its intended purpose [14].
For each model-generated record, require the generator name and checkpoint, the license or terms version and date, the account type, sampling settings, generation date and any human edits. Wider diligence is in due diligence for purchased synthetic fine-tuning data and provenance records for synthetic training data.
For disclosure, California's AB 2013 requires developers of generative AI systems made available to Californians to post training-data documentation, including whether synthetic data generation was used, first due by 1 January 2026 [15]. As of October 2026, providers of general-purpose AI models under the EU AI Act must publish a sufficiently detailed training-content summary using the AI Office template under Article 53(1)(d) [16].
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Natural rationales: the reasoning businesses already write down
Many operational systems store a short, human-written reason next to each decision, and these traces record real judgment under real constraints. To work as training data, each needs the inputs the decision used, as of the decision date, and the outcome that followed.
An actuary writing for the Society of Actuaries notes that underwriting evidence and decision rationale are typically locked in legacy systems, workbenches and rules engines, often unstructured, with only a limited set of codes passed to administration systems [17]. The same pattern holds elsewhere:
| Workflow | Where the rationale is written | Inputs to pair with it | Later signal to check it against |
|---|---|---|---|
| Commercial underwriting | Referral, decline or condition note in the workbench | Submission, loss runs, inspection report, guideline version | Bound or declined; later loss experience |
| Claims | Coverage position or reservation-of-rights letter; adjuster diary notes | Policy form and endorsements, first notice of loss, estimates | Payment, reopening, litigation |
| Commercial lending | Recommendation and exceptions sections of the credit memo | Financial statements, covenant calculations, collateral | Approval terms; later performance |
| Software engineering | Pull request description, review comments, decision records | Diff, failing tests, linked issue | Merge, revert, follow-up defect |
| Support escalations | Exception reason or supervisor approval note | Ticket thread, account state, policy text | Resolution, reopen, chargeback |
In a 2025 study of policy-exception decisions, LLMs applied policies rigidly even when doing so was impractical, and supervised fine-tuning on human explanations worked better than ethical-framework prompting or chain-of-thought prompting [18]. Natural rationales may justify rather than explain, so the caution in [10] applies to them too.
Use them in one of three ways: train on the rationale as written, placed before the decision; have a domain expert expand it into explicit steps, keeping the original as an anchor; or keep a model-generated trace only when it reaches the recorded decision for compatible reasons. Field-level detail is in decision records with rationale, underwriting decision rationale data and commit histories with change rationale.
SourceX sources operational datasets from US companies, including support histories, engineering records, and finance and legal workflows, where rationale fields like these are kept. Datasets are sourced on request, not held in stock, so a request does not guarantee a match, and the supplying company approves every release. Each dataset goes through rights review, and names, emails, phone numbers and account numbers are removed or replaced before delivery; no de-identification method is perfect, so scan free-text rationales yourself. You can describe the decision workflows and rationale fields you need.
A reasoning-trace record with provenance and checks
Deliver the trace in its own field, separate from the prompt and the final answer, with per-record origin and verification metadata. You can then train with or without it and filter without re-parsing text. The record below shows an expert-expanded natural rationale.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"id": "rt-uw-000231",
"task_family": "commercial_property_referral",
"messages": [
{"role": "system", "content": "You are a commercial property underwriting assistant. Apply the guideline, then decide: accept, refer or decline."},
{"role": "user", "content": "Submission [SUB_1]: 3-story mixed-use frame building, built 1958. Loss runs: 2 water-damage claims in the last 3 years. Inspection: knob-and-tube wiring partly replaced; electrical recommendation open. Guideline [G_4]: refer frame buildings over 50 years old with any open electrical recommendation."}
],
"reasoning": [
"Frame construction built in 1958 is over 50 years old, so the age condition in [G_4] is met.",
"The inspection lists an open electrical recommendation, so [G_4] requires referral, not acceptance.",
"Two water-damage claims in 3 years point to plumbing condition; ask for the plumbing update date.",
"No decline condition applies, so the outcome is referral with conditions."
],
"final_answer": {"decision": "refer", "conditions": ["licensed electrician's certificate before binding", "plumbing update date"]},
"meta": {
"trace_origin": "expert_expanded_natural_rationale",
"source_rationale": "Refer. Old frame, open elec rec, 2 water losses. Need electrician cert + plumbing date.",
"rationale_written": "at_decision",
"inputs_as_of": "2025-03-14",
"guideline_version": "[G_4] rev [R]",
"editor_role": "senior commercial property underwriter",
"generator": null,
"answer_check": {"method": "matches_recorded_decision", "result": "pass"},
"outcome": {"field": "bind_status", "value": "bound_with_conditions"},
"step_review": {"reviewer_role": "underwriting manager", "steps_flagged": []},
"trace_tokens": 118,
"tokenizer": "[BUYER_TOKENIZER]",
"prompt_lineage": "licensed_submission_file",
"eval_overlap": "none_found",
"deidentification": "names_addresses_policy_numbers_replaced",
"rights_ref": "LIC-[ID]"
}
}
For a model-generated record, generator carries the provenance fields listed under provider terms above, and answer_check names the reference answer or test suite.
Trace length, delimiters and loss masking
Fix rendering and length rules in the request, because both are hard to change once a supplier has produced the data. Ask for token counts measured with your tokenizer, not word counts.
- Delimiters. If traces arrive inline, require one delimiter convention (for example, tags such as
<think>and</think>) and a parseable final-answer marker, so answer checks run automatically (chat fine-tuning data format). - Loss masking. Prompt tokens are normally masked from the loss; decide whether trace tokens are trained (reasoning SFT) or masked (answer-only tuning).
- Length distribution. Ask for the median, 95th percentile and maximum trace length per task family. Traces longer than your training sequence length get truncated and teach an unfinished argument (long-context fine-tuning data).
- Verbosity and difficulty tags. Per-record scores like OmniThought's [6] let you choose the mix rather than accept the supplier's.
- Brevity where you want it. If every example carries a long trace, expect the student to reason at length on easy prompts too; include short-trace or answer-only examples where latency matters.
Red flags and a request checklist for reasoning data
Missing verification records, missing generator metadata and broken license chains are the usual signs of unusable reasoning data. Watch for:
- "Verified" with no method named and no per-record verification field.
- Model-generated rows with no generator, checkpoint or terms recorded.
- A dataset-card license with no chain back to the upstream prompt sources. The Data Provenance Initiative's audit of more than 1,800 text datasets reported license omission rates above 70% and error rates above 50% on popular dataset hosting sites [19].
- "Human-written" traces with uniform length and phrasing (detecting model-generated content in purchased human data).
A request for reasoning data should state:
- Task families, the target share of each, and whether answers are machine-checkable or judgment-based.
- Allowed origin and hybrid edits per family.
- Verification method, pass criteria and step-review sample size per family.
- Generator metadata and terms for every model-generated record.
- Trace layout (fields, delimiters, step segmentation, final-answer marker) and a length cap measured with your tokenizer.
- A prompt lineage manifest and decontamination against named evaluation sets.
- For natural rationales: as-of inputs, outcome fields and the de-identification method.
- License scope, including whether the trained model may be distilled or used to generate derived traces (fine-tuning-only data licenses).
- Volume sized from a learning curve (how much data you need to fine-tune an LLM).
Looking for reasoning traces grounded in real decisions?
Describe the decisions or problems the traces should cover, the rationale fields, inputs and outcome data you need, and the license scope you need. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Submit your reasoning data requirements to SourceX.
Sources
- arXiv:2503.06072, "A Survey on Post-training of Large Language Models" (2025). https://arxiv.org/pdf/2503.06072
- arXiv:2604.10480, "Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs" (2026). https://arxiv.org/pdf/2604.10480
- a-m-team, via hyper.ai paper mirror, "AM-DeepSeek-R1-Distilled 1.4M (paper page for arXiv:2503.19633)" (2025). https://hyper.ai/papers/2503.19633
- hyper.ai paper mirror, "MMFineReason (paper page for arXiv:2601.21821)" (2026). https://hyper.ai/en/papers/2601.21821
- SVECTOR Corporation, Hugging Face, "ThinkChain-20M dataset card (README.md)". https://huggingface.co/datasets/SVECTOR-CORPORATION/ThinkChain-20M/blob/main/README.md
- arXiv:2505.20888, "EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models" (2025). https://arxiv.org/pdf/2505.20888
- arXiv:2606.02113, "A Primer in Post-Training Reasoning Data: What We Know About How It Works" (2026). https://arxiv.org/pdf/2606.02113
- Hugging Face (Zigeng), "dParallel_LLaDA_Distill_Data dataset card (README.md)". https://huggingface.co/datasets/Zigeng/dParallel_LLaDA_Distill_Data/blob/main/README.md
- arXiv:2605.26971, "RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data" (2026). https://arxiv.org/pdf/2605.26971
- arXiv:1901.03729, "Automated Rationale Generation: A Technique for Explainable AI and its Effects on Human Perceptions" (2019). https://arxiv.org/pdf/1901.03729
- Anthropic Help Center, "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
- BatsResearch, Hugging Face, "NusaX-senti-LexC-Gen dataset card (commit 260c323)". https://huggingface.co/datasets/BatsResearch/NusaX-senti-LexC-Gen/commit/260c3230882c4337c063ca2458bb67ba74667fa4
- SpicyIP, "Discussing Lemley and Henderson's 'The Mirage of Artificial Intelligence Terms of Use Restrictions'" (2025). https://spicyip.com/2025/01/discussing-lemley-and-hendersons-the-mirage-of-artificial-intelligence-terms-of-use-restrictions.html
- Mayer Brown, "Synthetic Data as a Deal Asset: Ownership, Provenance and Diligence Considerations in AI Acquisitions" (2026). https://www.mayerbrown.com/en/insights/publications/2026/07/synthetic-data-as-a-deal-asset-ownership-provenance-and-diligence-considerations-in-ai-acquisitions
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- Society of Actuaries, Reinsurance News, "An Actuary's Perspective: Unlocking the Power of Structured Underwriting Data" (2024). https://www.soa.org/sections/reinsurance/reinsurance-newsletter/2024/december/rsn-2024-12-ma/
- arXiv:2503.02976, "Teaching AI to Handle Exceptions: Supervised Fine-tuning with Human-aligned Judgment" (2025). https://arxiv.org/html/2503.02976v2
- Longpre et al. (arXiv:2310.16787; journal version in Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.