Fine-tuning and post-training data
Fine-tuning and post-training datasets: a buyer's guide
Quick answer
Fine-tuning datasets come in seven shapes, each matched to a training method: domain text for continued pre-training, prompt-response demonstrations for supervised fine-tuning (SFT), chosen-and-rejected pairs or binary ratings for preference optimization, reasoning traces, tasks with checkable answers for reinforcement learning, and teacher-model outputs for distillation. Pick the type from the behavior you need to change, then the sourcing route: open datasets after a license check, commissioned experts, licensed business records or synthetic generation.
By SourceX Editorial · Updated
For terms, see post-training, supervised fine-tuning and preference data. This hub, part of the AI data buyer's guide, covers which data to buy for each stage, where it comes from and what to check before it enters a training run.
Seven kinds of post-training data and the failure each one addresses
Each post-training method consumes a different record, so the first specification decision is the record shape, chosen from the failure you observe, not the topic. InstructGPT, for example, fine-tuned GPT-3 on labeler-written demonstrations, then applied reinforcement learning from human feedback (RLHF) using human rankings of model outputs, and evaluators preferred its outputs to those of the much larger GPT-3 [1]. Direct Preference Optimization (DPO) later fitted a model to preference pairs directly, without a separate reward model [2].
| Data type | Method it feeds | Typical record | Failure it addresses | Read next |
|---|---|---|---|---|
| Domain corpus | Continued pre-training, mid-training | Unlabeled documents with source, date and license metadata | Misused domain vocabulary, missing domain knowledge | Domain corpora for continued pre-training |
| Demonstrations | SFT, instruction tuning | messages list of system, user and assistant turns, or prompt and target | Right content, wrong format, fields or task behavior | Sourcing SFT data |
| Preference pairs | Reward models for RLHF, DPO | Prompt, chosen, rejected, annotator ID, optional margin and rationale | Acceptable answers that experts would rank lower | Preference datasets for DPO |
| Binary feedback | KTO-style methods, reward modeling | Prompt, one response, desirable or undesirable label, signal source | Responses reviewers already reject in production | Binary feedback for KTO |
| Reasoning traces | SFT on worked solutions, process supervision | Problem, ordered steps, final answer, optional step labels | Right answers reached by unreliable steps | Reasoning trace datasets |
| Verifiable tasks | Reinforcement learning with verifiable rewards (RLVR) | Prompt, reference answer or test, verifier specification | Wrong totals, codes or results a program can check | RLVR datasets |
| Distillation outputs | SFT on a teacher model's responses | Prompt, teacher response, teacher model and version, generation settings | A small model that should behave like a larger one | Synthetic data due diligence |
Volume needs differ sharply between rows. Continued pre-training corpora run to billions of words: one financial-domain study built a corpus in which SEC filings supplied 3.3 billion words, 16.5% of the total, with financial news from Common Crawl as the rest [3]. SFT can work with far fewer examples: LIMA tuned a 65-billion-parameter LLaMA model on 1,000 curated prompt-response pairs, and its authors concluded that almost all knowledge is learned in pre-training while instruction data mainly teaches format [4]. That is one reason SFT is an unreliable way to fill a knowledge gap; fine-tuning vs RAG vs continued pre-training covers the choice, and the capability pages on domain-specific fine-tuning, reward models, process supervision and RL environments list record examples.
Four sourcing routes and where each one breaks
Post-training data comes from open datasets, commissioned expert or annotation work, licensed business records, or synthetic and distilled generation. Most programs blend all four, so decide which route supplies which record type and what defect to test for.
| Route | Best fit | Typical defect | First check |
|---|---|---|---|
| Open instruction and preference datasets | Baselines, general instruction following, mixture filler | Misstated or non-commercial licenses; embedded model outputs; low-quality examples | Trace every component to its original license text |
| Commissioned experts or annotators | Comparisons of your own model's outputs, rubrics, demonstrations for tasks nobody records | Cost per item; annotator disagreement; templated writing; missing IP assignment | Pilot batch with measured agreement and a signed rights assignment |
| Licensed business records | Demonstrations and outcome signals from real work, at volume | Conversion effort; noisy outcome fields; personal data; policy drift | Which field is the label, and who or what set it |
| Synthetic or distilled data | Rare formats, scaling a seed set, smaller models | Inherits the generator's errors; provider output terms | Generator model, version and terms on the generation date |
Open datasets. The Data Provenance Initiative audited more than 1,800 text datasets, including widely used fine-tuning collections, and its arXiv version reports license omission above 70% and license error rates above 50% on popular hosting sites [5]. Quality varies as much as licensing: the AlpaGasus authors found many low-quality responses in Alpaca's 52,000 instruction examples, and a model trained on about 9,000 examples kept by an LLM grader outperformed the original [6]. See which open instruction and preference datasets allow commercial fine-tuning.
Commissioned work. LIMA's authors note that curating examples of that quality is labor-intensive [4]. Commission when no existing record captures the judgment you need, and compare managed RLHF collection with licensed preference data first. For blending real and generated records, see combining licensed and synthetic data.
Business records that already contain post-training signal
Many companies store demonstrations, preferences and verified outcomes as a by-product of work: a resolved ticket pairs a request with an accepted answer, a QA scorecard grades a response, and a supervisor's edit turns a draft into a better final. The buying task is converting those records without mistaking an administrative status for a human judgment.
| Business record | Post-training data it can yield | Verify before buying |
|---|---|---|
| Resolved support and service tickets | SFT pairs from the customer message and final agent reply | Closed by a person, not an auto-close rule; policy version in force |
| QA scorecards and reviewer corrections | Rubric scores, binary labels, corrected targets | Rubric version, reviewer role, calibration across reviewers |
| Supervisor approvals and rejections | Binary feedback | Whether rejections carry a reason and a named decision-maker |
| Draft-to-final edits of letters, reports and contracts | Preference pairs (final chosen, draft rejected) or rewriting SFT | The final was approved, not just the last save |
| Reconciled ledger entries, adjudicated claims | Verifiable tasks with known correct outcomes | Outcome stable after audit or appeal |
| Expert reports with their source inputs | Long-form SFT for report generation | Inputs complete enough to justify the report |
Similar record families appear in human feedback datasets of QA scores and corrections. Conversion is covered in turning business records into instruction-response pairs, draft-to-final pairs for rewriting models and RL environments from business workflows.
SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing agreement. Datasets are sourced on request rather than held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Each dataset goes through rights review, and names, emails, phone numbers and account numbers are removed or replaced before delivery with the method recorded, though no de-identification method is perfect. To source records like these, describe the post-training signal you need.
Illustrative example: invented to show structure; it does not describe an available dataset.
source_record:
record_id: case-2025-118204
systems: [helpdesk_ticket, qa_scorecard, order_lookup]
policy_version: returns-2025-06 # answers are only correct under this policy
closed_by: agent # not an auto-close rule
reopened_within_14_days: false
privacy: {deid_method: typed_placeholders, replaced: [name, email, order_id]}
derived_examples:
sft:
messages:
- {role: system, content: "Support agent for an appliance retailer. Apply returns policy 2025-06."}
- {role: user, content: "My dishwasher arrived with a cracked door. Order [ORDER_ID]."}
- {role: assistant, content: "<final reply approved by QA>"}
preference:
prompt_ref: sft.messages[0:2]
chosen: "<final reply after supervisor edit>"
rejected: "<agent's first draft>"
signal_source: qa_edit
reviewer_role: support_qa_lead
binary:
response_ref: agent_first_draft
label: undesirable
rubric_item: offered_all_policy_options
license: {scope: fine_tuning, contains_third_party_model_output: false}
Rights questions that are specific to post-training data
Post-training data carries rights problems that raw pre-training text rarely does: outputs from other models, research-only licenses on popular instruction sets, and authored work by annotators. Settle each in writing before the data enters a training run.
Model-output terms. Distilled and synthetic records can carry the generating provider's output terms, whatever their copyright status. As of October 2026, Anthropic's help center, for example, says its terms do not allow outputs to be used to train competing models, gives general-purpose chatbots and open-ended text generation models as prohibited examples, and lists non-competing uses such as content categorization, summarization and information extraction tools as allowed [7]. Lemley and Henderson argue that such restrictions rest on contract rather than copyright, because model output generally lacks human authorship, and commentators question whether parties that never accepted the terms are bound [8]. Record the generating model, version and terms for every synthetic record.
Inherited non-commercial licenses. Sets built from shared chatbot logs combine both problems: a third-party catalog lists ShareGPT 52K, a set of user-shared ChatGPT conversations, under CC BY-NC 4.0 for research and non-commercial use [9]. Confirm terms on the original dataset card, not a mirror.
Annotator and expert IP. Demonstrations, rationales and rubrics are authored works. Contracts should assign or license them to you, bar contributors from pasting employer-confidential or model-generated text, and state whether the vendor may reuse the items for other clients; see contracting domain experts for post-training data.
License scope. Confirm the grant covers fine-tuning, release of the tuned weights and any hosted fine-tuning service that receives the data; see fine-tuning-only data licenses and processor and API terms for hosted fine-tuning.
Personal data. The Janus Interface authors report that fine-tuning can amplify privacy risk, showing that tuning on a few personal-data examples can make a model recover personal information from its pre-training data [10], and Nasr et al. extracted memorized training data from an aligned production chat model [11]. Alignment is not a privacy control, so scan prompts and context fields as well as targets; see scanning a corpus for PII before fine-tuning, differential privacy for LLM fine-tuning and the privacy and de-identification hub.
Disclosure. As of October 2026, California's AB 2013 requires developers of generative AI systems offered to Californians to post a high-level summary of training datasets, due by 1 January 2026 and again before each later release of a covered system or substantial modification [12]. In the EU, Article 53(1)(d) of the AI Act requires providers of general-purpose AI models to publish a summary of training content, according to a template provided by the AI Office [13]. Whether a fine-tune makes you a provider or developer is covered in fine-tuning with acquired data under the EU AI Act and AB 2013.
Scope a post-training purchase in six decisions
A scoped request names the behavior to change, the record shape, a pilot size, how labels were produced, the acceptance test and the licensed uses. Use these six decisions as the skeleton of a data request or RFP.
- Objective. The failing behavior and the held-out evaluation that shows it, fixed before any data arrives; see targeted data acquisition from model failures and the evaluation datasets hub.
- Record shape and format. The row from the first table, the file format (usually JSONL), the chat template and which turns carry loss; see chat fine-tuning data format.
- Size. A pilot first, then volume. LIMA [4] and AlpaGasus [6] both suggest selection quality can matter more than raw count; see how much data you need to fine-tune an LLM.
- Label provenance. For preference and binary data: reviewer role and qualification, rubric version, whether ties were allowed and per-item agreement. One ICML 2026 paper notes that standard methods such as DPO treat high-disagreement pairs the same as unanimous ones, and that inconsistent labels can severely degrade aligned models [14].
- Acceptance test. Schema validation, deduplication against your eval sets, a PII scan of every field, a gold-label audit on a random sample, and a measured gain on the held-out eval from a pilot run. Ask for machine-readable documentation too: NeurIPS 2026 requires Croissant-RAI-based metadata for its Evaluations and Datasets Track [15]. See evaluating a fine-tuning dataset before you buy it and running a data pilot with a supplier.
- Licensed uses. Fine-tuning, weight release, customer-specific models, term, and whether tuned models may be kept after the license ends.
For RFP and contract mechanics, see how data procurement differs by training stage.
Start here: fine-tuning guides by stage
Each guide below goes deeper on one record type, domain or buying step.
- SFT and instruction data: expert demonstrations, real-world prompt sets, structured outputs, multi-turn conversations, function calling, summarization, report generation, long context, retrieval-augmented fine-tuning, classification and routing, style and voice, non-English instruction data, hallucination reduction.
- Continued pre-training and mid-training: mid-training and annealing data, data mixtures that prevent catastrophic forgetting.
- Preference and RLHF: domain-expert preference data, on-policy vs off-policy preference data, rubrics as rewards.
- Reasoning and RL: expert worked solutions; RLVR and reasoning-trace guides are linked in the first table.
- Safety: safety-tuning and refusal data, keeping safety intact when fine-tuning, compliance-reviewed communications.
- Domain data: legal, financial services, healthcare under HIPAA; for other sectors, the industry data hub.
- Budget and upkeep: allocating a post-training data budget, refresh cadence.
Mistakes that waste a post-training data budget
These errors produce a dataset that trains cleanly and changes the wrong thing.
- Buying demonstrations to teach facts. SFT mainly shapes format and behavior [4]; missing knowledge calls for a domain corpus or retrieval.
- Treating status fields as labels. Auto-closed tickets, bulk approvals and migrated records look like outcomes but record no judgment.
- Mixing in model outputs without provenance. One unrecorded generator can put a whole set under terms nobody reviewed.
- Tuning on narrow data alone. Narrow sets risk eroding general skills and safety behavior, which is why replay and safety data belong in the mixture.
- Letting eval items into the training purchase. Near-duplicates inflate every later score; the data quality hub covers overlap checks.
Describe the fine-tuning data your model needs
Tell SourceX the behavior you want to change, the record type (demonstrations, preference or approval signals, outcome-labeled tasks or domain text), the volume and the uses you need licensed. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Specify your fine-tuning dataset.
Guides in this section
- Continued Pre-Training Datasets: Sourcing Domain CorporaHow to source, qualify and license in-domain text for continued pre-training: source systems, token audits, deduplication, privacy, replay and rights.
- DPO Dataset: Preference Pair Format, Sources and BuyingWhat a DPO dataset contains (prompt, chosen, rejected, label metadata), where preference pairs come from, and the checks to run before you buy one.
- Evaluate a Fine-Tuning Dataset Before Buying: Ablation PlanA pre-purchase test for SFT and preference data: evaluation rights, sample rules, ablation arms, private evals, decision thresholds and acceptance terms.
- Expert Demonstration Data for LLMs: Commission or LicenseHow to source expert demonstration data for LLM fine-tuning: commission credentialed writers or license work product, then verify experts and rights.
- Expert Preference Data for RLHF: Raters, Rubrics, ReviewHow to source expert preference data for RLHF and DPO: match rater credentials, rank errors in a domain rubric, adjudicate disagreement and vet samples.
- How Much Data to Fine-Tune an LLM: Examples and TokensHow many examples, preference pairs or tokens an LLM fine-tune needs, by objective: published reference points, a worked sizing example and a checklist.
- How to Create an Instruction Dataset from Company DataHow to turn resolved tickets, approved drafts, posted forms and expert answers into instruction pairs: context, filtering, de-identification and splits.
- Instruction Tuning Datasets: Commercial Use License ChecksWhich open SFT and preference datasets allow commercial fine-tuning, why model-written responses carry provider terms, and the license checks to run.
- Prompt Datasets for RLHF: Sourcing Real User RequestsHow to source a prompt dataset for RLHF from real customer and employee requests: distribution, tagging, deduplication, de-identification and rights.
- Reasoning Datasets for Fine-Tuning: Human vs Model TracesCompare expert-written, model-generated and business-rationale reasoning traces for SFT: answer checks, teacher terms, lineage and trace format.
- RLHF Data Collection Services vs Licensed Preference DataCompare RLHF data collection services, expert networks and licensed preference judgments, and what to put in the RLHF statement of work and pilot.
- RLVR Datasets: Verifiers, Difficulty and ContaminationWhat an RLVR dataset holds (prompt, reference answer, verifier), how verifiers fail, and how to check pass rates, lineage and contamination before buying.
- Sourcing SFT Data: Commission, License, Open or GenerateCompare commissioned, licensed, open and model-generated SFT data, then specify an order: task mix, prompt source, response rules, QA gates and rights.
- Structured Output Fine-Tuning Dataset: Sourcing JSON PairsBuild structured-output fine-tuning pairs: document, schema and verified JSON from systems of record, with null cases, field metrics and PII surrogates.
- Synthetic Fine-Tuning Data License: Vendor Due DiligenceVet a synthetic SFT or preference dataset before you license it: generator and judge terms, seed rights, filtering evidence, real-data pilots, warranties.
- Chat Fine-Tuning Data Format: Messages, Roles, Loss MasksHow to specify chat SFT records for suppliers: messages arrays, roles, system prompts, loss masking, metadata and JSONL checks that load without rework.
- Compliance-Reviewed Communications as Fine-Tuning DataHow to source compliance review records (drafts, decisions, required edits, cited rules) as SFT and preference data that teaches domain policies.
- Contracting Domain Experts for Post-Training DataHow to hire domain experts for AI training data: engagement models, IP assignment, confidentiality, credential checks, QA layers and payment terms.
- Expert Worked Solutions Data for Applied ReasoningHow to source step-by-step expert solutions for finance, engineering and tax reasoning: record schema, verifiable answers, masking and acceptance tests.
- Financial LLM Fine-Tuning Data: Sourcing GuideHow to source financial LLM fine-tuning data: reconciliations, analyses and client communications with numeric fidelity, as-of context and GLBA limits.
- Fine-Tuning Compromises Safety: Data Mixes That Reduce ItWhy fine-tuning an aligned model can erode its safety, and which safety, refusal and over-refusal data to mix in and evaluate to keep guardrails intact.
- Fine-Tuning Data to Reduce Hallucinations: Abstention DesignHow to build SFT and preference data that cuts hallucinations: known-fact filtering, unanswerable and insufficient-context cases, and abstention balance.
- Fine-tuning vs RAG vs Continued Pre-training: Data NeedsFine-tuning vs RAG vs continued pre-training compared by the data each consumes: record shapes, volumes, failure modes and the license rights to acquire.
- Function-Calling Fine-Tuning Data Format and CoverageThe record anatomy and coverage spec for function-calling SFT data: tool schemas, calls, tool results, parallel calls, no-call cases and error recovery.
- Healthcare LLM Fine-Tuning Data: Clinical vs Admin RecordsHow to source medical LLM fine-tuning data: clinical notes vs administrative records, HIPAA de-identification routes, BAAs, and evaluation across sites.
- KTO Dataset Format: Training on Binary Feedback DataHow to structure approvals, rejections and thumbs signals as a KTO dataset: record format, label thresholds, class balance and bias checks.
- Legal LLM Fine-Tuning Data: Redlines, Memos, PrivilegeHow to turn legal work product (redlines, negotiated clauses, memos) into SFT and preference pairs while screening privilege and client confidences.
- LLM Distillation Datasets: Teacher Sampling and FilteringHow to build an LLM distillation dataset: prompt pool coverage, teacher sampling, rejection filtering, teacher terms and fitting data to the student.
- Long-Context Fine-Tuning Data: Sourcing Long DocumentsHow to source naturally long business documents and tasks for long-context SFT: length specs, document families, task types, packing and privacy checks.
- Mid-Training and Annealing Data for LLMs: What QualifiesWhat labs put in mid-training and annealing mixes, which quality signals qualify licensed domain text, and how to test a candidate source before you buy.
- Multi-Turn Conversation Datasets for Chat Fine-TuningHow to source real multi-turn conversations for chat fine-tuning: role mapping, turn segmentation, outcome filtering, macro tagging and de-identification.
- Multilingual Instruction Tuning Data: Native vs TranslatedDecide between translated English SFT data and native-language instruction data, where native data comes from, and how to check it before you buy.
- On-Policy vs Off-Policy Preference Data for DPO and RLHFWhy fixed preference pairs drift from the policy you train, what licensed off-policy pairs still do well, and when to buy prompts, rubrics and judges.
- Report Generation Datasets: Inputs Paired With Expert TextHow to source report-generation fine-tuning data: input bundles paired with expert-written reports, faithfulness checks, boilerplate dedup and masking.
- Retrieval-Augmented Fine-Tuning Data: RAFT RecordsHow to build retrieval-augmented fine-tuning data: questions, oracle and distractor documents, cited answers, no-answer cases, splits and training rights.
- RL Environments from Business Workflows: Data to LicenseThe records you need to build RL environments from real business workflows: task histories, system states, policies and outcomes, and how to license them.
- RLAIF vs RLHF: When Human Preference Data Is Worth ItRLAIF vs RLHF for buyers: where LLM-judge preference labels hold up, where human or expert judgments earn their cost, and how to calibrate a judge first.
- RLHF Data Cost: What Drives Preference and SFT PricingWhat drives RLHF and SFT data cost: pricing units, expertise, task length, review layers and rework, plus a worksheet to compare quotes per accepted item.
- Rubric-Based Rewards: Expert Rubric Data for RL TrainingHow rubric-based rewards work in RL for non-verifiable tasks, what expert rubric data must contain, and how to source and calibrate it before training.
- Safety Fine-Tuning Datasets: Refusal and Borderline DataHow to source safety fine-tuning data: balance harmful, borderline and benign prompts, write target responses, and avoid over-refusal and eval leakage.
- SFT Data Mixtures to Prevent Catastrophic ForgettingHow to design a domain SFT data mixture that avoids catastrophic forgetting: replay data, safety data, task caps, ratio sweeps and regression evals.
- Summarization Fine-Tuning Data From Business RecordsHow to source summarization fine-tuning data from business records: natural source-summary pairs, faithfulness filters, audience tags, de-identification.
- Targeted Data Acquisition: Model Failures to Data RequestsTurn fine-tuned model failures into purchasable data requests: build an error taxonomy, spec each slice, protect held-out sets and measure the fix.
- Text Revision Datasets: Draft-to-Final Pairs for Edit ModelsHow to source text revision datasets: draft-to-final document pairs, tracked changes and review comments for training rewriting and editing models.
Sources
- Ouyang et al., OpenAI (arXiv:2203.02155), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
- Rafailov et al., Stanford (arXiv:2305.18290; NeurIPS 2023), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
- Xie et al. (arXiv:2311.08545), "Efficient Continual Pre-training for Building Domain Specific Large Language Models" (2023). https://arxiv.org/pdf/2311.08545
- Zhou et al., Meta AI and collaborators (arXiv:2305.11206; NeurIPS 2023), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- Longpre et al. (arXiv:2310.16787; journal version in Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Chen et al. (arXiv:2307.08701), "AlpaGasus: Training A Better Alpaca with Fewer Data" (2023). https://arxiv.org/pdf/2307.08701v1
- Anthropic Help Center (provider terms guidance), "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
- SpicyIP, "Discussing Lemley and Henderson's 'The Mirage of Artificial Intelligence Terms of Use Restrictions'" (2025). https://spicyip.com/2025/01/discussing-lemley-and-hendersons-the-mirage-of-artificial-intelligence-terms-of-use-restrictions.html
- LLM Configurator (third-party dataset catalog), "ShareGPT 52K: LLM Instruction / SFT Dataset". https://llmconfigurator.com/en/datasets/sharegpt
- Chen et al. (arXiv:2310.15469), "The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks" (2023). https://arxiv.org/pdf/2310.15469
- Nasr et al. (ICLR 2025), "Scalable Extraction of Training Data from Aligned, Production Language Models" (2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/cce0e917b050208170151f77b497fc71-Abstract-Conference.html
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- ICML 2026 (poster), "Reliability-Aware LLM Alignment from Inconsistent Human Feedback" (2026). https://icml.cc/virtual/2026/poster/66787
- NeurIPS blog, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.