Skip to content

Fine-tuning and post-training data

Fine-tuning and post-training datasets: a buyer's guide

Quick answer

Fine-tuning datasets come in seven shapes, each matched to a training method: domain text for continued pre-training, prompt-response demonstrations for supervised fine-tuning (SFT), chosen-and-rejected pairs or binary ratings for preference optimization, reasoning traces, tasks with checkable answers for reinforcement learning, and teacher-model outputs for distillation. Pick the type from the behavior you need to change, then the sourcing route: open datasets after a license check, commissioned experts, licensed business records or synthetic generation.

By SourceX Editorial · Updated

For terms, see post-training, supervised fine-tuning and preference data. This hub, part of the AI data buyer's guide, covers which data to buy for each stage, where it comes from and what to check before it enters a training run.

Seven kinds of post-training data and the failure each one addresses

Each post-training method consumes a different record, so the first specification decision is the record shape, chosen from the failure you observe, not the topic. InstructGPT, for example, fine-tuned GPT-3 on labeler-written demonstrations, then applied reinforcement learning from human feedback (RLHF) using human rankings of model outputs, and evaluators preferred its outputs to those of the much larger GPT-3 [1]. Direct Preference Optimization (DPO) later fitted a model to preference pairs directly, without a separate reward model [2].

Data typeMethod it feedsTypical recordFailure it addressesRead next
Domain corpusContinued pre-training, mid-trainingUnlabeled documents with source, date and license metadataMisused domain vocabulary, missing domain knowledgeDomain corpora for continued pre-training
DemonstrationsSFT, instruction tuningmessages list of system, user and assistant turns, or prompt and targetRight content, wrong format, fields or task behaviorSourcing SFT data
Preference pairsReward models for RLHF, DPOPrompt, chosen, rejected, annotator ID, optional margin and rationaleAcceptable answers that experts would rank lowerPreference datasets for DPO
Binary feedbackKTO-style methods, reward modelingPrompt, one response, desirable or undesirable label, signal sourceResponses reviewers already reject in productionBinary feedback for KTO
Reasoning tracesSFT on worked solutions, process supervisionProblem, ordered steps, final answer, optional step labelsRight answers reached by unreliable stepsReasoning trace datasets
Verifiable tasksReinforcement learning with verifiable rewards (RLVR)Prompt, reference answer or test, verifier specificationWrong totals, codes or results a program can checkRLVR datasets
Distillation outputsSFT on a teacher model's responsesPrompt, teacher response, teacher model and version, generation settingsA small model that should behave like a larger oneSynthetic data due diligence

Volume needs differ sharply between rows. Continued pre-training corpora run to billions of words: one financial-domain study built a corpus in which SEC filings supplied 3.3 billion words, 16.5% of the total, with financial news from Common Crawl as the rest [3]. SFT can work with far fewer examples: LIMA tuned a 65-billion-parameter LLaMA model on 1,000 curated prompt-response pairs, and its authors concluded that almost all knowledge is learned in pre-training while instruction data mainly teaches format [4]. That is one reason SFT is an unreliable way to fill a knowledge gap; fine-tuning vs RAG vs continued pre-training covers the choice, and the capability pages on domain-specific fine-tuning, reward models, process supervision and RL environments list record examples.

Four sourcing routes and where each one breaks

Post-training data comes from open datasets, commissioned expert or annotation work, licensed business records, or synthetic and distilled generation. Most programs blend all four, so decide which route supplies which record type and what defect to test for.

RouteBest fitTypical defectFirst check
Open instruction and preference datasetsBaselines, general instruction following, mixture fillerMisstated or non-commercial licenses; embedded model outputs; low-quality examplesTrace every component to its original license text
Commissioned experts or annotatorsComparisons of your own model's outputs, rubrics, demonstrations for tasks nobody recordsCost per item; annotator disagreement; templated writing; missing IP assignmentPilot batch with measured agreement and a signed rights assignment
Licensed business recordsDemonstrations and outcome signals from real work, at volumeConversion effort; noisy outcome fields; personal data; policy driftWhich field is the label, and who or what set it
Synthetic or distilled dataRare formats, scaling a seed set, smaller modelsInherits the generator's errors; provider output termsGenerator model, version and terms on the generation date

Open datasets. The Data Provenance Initiative audited more than 1,800 text datasets, including widely used fine-tuning collections, and its arXiv version reports license omission above 70% and license error rates above 50% on popular hosting sites [5]. Quality varies as much as licensing: the AlpaGasus authors found many low-quality responses in Alpaca's 52,000 instruction examples, and a model trained on about 9,000 examples kept by an LLM grader outperformed the original [6]. See which open instruction and preference datasets allow commercial fine-tuning.

Commissioned work. LIMA's authors note that curating examples of that quality is labor-intensive [4]. Commission when no existing record captures the judgment you need, and compare managed RLHF collection with licensed preference data first. For blending real and generated records, see combining licensed and synthetic data.

Business records that already contain post-training signal

Many companies store demonstrations, preferences and verified outcomes as a by-product of work: a resolved ticket pairs a request with an accepted answer, a QA scorecard grades a response, and a supervisor's edit turns a draft into a better final. The buying task is converting those records without mistaking an administrative status for a human judgment.

Business recordPost-training data it can yieldVerify before buying
Resolved support and service ticketsSFT pairs from the customer message and final agent replyClosed by a person, not an auto-close rule; policy version in force
QA scorecards and reviewer correctionsRubric scores, binary labels, corrected targetsRubric version, reviewer role, calibration across reviewers
Supervisor approvals and rejectionsBinary feedbackWhether rejections carry a reason and a named decision-maker
Draft-to-final edits of letters, reports and contractsPreference pairs (final chosen, draft rejected) or rewriting SFTThe final was approved, not just the last save
Reconciled ledger entries, adjudicated claimsVerifiable tasks with known correct outcomesOutcome stable after audit or appeal
Expert reports with their source inputsLong-form SFT for report generationInputs complete enough to justify the report

Similar record families appear in human feedback datasets of QA scores and corrections. Conversion is covered in turning business records into instruction-response pairs, draft-to-final pairs for rewriting models and RL environments from business workflows.

SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing agreement. Datasets are sourced on request rather than held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Each dataset goes through rights review, and names, emails, phone numbers and account numbers are removed or replaced before delivery with the method recorded, though no de-identification method is perfect. To source records like these, describe the post-training signal you need.

Illustrative example: invented to show structure; it does not describe an available dataset.

source_record:
  record_id: case-2025-118204
  systems: [helpdesk_ticket, qa_scorecard, order_lookup]
  policy_version: returns-2025-06          # answers are only correct under this policy
  closed_by: agent                          # not an auto-close rule
  reopened_within_14_days: false
  privacy: {deid_method: typed_placeholders, replaced: [name, email, order_id]}
derived_examples:
  sft:
    messages:
      - {role: system, content: "Support agent for an appliance retailer. Apply returns policy 2025-06."}
      - {role: user, content: "My dishwasher arrived with a cracked door. Order [ORDER_ID]."}
      - {role: assistant, content: "<final reply approved by QA>"}
  preference:
    prompt_ref: sft.messages[0:2]
    chosen: "<final reply after supervisor edit>"
    rejected: "<agent's first draft>"
    signal_source: qa_edit
    reviewer_role: support_qa_lead
  binary:
    response_ref: agent_first_draft
    label: undesirable
    rubric_item: offered_all_policy_options
license: {scope: fine_tuning, contains_third_party_model_output: false}

Rights questions that are specific to post-training data

Post-training data carries rights problems that raw pre-training text rarely does: outputs from other models, research-only licenses on popular instruction sets, and authored work by annotators. Settle each in writing before the data enters a training run.

Model-output terms. Distilled and synthetic records can carry the generating provider's output terms, whatever their copyright status. As of October 2026, Anthropic's help center, for example, says its terms do not allow outputs to be used to train competing models, gives general-purpose chatbots and open-ended text generation models as prohibited examples, and lists non-competing uses such as content categorization, summarization and information extraction tools as allowed [7]. Lemley and Henderson argue that such restrictions rest on contract rather than copyright, because model output generally lacks human authorship, and commentators question whether parties that never accepted the terms are bound [8]. Record the generating model, version and terms for every synthetic record.

Inherited non-commercial licenses. Sets built from shared chatbot logs combine both problems: a third-party catalog lists ShareGPT 52K, a set of user-shared ChatGPT conversations, under CC BY-NC 4.0 for research and non-commercial use [9]. Confirm terms on the original dataset card, not a mirror.

Annotator and expert IP. Demonstrations, rationales and rubrics are authored works. Contracts should assign or license them to you, bar contributors from pasting employer-confidential or model-generated text, and state whether the vendor may reuse the items for other clients; see contracting domain experts for post-training data.

License scope. Confirm the grant covers fine-tuning, release of the tuned weights and any hosted fine-tuning service that receives the data; see fine-tuning-only data licenses and processor and API terms for hosted fine-tuning.

Personal data. The Janus Interface authors report that fine-tuning can amplify privacy risk, showing that tuning on a few personal-data examples can make a model recover personal information from its pre-training data [10], and Nasr et al. extracted memorized training data from an aligned production chat model [11]. Alignment is not a privacy control, so scan prompts and context fields as well as targets; see scanning a corpus for PII before fine-tuning, differential privacy for LLM fine-tuning and the privacy and de-identification hub.

Disclosure. As of October 2026, California's AB 2013 requires developers of generative AI systems offered to Californians to post a high-level summary of training datasets, due by 1 January 2026 and again before each later release of a covered system or substantial modification [12]. In the EU, Article 53(1)(d) of the AI Act requires providers of general-purpose AI models to publish a summary of training content, according to a template provided by the AI Office [13]. Whether a fine-tune makes you a provider or developer is covered in fine-tuning with acquired data under the EU AI Act and AB 2013.

Scope a post-training purchase in six decisions

A scoped request names the behavior to change, the record shape, a pilot size, how labels were produced, the acceptance test and the licensed uses. Use these six decisions as the skeleton of a data request or RFP.

  1. Objective. The failing behavior and the held-out evaluation that shows it, fixed before any data arrives; see targeted data acquisition from model failures and the evaluation datasets hub.
  2. Record shape and format. The row from the first table, the file format (usually JSONL), the chat template and which turns carry loss; see chat fine-tuning data format.
  3. Size. A pilot first, then volume. LIMA [4] and AlpaGasus [6] both suggest selection quality can matter more than raw count; see how much data you need to fine-tune an LLM.
  4. Label provenance. For preference and binary data: reviewer role and qualification, rubric version, whether ties were allowed and per-item agreement. One ICML 2026 paper notes that standard methods such as DPO treat high-disagreement pairs the same as unanimous ones, and that inconsistent labels can severely degrade aligned models [14].
  5. Acceptance test. Schema validation, deduplication against your eval sets, a PII scan of every field, a gold-label audit on a random sample, and a measured gain on the held-out eval from a pilot run. Ask for machine-readable documentation too: NeurIPS 2026 requires Croissant-RAI-based metadata for its Evaluations and Datasets Track [15]. See evaluating a fine-tuning dataset before you buy it and running a data pilot with a supplier.
  6. Licensed uses. Fine-tuning, weight release, customer-specific models, term, and whether tuned models may be kept after the license ends.

For RFP and contract mechanics, see how data procurement differs by training stage.

Start here: fine-tuning guides by stage

Each guide below goes deeper on one record type, domain or buying step.

Mistakes that waste a post-training data budget

These errors produce a dataset that trains cleanly and changes the wrong thing.

  • Buying demonstrations to teach facts. SFT mainly shapes format and behavior [4]; missing knowledge calls for a domain corpus or retrieval.
  • Treating status fields as labels. Auto-closed tickets, bulk approvals and migrated records look like outcomes but record no judgment.
  • Mixing in model outputs without provenance. One unrecorded generator can put a whole set under terms nobody reviewed.
  • Tuning on narrow data alone. Narrow sets risk eroding general skills and safety behavior, which is why replay and safety data belong in the mixture.
  • Letting eval items into the training purchase. Near-duplicates inflate every later score; the data quality hub covers overlap checks.

Describe the fine-tuning data your model needs

Tell SourceX the behavior you want to change, the record type (demonstrations, preference or approval signals, outcome-labeled tasks or domain text), the volume and the uses you need licensed. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Specify your fine-tuning dataset.

Guides in this section

Sources

  1. Ouyang et al., OpenAI (arXiv:2203.02155), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  2. Rafailov et al., Stanford (arXiv:2305.18290; NeurIPS 2023), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  3. Xie et al. (arXiv:2311.08545), "Efficient Continual Pre-training for Building Domain Specific Large Language Models" (2023). https://arxiv.org/pdf/2311.08545
  4. Zhou et al., Meta AI and collaborators (arXiv:2305.11206; NeurIPS 2023), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  5. Longpre et al. (arXiv:2310.16787; journal version in Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  6. Chen et al. (arXiv:2307.08701), "AlpaGasus: Training A Better Alpaca with Fewer Data" (2023). https://arxiv.org/pdf/2307.08701v1
  7. Anthropic Help Center (provider terms guidance), "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
  8. SpicyIP, "Discussing Lemley and Henderson's 'The Mirage of Artificial Intelligence Terms of Use Restrictions'" (2025). https://spicyip.com/2025/01/discussing-lemley-and-hendersons-the-mirage-of-artificial-intelligence-terms-of-use-restrictions.html
  9. LLM Configurator (third-party dataset catalog), "ShareGPT 52K: LLM Instruction / SFT Dataset". https://llmconfigurator.com/en/datasets/sharegpt
  10. Chen et al. (arXiv:2310.15469), "The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks" (2023). https://arxiv.org/pdf/2310.15469
  11. Nasr et al. (ICLR 2025), "Scalable Extraction of Training Data from Aligned, Production Language Models" (2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/cce0e917b050208170151f77b497fc71-Abstract-Conference.html
  12. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  13. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  14. ICML 2026 (poster), "Reliability-Aware LLM Alignment from Inconsistent Human Feedback" (2026). https://icml.cc/virtual/2026/poster/66787
  15. NeurIPS blog, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data