Fine-tuning and post-training data
How to source a supervised fine-tuning (SFT) dataset
Quick answer
A supervised fine-tuning dataset pairs prompts with the exact responses you want a model to learn to produce. Buyers get one in four ways: commission new demonstrations from writers or domain experts, license existing work product such as resolved tickets or approved replies, adopt an open dataset, or generate responses with another model. Choose the route by what the model must learn and which rights you need, then write an order that fixes the task mix, prompt source, response rules, quality gates and rights basis.
By SourceX Editorial · Updated
Definitions are in the glossary entries for supervised fine-tuning and instruction tuning; the fine-tuning and post-training data guide covers the other adaptation data types.
What SFT data can and cannot teach a model
SFT data mainly teaches a model how to respond: format, tone, task procedure, and when to ask a clarifying question or decline. It is a weak way to add factual knowledge.
InstructGPT fine-tuned GPT-3 on demonstrations written by contracted labelers before applying reinforcement learning from human feedback, and human evaluators preferred outputs from the resulting 1.3B-parameter model to those of the 175B-parameter GPT-3 [1]. LIMA fine-tuned a 65B-parameter LLaMA model on only 1,000 curated prompt-response pairs with no reinforcement learning; in the authors' human study its answers were equivalent to or preferred over GPT-4's in 43% of cases [2]. The LIMA authors conclude that almost all knowledge is learned in pre-training, that a small, high-quality set mostly teaches output format, and that curating such examples is labor-intensive [2].
So size the order with a learning curve rather than a headline count (how much data you need to fine-tune an LLM), and if the gap is knowledge rather than behavior, look at domain corpora for continued pre-training or retrieval instead.
Four routes to SFT data, compared
Commissioned demonstrations give the most control, licensed work product gives the most realistic prompts and expert answers, open datasets are the fastest start with the weakest rights record, and model-generated data is the cheapest to scale but inherits the generator's errors and terms.
| Commissioned demonstrations | Licensed existing work product | Open datasets | Model-generated responses | |
|---|---|---|---|---|
| What you receive | New prompts and responses written to your guidelines | Real requests paired with the answer an employee gave: resolved tickets, approved email replies, completed reports | Public instruction collections under a stated license | Responses, sometimes prompts too, written by an LLM and then filtered |
| Prompt realism | Medium: writers invent tidy, single-intent requests unless given real ones | High: real phrasing, missing details, attachments and prior thread | Varies; many prompts are themselves generated | Mirrors the seed prompts |
| Response quality control | Your guidelines, calibration and review | Fixed at source; filter by outcome (resolved, approved, not reopened) | Inherited, often undocumented | Grader model or human review |
| Main cost driver | Writer expertise, review layers | Selection, prompt rebuilding, de-identification | Your audit time | Filtering and review |
| Rights question | Did each writer assign rights, and were AI tools used | May the business license the records; client confidentiality; third-party content | Is the license tag correct and complete | Do the generator's terms allow training your model |
| Typical failure | Polished answers unlike production traffic | Signatures, macros and personal data left in | Non-commercial or mislabeled components | Confident errors and uniform style |
Match the route to the gap. Output structure needs small, perfectly consistent sets that generation or commissioning can supply (structured-output fine-tuning data); domain judgment, such as resolving an accounts-payable exception, is where licensed work product or expert writing earns its cost. Most orders combine routes (post-training data sourcing by team).
Commissioned demonstrations
Commissioning means a vendor or your own contractors write responses to a guideline you control, as in InstructGPT [1]. The risks are prompts that drift toward easy, fully specified requests, and writers who quietly draft with a chatbot, turning "human-written SFT data" into undisclosed distillation. State an AI-assistance policy and require a per-record disclosure. See expert demonstration data for SFT, writing guidelines for SFT demonstration writers and cost drivers of human SFT and preference data.
Licensed existing work product
Businesses already hold request-and-answer pairs produced in the course of work: a ticket and its resolution, an RFP question and the approved answer, a draft and the signed-off version. The prompt usually has to be rebuilt from the record, and responses filtered to accepted outcomes. LIMA itself drew most of its examples from existing community content such as Stack Exchange answers and wikiHow articles, alongside examples its authors wrote [2]; LongForm pairs human-written documents from corpora such as C4 and Wikipedia with instructions generated by an LLM [3]. The record-to-pair mapping is covered in turning business records into instruction-response pairs.
SourceX works this route: it sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing agreement. Datasets are sourced on request rather than held in stock, so a request does not guarantee a match. You can describe the work product you want as SFT targets or browse training data for domain-specific fine-tuning.
Open datasets
Open instruction collections are the quickest start, but their license metadata is unreliable. The Data Provenance Initiative traced more than 1,800 text datasets, starting with widely used instruction and alignment collections, and reported license omission above 70% and license error rates above 50% on popular hosting sites [4]. Treat a hub license tag as a lead, then check the dataset card, the original source terms and any model-generated components (open instruction datasets that allow commercial fine-tuning).
Model-generated responses
Generated data scales cheaply, but its quality varies and the generator's terms travel with it. AlpaGasus found that Alpaca's 52,000 instruction examples contained many low-quality instances with incorrect or irrelevant responses, and a 9,000-example subset selected by a ChatGPT grader outperformed the full set in GPT-4-judged and human evaluations [5].
On terms, as of October 2026 Anthropic's help center says its terms do not allow using outputs to train models that compete with its own and lists general-purpose chatbots and open-ended text-generation models among prohibited uses, while describing narrower tools such as summarization or classification systems as permitted [6]. Lemley and Henderson argue that such restrictions rest on contract rather than copyright, because model output generally lacks human authorship; whether they bind a party who never accepted them is debated [7]. Record the generating model, version, date and terms per record; see due diligence for purchased synthetic fine-tuning data.
Where the prompts come from matters as much as the answers
The prompt distribution decides which behaviors improve, so an order must name the prompt source and sampling rule, not only a count. Writer-invented prompts are tidy; real requests are multi-intent, under-specified, and arrive with a prior thread, attachment or account state the answer depends on.
Authentic prompts come from inbound requests held by businesses (support, IT helpdesk, procurement inquiries), your own product logs, and public Q&A. Before mining your own logs, check your privacy policy and contracts: FTC technology staff warned in January 2024 that model-as-a-service companies that use customer data for training contrary to their commitments may be liable under the laws the FTC enforces [8]. Hold out a time-split prompt slice for evaluation before responses are written; sampling and taxonomy are in real-world prompt sets for post-training.
What an SFT order specification must state
An SFT order should define the task mix, prompt source, response rules, per-record metadata, quality gates, rights basis and delivery format. Anything left open defaults to whatever is fastest to produce, usually short single-turn pairs.
| Field | What to specify | Why it matters |
|---|---|---|
| Task taxonomy and mix | Named task families with a target share each, plus a floor for rare but important tasks | Stops easy tasks crowding out hard ones |
| Prompt source | Licensed or logged real requests, writer-authored, or generated, with the sampling rule | Sets realism and evaluation-leakage risk |
| Context | System prompt, documents or prior turns sent with each prompt | Answers written without that context teach guessing |
| Multi-turn share | Share of conversations with more than one assistant turn, and maximum length | Single-turn data under-trains follow-ups (multi-turn conversation data) |
| Response rules | Length range per task, structure (prose, list, JSON), citation behavior, when to ask or decline | Inconsistent rules teach inconsistent behavior |
| Author qualifications | Role or credential per task family; conflict declarations | Expert tasks need expert writers and reviewers (expert annotations and labels) |
| AI-assistance policy | Tools allowed, if any, and a per-record disclosure | Keeps a "human-written" claim checkable |
| Record metadata | Prompt and response origin, author role, source date, domain, review result, de-identification method, rights reference | Filtering, audit and disclosures |
| Quality gates | Independent review share, pass threshold per task family, duplicate and contamination reports | Defines acceptance before delivery |
| Rights basis | Writer assignments, record-holder license, generator terms | Each route has its own paper trail |
| Delivery | UTF-8 JSON Lines, one conversation per line [9]; template-neutral roles; train, validation and held-out splits; a datasheet covering motivation, composition and collection process [10] | Loads into trainers without rework (chat fine-tuning data format) |
General request structure is in how to write a data request for suppliers. Below is one record as an order might require it, pretty-printed but delivered as one JSONL line.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"id": "sft-ap-000213",
"messages": [
{"role": "system", "content": "You are an accounts-payable assistant for a wholesale distributor. Apply the duplicate-payment policy."},
{"role": "user", "content": "[VENDOR_1] says invoice [INV_1] was paid twice in March. Ledger extract attached. What do we do, and what do I tell them?"},
{"role": "assistant", "content": "The extract shows two ACH payments for [INV_1], on 4 and 11 March, for the same amount, so the second is a duplicate. 1) Ask [VENDOR_1] to apply it as a credit to their next invoice or return it. 2) Place a hold on the vendor record until the credit posts. Draft reply: ..."}
],
"context_files": ["ledger_extract_[INV_1].csv"],
"meta": {
"task_family": "ap.duplicate_payment",
"prompt_origin": "licensed_record_reconstructed",
"response_origin": "expert_edit_of_approved_reply",
"author_role": "AP specialist",
"reviewer_role": "AP manager",
"review_result": "pass_minor_edits",
"ai_assistance": "none_declared",
"source_period": "2025-Q2",
"deid_method": "surrogate_tokens_plus_manual_sample",
"rights_ref": "LIC-0193",
"split": "train",
"near_dup_cluster": "c-88412"
}
}
Quality gates to write into the order
Accept SFT data on gates measured per tranche (review pass rates, consistency with the record, duplicate, contamination and personal-data screens), not on volume delivered.
- Independent review. A second qualified reviewer checks a stated share of every task family, against a pass threshold with replacement at the supplier's cost. Report rates per family; a good overall rate can hide a weak one.
- Consistency with the record. For licensed work product, check the response against the recorded outcome (the credit actually issued, the fix actually merged); for commissioned data, reviewers verify claims, not just tone.
- Near-duplicates. Lee and colleagues found many near-duplicates in language-modeling datasets, and deduplicated training made models emit memorized text about ten times less often [11]. In SFT data built from business records they often come from macros and templated replies; require a near-duplicate rate and the clustering method.
- Contamination against your evaluations. Yang and colleagues showed that paraphrased test items slip past n-gram decontamination, and a 13B model trained on rephrased benchmark samples reached scores on par with GPT-4 [12]. Ask for embedding or model-based screens against your held-out sets and any public benchmarks you report.
- Personal data. Researchers extracted hundreds of verbatim training sequences, including contact details, from GPT-2 [13], and SFT trains a model to reproduce its targets. The open-source Presidio detector states there is no guarantee it finds all sensitive information [14], so require the de-identification method plus a manual sample check (scanning a corpus for PII before fine-tuning).
- Grader filtering, validated. If the supplier filters with an LLM grader in the AlpaGasus style [5], ask for the rubric and threshold, then check the effect on your own evaluation (what LIMA and AlpaGasus show about filtering).
Before ordering full volume, run a sample ablation (evaluating a fine-tuning dataset before you buy it).
Rights evidence and disclosure records by route
Each route needs different rights evidence, and what you collect at purchase, per record rather than per vendor, is what training-data disclosures will later draw on.
- Commissioned: written IP assignment or license from every writer and reviewer, including the vendor's subcontractors, plus the AI-assistance disclosure.
- Licensed work product: the holder's authority to license the records, client confidentiality limits, third-party material in attachments, and consents; for protected health information under HIPAA, de-identification by Safe Harbor (18 listed identifiers removed and no actual knowledge that the rest identifies anyone) or Expert Determination [15].
- Open: the license of every component and the original source terms.
- Generated: the generator's terms in force on the generation date.
- Scope: whether the grant covers fine-tuning only, which base models, and adapters or merged weights (fine-tuning-only data licenses).
As of October 2026, providers of general-purpose AI models placed on the EU market must publish a sufficiently detailed summary of training content under Article 53(1)(d) of the AI Act, using the AI Office template [16]. California's AB 2013 requires developers of generative AI systems made available to Californians to post training-data documentation, including whether datasets contain personal information and whether synthetic data was used; it was due by 1 January 2026 and is due again before each new system or substantial modification is made available [17]. Per-record origin fields make both easier; when fine-tuning triggers them is in provider and developer duties when fine-tuning with acquired data.
For records sourced through SourceX, rights review checks that the business owns or may share them and that required consents are in place. Personal details such as names, emails and account numbers are removed or replaced before delivery, with the method recorded per dataset and a sample checked; no de-identification method is perfect, so run your own scan. Health records must be HIPAA de-identified before SourceX considers them for a license.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Red flags in an SFT data proposal
These signs in a proposal or sample usually mean the delivered data will need rework.
- Responses of near-identical length and structure across different task families.
- Prompts that are all one-sentence, fully specified requests with no attached context.
- Licensed records with signatures, ticket macros, internal IDs or customer names still in the text.
- A sample hand-picked by the supplier rather than drawn the way the full delivery will be.
Need SFT targets drawn from real expert work?
Describe the task families, the kinds of records whose answers should become targets, the metadata you need on each record and the license scope. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Specify the SFT data you need.
Sources
- Ouyang et al., OpenAI (arXiv:2203.02155), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
- Zhou et al., Meta AI and collaborators (arXiv:2305.11206; NeurIPS 2023), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- Papers with Code, "LongForm (dataset page)". https://paperswithcode.com/dataset/longform
- Longpre et al. (arXiv:2310.16787; journal version in Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Chen et al. (arXiv:2307.08701), "AlpaGasus: Training A Better Alpaca with Fewer Data" (2023). https://arxiv.org/pdf/2307.08701v1
- Anthropic (Claude Help Center; provider terms, market practice), "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
- SpicyIP (blog), "Discussing Lemley and Henderson's \"The Mirage of Artificial Intelligence Terms of Use Restrictions\"" (2025). https://spicyip.com/2025/01/discussing-lemley-and-hendersons-the-mirage-of-artificial-intelligence-terms-of-use-restrictions.html
- Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- Gebru et al. (arXiv:1803.09010; Communications of the ACM 2021), "Datasheets for Datasets" (2021). https://arxiv.org/pdf/1803.09010
- Lee et al. (arXiv:2107.06499; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- Yang et al., LMSYS (arXiv:2311.04850), "Rethinking Benchmark and Contamination for Language Models with Rephrased Samples" (2023). https://arxiv.org/pdf/2311.04850v1
- Carlini et al. (USENIX Security 2021), "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
- Microsoft presidio project (pkg.go.dev), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.