Skip to content

Fine-tuning and post-training data

Fine-tuning vs RAG vs continued pre-training: what data each approach needs

Quick answer

Fine-tuning, RAG and continued pre-training consume different data and need different license rights. Supervised fine-tuning needs hundreds to tens of thousands of curated prompt-response pairs that teach format, behavior or task skill. RAG needs an indexed, access-controlled document corpus with clean metadata, licensed for indexing, caching and excerpt display. Continued pre-training needs a large volume of unlabeled in-domain text, licensed for training. Choose by objective first: retrieval for facts and freshness, fine-tuning for behavior, continued pre-training for domain fluency.

By SourceX Editorial · Updated

Which approach fits which objective

Match the approach to the gap you are closing, because each one changes a different part of the system. Fine-tuning changes weights so the model behaves differently; RAG leaves weights alone and changes what the model sees at query time; continued pre-training (also called domain-adaptive pre-training, or DAPT) shifts the model's underlying distribution toward a domain's vocabulary and structure. Market guidance puts fine-tuning sets at roughly 100 to 100,000+ curated examples plus GPU time and model versioning, while RAG leaves weights unchanged [3]. Platform guidance converges on the same split: fine-tuning for fixed output formats and style, retrieval for current, citable facts [5].

The table below is a starting hypothesis for planning, not a rule. Validate it on your own held-out evaluation set before committing a data budget.

Illustrative example: invented to show structure; it does not describe an available dataset.

ObjectiveFirst choiceData you acquireRights the license must grant
Fixed output format (JSON schema, report template)SFTInput-output pairs with schema-valid targetsTraining, derivative model use
Tone, policy adherence, refusal behaviorSFT or preference tuningPrompt-response pairs; chosen/rejected pairsTraining
Task skill (triage, extraction, summarization)SFTReal task inputs with expert outputsTraining
New or proprietary factsRAGDocument corpus with metadata and ACLsIndexing, caching, excerpt display
Facts that change weeklyRAGCorpus plus an update feedIndexing over the term, ongoing deliveries
Domain vocabulary and notation (clinical, legal, CAD logs)Continued pre-trainingLarge unlabeled in-domain textTraining at scale
Grounded answers that cite sourcesRAG plus retrieval-aware SFTCorpus plus question, document, distractor, cited-answer recordsBoth training and indexing

Can fine-tuning add new knowledge to a model?

Fine-tuning can inject facts, but it is an unreliable way to do it and can make hallucination worse. Published research [11] reports that in a controlled closed-book QA setup, examples introducing unfamiliar facts were learned more slowly and, once learned, raised the model's tendency to hallucinate. Practitioner guides building RAG fine-tuning sets from company documents make a related point: base-model answers can be verbose, hallucination-prone or off-tone, which fine-tuning fixes as behavior while the documents still supply the facts [1].

For a buyer, the practical consequence is a data-routing rule. Records whose value is the fact itself (a part number's tolerance, a policy clause, a customer's contract terms) belong in a retrieval index, where they can be updated, permissioned and cited. Records whose value is the pattern (how an expert writes a claim denial, how an engineer structures a root-cause note) belong in an SFT set. Mixing the two in one SFT file is a common failure mode: the model memorizes stale facts and learns to answer confidently about things it never saw.

What data supervised fine-tuning consumes

SFT consumes curated demonstration pairs, and quality matters more than raw count. LIMA fine-tuned a 65B-parameter LLaMA model on only 1,000 carefully curated prompt-response pairs, while noting that curating such examples is labor-intensive [4]. Cloud fine-tuning services define an SFT dataset as labeled examples that specialize a pre-trained model to a task or domain, each with a defined structure [6].

The usual delivery format is JSON Lines: UTF-8, no byte order mark, one valid JSON object per line, no blank lines [7]. Chat-style records carry a messages array with system, user and assistant roles; tool-use records add function schemas, calls and results. For format detail see the chat fine-tuning data format guide, and for sourcing see how to source supervised fine-tuning data.

What to check in an SFT purchase:

  • Target quality. Who wrote the assistant turns, and were they reviewed? An expert-written support resolution beats a model-generated paraphrase of it.
  • Input realism. Prompts should reflect the noise of production inputs: typos, partial context, forwarded threads.
  • Coverage. Count examples per intent or task type, not just in total; long-tail categories are where fine-tuned models fail.
  • Contamination. Deduplicate against your evaluation sets before training.
  • Mixture risk. Plan general-domain replay data to limit regression; see data mixtures for preventing catastrophic forgetting.

What data a RAG system consumes

RAG consumes a document corpus plus the metadata that makes retrieval precise, permissioned and auditable. The model is not trained on the corpus; it reads chunks at query time [5]. That shifts the data requirements away from labels and toward structure.

A usable RAG corpus has stable document IDs, version or effective dates, source system, document type, section structure preserved through parsing (headings, tables, lists), and access-control attributes that map to your users. Without effective dates, the retriever returns superseded policies; without ACL fields, you either over-share or cannot deploy. Scanned PDFs need OCR with layout retained, or tables collapse into unreadable token streams.

You also need evaluation data that is separate from the corpus: questions, the gold passages that answer them, and reference answers. The RAG evaluation datasets from real company documents page covers that use case; the glossary entry on retrieval-augmented generation defines the pattern.

Illustrative example: invented to show structure; it does not describe an available dataset.

{"doc_id": "pol-0412", "version": "3.2", "effective_date": "2026-03-01", "superseded_by": null,
 "source_system": "policy_repository", "doc_type": "warranty_policy", "section_path": ["4", "4.2"],
 "acl_groups": ["support_tier2", "legal"], "language": "en", "pii_treatment": "names and account numbers replaced",
 "text": "Section 4.2: Claims submitted more than 90 days after delivery require..."}

What data continued pre-training consumes

Continued pre-training consumes large volumes of unlabeled in-domain text, and its payoff is fluency in a domain's language rather than any single task. The domain-adaptive pretraining literature [12] reported gains from a second phase of in-domain pretraining across several domains, and further gains from pretraining on a task's own unlabeled text. Unlike SFT, no labels or targets are needed; the value sits entirely in coverage and cleanliness of the text.

The data shape is simple; the sourcing is not. You need volume (typically far more tokens than an SFT set), deduplication at document and near-duplicate level, boilerplate removal (email signatures, legal footers, ticket templates), and a mix that preserves general capability. Operational text such as engineering tickets, maintenance logs, contract redlines and support histories is valuable here precisely because it is underrepresented on the public web. For sourcing detail, see domain corpora for continued pre-training.

Continued pre-training is the most expensive path in compute and in rights. Treat it as justified when retrieval plus SFT has measurably plateaued on domain notation, not as a default.

Hybrid patterns: retrieval-aware fine-tuning

Most production systems combine approaches, and the hybrid needs its own training records. RA-DIT, for example, fine-tunes the language model to use retrieved context and separately tunes the retriever toward passages that help the model answer [2]. Retrieval-augmented fine-tuning recipes such as RAFT [13] go further, pairing each question with the relevant ("oracle") document, irrelevant distractors and an answer that quotes the supporting passage, so the model learns to ignore unhelpful context. Enterprise guides describe building such sets directly from a company's own documents [1].

This means a retrieval-aware SFT record is a join: question, oracle passage ID, distractor passage IDs, and a cited answer. Buying it requires rights to train on the passages, not just to index them. The retrieval-augmented fine-tuning data guide covers construction and quality checks.

Training rights versus indexing rights

License the use you will actually make, because training and indexing are different acts with different legal analyses. Training turns the data into model weights; RAG stores copies (often embeddings plus raw chunks), caches them, and displays excerpts to end users. The U.S. Copyright Office's Part 3 report on generative AI training, still a pre-publication version as of October 2026, analyzes the copying involved in training and related acts [8]. In the EU, providers of general-purpose AI models must maintain a copyright policy that honors rights reservations under Article 4(3) of the DSM Directive [10].

Rights gaps are common in practice: the Data Provenance Initiative found license information omitted for more than 70% of audited datasets on popular hosting sites and license errors above 50% [9]. A license written for one approach rarely covers another.

Illustrative example: invented to show structure; it does not describe an available dataset.

License questionSFTRAGContinued pre-training
Training on the records permitted?RequiredOnly if you also fine-tuneRequired
Persistent storage of copies and embeddings?During trainingRequired for the termDuring training
Displaying excerpts to end users?Rarely neededRequired, define length limitsNot needed
Model weights survive license expiry?Negotiate explicitlyNot applicable; index is deletedNegotiate explicitly
Ongoing updates to the data?OptionalUsually requiredOptional
Personal data treatment documented?RequiredRequired, plus ACL mappingRequired

For a broader view of terms, see the AI training data licensing guide and how procurement differs by training stage.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

How SourceX supports each data shape

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents, and finance and legal workflows, which can serve as RAG corpora, SFT material or continued pre-training text. Data is not held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, so you can specify whether you need training rights, indexing rights or both. To scope a request, start at the buyer intake.

Sourcing data for fine-tuning, RAG or continued pre-training

SourceX looks for US businesses that hold the data you describe and manages the process from assessing data and licensing permissions to agreeing pricing and allowed uses in a license. Nothing is contracted until a supplier agrees, and every release is approved by the supplying company. Describe the records, approach and rights you need at sourcex.si/buyers.

More in this cluster: fine-tuning and post-training datasets: a buyer's guide and the AI data hub.

Frequently asked questions

Is RAG always cheaper than fine-tuning?

Not on data. RAG avoids labeling, but it needs parsing, metadata, ACL mapping and an update feed for the license term, plus a separate evaluation set. SFT needs fewer records but expert-written targets.

Can I use the same records for both RAG and fine-tuning?

Yes, if the license grants both uses. Route fact-bearing documents to the index and convert expert work product into SFT pairs or retrieval-aware records [1][2]; avoid training on facts you expect to change.

When does continued pre-training beat SFT alone?

When the domain's notation or vocabulary is poorly represented in the base model and SFT gains have plateaued. Domain-adaptive pretraining research has reported gains in specialized domains, but the cost in tokens, compute and training rights is the highest of the three.

Sources

  1. Lenovo Press, "Making LLMs Work for Enterprise, Part 2: RAG Fine-Tuning Dataset Creation (LP1954)". https://lenovopress.lenovo.com/lp1954-making-llms-work-for-enterprise-part-2-rag-fine-tuning-dataset-creation
  2. LlamaIndex blog, "Improving RAG effectiveness with Retrieval-Augmented Dual Instruction Tuning (RA-DIT)". https://llamaindex.ai/blog/improving-rag-effectiveness-with-retrieval-augmented-dual-instruction-tuning-ra-dit-01e73116655d
  3. Azion Learning Center, "Fine-tuning vs RAG". https://www.azion.com/en/learning/ai/fine-tuning-vs-rag/
  4. Zhou et al. (arXiv:2305.11206), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  5. Domino Data Lab documentation, "RAG vs fine-tuning". https://docs.dominodatalab.com/en/6.0/user_guide/94beaa/rag-vs-fine-tuning/
  6. Google Cloud Vertex AI documentation, "Prepare supervised fine-tuning data for Gemini models". https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini-supervised-tuning-prepare
  7. jsonlines.org, "JSON Lines". https://jsonlines.org/
  8. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  9. Longpre et al. (Nature Machine Intelligence 6, 2024), "A large-scale audit of dataset licensing and attribution in AI". https://www.nature.com/articles/s42256-024-00878-8
  10. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  11. Gekhman et al. (arXiv:2405.05904), "Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?" (2024). https://arxiv.org/abs/2405.05904v3
  12. Gururangan et al. (arXiv:2004.10964), "Don't Stop Pretraining: Adapt Language Models to Domains and Tasks" (2020). https://arxiv.org/abs/2004.10964
  13. Zhang et al. (arXiv:2403.10131), "RAFT: Adapting Language Models for Retrieval Augmented Generation" (2024). https://arxiv.org/abs/2403.10131

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data