Skip to content

Fine-tuning and post-training data

How much data do you need to fine-tune an LLM?

Quick answer

How much data you need to fine-tune an LLM depends on what the fine-tune must change. Output format and response style can be taught with about a thousand curated examples: LIMA used 1,000 [1]. A new task skill needs enough distinct, verified examples to cover the task's real variation, and published instruction sets sit in the thousands to tens of thousands [2][3]. Preference tuning is sized in distinct prompts; new domain knowledge needs continued pre-training or retrieval, and published runs use hundreds of millions of tokens or more [4][5].

By SourceX Editorial · Updated

Size follows the objective, not the training method

Each fine-tuning objective has its own counting unit, so decide what to count before you ask anyone for a quantity. Here an "example" is a prompt with its target response or, for preference data, a prompt with a chosen and a rejected response.

ObjectiveWhat the model must learnCount thisPublished reference pointWhat pushes the number up
Output format or schema (JSON fields, section order, citation style)A consistent surface patternDistinct examples covering every field and edge case1,000 curated pairs taught response format and style to a 65B base model [1]Optional and nested fields, rare null cases
Tone, register or personaHouse wording and structureDistinct examples per document type and audienceNo separate figure; LIMA's result covers style as well as format [1]Many document types, mixed audiences
Narrow task skill (classify, extract, route, triage)A decision boundary or field mappingExamples per label or field, with a floor for the rarestNo universal figure; set by label count and the long tailMany labels, class imbalance, ambiguous policy
Broad instruction followingMany task typesDistinct prompts per task categoryAbout 9,000 filtered examples beat the full 52,000-example Alpaca set [2]; InstructGPT's SFT set had about 13,000 training prompts [3]Breadth of categories, multi-turn, safety behavior
Preference alignment (DPO, reward models)Which of two responses is betterDistinct prompts, then pairs per promptInstructGPT's reward-model set had about 33,000 training prompts [3]Subtle quality differences, low annotator agreement
Domain knowledge (vocabulary, facts, document conventions)New distributional knowledgeDeduplicated tokens400 million tokens of SEC filings in one financial study [4]; 3.3 billion words of filings in another [5]Breadth of the domain, need to keep general skills

The last row is a different kind of purchase. LIMA's authors concluded that almost all of a model's knowledge is learned in pre-training and that instruction data mostly teaches the format of responses [1]. If the model must know things it never saw, budget for continued pre-training or retrieval rather than more instruction pairs; see data needs for fine-tuning, RAG and continued pre-training. Images and agents use other units (VLM data volume, agent trajectories).

What the small-data results prove, and what they do not

The small-data papers show that curation beats volume on a strong base model; they do not show that a few hundred examples will hit your accuracy target on a narrow business task. LIMA fine-tuned a 65-billion-parameter LLaMA model on 1,000 curated prompt-response pairs with no reinforcement learning, and in a human study its responses were rated equal to or better than GPT-4's in 43% of cases [1]. The authors also note that curating such examples is labor-intensive and that LIMA is less robust than product-grade models [1].

AlpaGasus reached a similar result by filtering instead of writing. Its authors found that Alpaca's 52,000-example instruction set contained many low-quality instances with incorrect or irrelevant responses. They had ChatGPT score each example from 0 to 5 and kept about 9,000 above a threshold, and the resulting model outperformed the original Alpaca in GPT-4 and human evaluations [2]. Training time for the 7B variant fell from 80 minutes to 14 [2].

Three limits matter when you turn these results into a purchase quantity:

  • Metric mismatch. Both studies mainly judged general assistant quality by preference, not by field-level extraction accuracy, routing F1 or policy compliance.
  • Base model. Both used strong base models. The floor is low for style; the results say nothing about how many examples your rarest label needs.
  • Judge dependence. An AlpaGasus-style filtered set is defined by its grader and cutoff; a different grader or threshold keeps a different subset. If a supplier quotes a "filtered" count, ask which grader and threshold produced it.

So pay for fewer verified examples rather than a larger unreviewed set, and size on the usable count left after deduplication, filtering, de-identification and label review.

Fine-tuning with a small dataset: where it breaks

Small fine-tuning sets fail in three predictable places: rare categories, memorization and evaluation noise. Which one limits you decides what extra data to buy.

  • Rare categories. With 30 routing queues and 450 examples, several queues may appear fewer than five times. Size by the rarest label the model must handle, not by the average.
  • Memorization. Small sets get more epochs, and duplicated records are seen more often still. Lee et al. found near-duplicates common in language-modeling datasets; deduplicated training emitted memorized text about ten times less often and reached the same or better accuracy in fewer steps [6]. A separate study reports that fine-tuning can amplify a model's privacy risks, including disclosure of personal information [7]. De-identify records before training rather than relying on output filters afterward.
  • Evaluation noise. A small held-out set cannot separate a real gain from run-to-run variance (see the learning-curve section).

A minimum dataset size enforced by a fine-tuning service only makes a job run; it does not predict quality. For the supplier-side view, see the minimum dataset size AI buyers accept, how many records AI labs want and the data volume estimator.

LoRA and QLoRA lower the compute bill, not the example count

LoRA and other parameter-efficient fine-tuning (PEFT) methods cut the number of trained weights and the GPU memory needed; they do not reduce how many distinct situations the model must see. LoRA freezes the base model and trains small low-rank adapter matrices. QLoRA also stores the frozen base model in quantized form to cut memory further. If 8 of your 30 queues are missing from the data, an adapter will not learn them any more than a full fine-tune would.

PEFT does change the sizing plan in three ways:

  • Pilots become cheap. Training a separate adapter on each slice of a pilot sample costs little, so measure the learning curve (described below) instead of guessing.
  • Overfitting shows up early. On small sets, track held-out loss per epoch and keep the checkpoint where it bottoms out.
  • Sizing can be per task. One base model can carry several adapters, so you can size extraction data and reply-drafting data separately instead of buying one blended set.

How many examples for DPO: count prompts first, then pairs

For DPO, the number that matters most is distinct prompts, because pairs drawn from the same prompt are correlated. DPO fits the policy directly to records of a prompt, a preferred response and a rejected response, without training a separate reward model [8]. InstructGPT's labelers ranked between 4 and 9 responses per prompt, which yields up to 36 pairwise comparisons from one prompt. Because those comparisons are highly correlated, the authors trained on all comparisons from a prompt as a single batch element to avoid overfitting [3].

Size a preference purchase in this order; the record structure is in preference datasets for DPO.

  1. Prompt categories. List the task types and difficulty levels the model must handle, ideally from real-world prompt sets.
  2. Distinct prompts per category. Allow enough that a held-out slice per category can still be scored.
  3. Pairs per prompt. Cap them so a few heavily annotated prompts cannot dominate the set.
  4. Pair quality. Specify the minimum margin between chosen and rejected, tie handling and acceptable annotator agreement.
  5. Policy match. Pairs generated by a model very different from yours carry a weaker signal; see on-policy vs off-policy preference data.

Tokens needed for fine-tuning: three counts to keep apart

Count tokens as well as examples, because suppliers may quote by record or by token, compute scales with tokens processed, and example lengths vary widely. A routing example can be a few hundred tokens with a one-word label, while a multi-turn support thread with a drafted reply can run to thousands.

CountWhat it measuresWhy it matters
Raw tokensEverything in the delivered recordsWhat a per-token price refers to
Supervised tokensTokens that receive loss, usually assistant turns only in chat SFTHow much target behavior the model actually sees
Tokens processedRaw training tokens times epochsWhat drives GPU time and cost

Loss masking is explained in chat fine-tuning data format. Token counts also depend on the tokenizer. In one language-adaptation study, replacing about 5,000 tokens, roughly 10% of the vocabulary, improved fertility (tokens per word) by 42% for Hungarian and 73% for Thai [9]. Ask suppliers to count with your model's tokenizer, or to report word and character counts you can convert.

How many tokens for continued pre-training

Continued pre-training is sized in deduplicated tokens, and published domain runs range from hundreds of millions to billions. One preprint on data-efficiency scaling laws for financial models built a 400-million-token corpus from 10-K, 10-Q and DEF 14A filings and kept narrative sections such as MD&A and Risk Factors. MinHash deduplication removed about 1.9% of its tokens [4]. An earlier continual pre-training study used 3.3 billion words of SEC filings, 16.5% of its financial corpus, with the rest taken from financial news in Common Crawl [5].

A continued pre-training budget has four inputs:

  • Domain tokens after text extraction, boilerplate removal and deduplication, counted with your tokenizer.
  • Replay tokens of general-domain text if the model must keep its general skills; mixing ratios are covered in data mixtures that prevent catastrophic forgetting.
  • Epochs over the domain set.
  • Compute. GPU-hours ≈ (domain tokens + replay tokens) × epochs ÷ (measured tokens per second per GPU × 3,600). Measure throughput in a short pilot run at your sequence length and hardware rather than borrowing a published figure.

Sourcing is covered in domain corpora for continued pre-training; new-language adaptation has its own corpus sizing guide.

Worked example: sizing a support-triage fine-tune

A worked plan turns objectives into usable counts, a raw request quantity and a token budget for a mixed SFT and DPO project.

Illustrative example: invented to show structure; it does not describe an available dataset.

A team fine-tunes an 8B open-weight model with LoRA to route inbound B2B support tickets to 30 queues and draft a first reply in the support team's register. It then runs DPO on supervisor edits.

ComponentCounting unitPlanning rule (the team's own choice)Usable records
Routing SFTTicket text → queue labelFloor of 120 per queue × 30 queues3,600
Reply SFTThread → first agent reply60 per queue, stratified1,800
DPO pairsPrompt, chosen (supervisor-edited), rejected (original draft)900 distinct prompts, one pair each900
Held-out evaluationFrozen, never trained on20 per queue plus 150 reply prompts750
Total usable7,050
  • Raw request. The pilot showed 72% of raw records surviving deduplication, de-identification QA and label-conflict review, so the request is 7,050 ÷ 0.72 ≈ 9,800 raw records.
  • Tokens. Routing examples average 400 tokens; reply examples average 1,600 tokens of thread plus 250 of reply. Raw SFT training tokens are 3,600 × 400 + 1,800 × 1,850 ≈ 4.8 million, and three epochs process about 14.3 million. Supervised tokens on the reply set are only 1,800 × 250 = 450,000.
  • Reading. Compute is trivial at this scale; cost sits in verified records per queue, especially rare queues, and that is the number to negotiate.

Run a learning curve before buying full volume

A learning curve on a pilot sample tells you whether more of the same data will help, which no rule of thumb can. Train on 10%, 25%, 50% and 100% of the pilot, score every run on the same frozen held-out set, and plot the metric against usable examples per category. A curve still rising at 100% says more of the same will likely pay. A flat curve says buy different data, such as new categories, harder cases or cleaner labels, rather than more.

The held-out set must be large enough to show the differences you care about. One practitioner analysis estimates that a 100-example set can detect only differences of about 10 to 12 points at 80% power [10]. An ICML 2025 position paper argues that standard CLT-based confidence intervals produce error bars that are too narrow below a few hundred datapoints [11]. Size it with eval set sizing for statistical power; sizing a data purchase turns the curve into tranches, and SourceX's guide to running a data pilot with a supplier covers the pilot.

Sizing checklist for a fine-tuning data request

A request that states the counting unit, the coverage floor and the usable-record definition gets comparable answers from suppliers. Use this list before you ask for a quantity:

  • Objective for each component (format, task skill, preference, knowledge) and the method (full SFT, LoRA or QLoRA, DPO, continued pre-training)
  • Base model and tokenizer, so token counts are comparable
  • Counting unit per component: examples, distinct prompts, a cap on pairs per prompt, or deduplicated tokens
  • Coverage targets per category or label, with a floor for the rarest
  • Usable-record definition: exact and near-duplicate rules, quality filters, de-identification, label-conflict handling
  • A held-out evaluation slice that no training delivery will include
  • Pilot sample size and the learning-curve slices
  • An option for more volume of the same specification if the curve is still rising
  • License scope covering the method and the resulting adapters or models; see fine-tuning-only data licenses
  • License status of any open instruction data used to top up volume. The Data Provenance Initiative found license omission rates above 70% and error rates above 50% on popular dataset hosting sites [12]; see open datasets that allow commercial fine-tuning

Many of the records that teach business tasks sit inside companies: support and sales histories, engineering records, documents, and finance and legal workflows. SourceX sources operational datasets like these from US companies on request rather than holding them in stock, and a request does not guarantee a matching dataset. Every dataset goes through rights review and is delivered under a license that defines which records are included and what they can be used for. You can send SourceX the counts and coverage targets from your sizing plan, and the fine-tuning data hub and SFT sourcing guide cover what to specify beyond volume.

Turn your sizing estimate into a data request

On the SourceX buyer page you can submit a data request with the objectives, counting units and coverage floors from your plan, or talk to SourceX about them first. SourceX looks for US companies that hold matching records, checks their licensing permissions, and manages the license and delivery. Send SourceX your fine-tuning data specification.

Sources

  1. Zhou et al. (Meta AI and collaborators), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  2. Chen et al., "AlpaGasus: Training A Better Alpaca with Fewer Data" (2023). https://arxiv.org/pdf/2307.08701v1
  3. Ouyang et al. (OpenAI), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  4. arXiv preprint 2512.12384, "The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining" (2025). https://arxiv.org/pdf/2512.12384
  5. arXiv preprint 2311.08545, "Efficient Continual Pre-training for Building Domain Specific Large Language Models" (2023). https://arxiv.org/pdf/2311.08545
  6. Lee, Ippolito, Nystrom, Zhang, Eck, Callison-Burch, Carlini, "Deduplicating Training Data Makes Language Models Better" (2021; ACL 2022). https://arxiv.org/abs/2107.06499v1
  7. arXiv preprint 2310.15469, "The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks" (2023). https://arxiv.org/pdf/2310.15469
  8. Rafailov, Sharma, Mitchell, Ermon, Manning, Finn (Stanford), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  9. arXiv preprint 2311.05741, "Efficiently Adapting Pretrained Language Models To New Languages" (2023). https://arxiv.org/pdf/2311.05741
  10. Tian Pan (practitioner blog), "Statistical power for LLM evals" (2026). https://tianpan.co/blog/2026/04/15/statistical-power-llm-evals
  11. ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  12. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version in Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data