Data quality, coverage and contamination
Filtering Instruction-Tuning Data for Quality: What LIMA and AlpaGasus Show Buyers
Quick answer
Instruction-tuning data quality matters more than raw volume once a base model is strong, but quantity does not stop mattering. LIMA fine-tuned a 65B LLaMA on 1,000 curated prompt-response pairs and stayed competitive in human preference tests; AlpaGasus kept about 9k of 52k Alpaca examples after LLM scoring and beat the full set. For buyers, the lesson is to specify quality criteria, scoring evidence and filtering rights in the deal, then validate every threshold on your own held-out evaluation set.
By SourceX Editorial · Updated
What LIMA actually showed about 1,000 examples
LIMA showed that a small, carefully curated SFT set can produce a strong assistant when the base model is already capable, not that data volume is irrelevant. The authors fine-tuned a 65B-parameter LLaMA on 1,000 prompt-response pairs with standard supervised loss and no reinforcement learning or preference modeling [1]. In their human study, LIMA's responses were judged equivalent to or better than GPT-4's in 43% of cases [1].
The paper's own framing is the important part for buyers. The authors argue that almost all knowledge is learned during pretraining and that instruction tuning mostly teaches the format and style of interaction, which they call the superficial alignment hypothesis [1]. They also state the costs: curating examples at that standard is labor-intensive, and LIMA is less robust than product-grade models [1].
Read the result as a statement about what SFT does, not a promise that 1,000 rows will close any capability gap. If your target behavior depends on domain knowledge the base model lacks, such as reading claim adjudication notes or interpreting PLC fault logs, a style-only SFT set will not supply it. For that case, see how teams source supervised fine-tuning data that carries real domain content.
How AlpaGasus scored and filtered 52k Alpaca pairs
AlpaGasus showed that an LLM grader with a simple threshold can remove most of a noisy instruction set and improve the resulting model. The authors prompted ChatGPT to rate each of the 52k Alpaca (instruction, input, response) triplets on a 0-5 scale for accuracy, then kept only high-scoring examples, about 9k in total [2]. The 7B and 13B models trained on the filtered subset outperformed the same models trained on the full 52k in their evaluations [2].
The efficiency gain is concrete. Training time for the 7B model dropped from 80 minutes to 14 minutes, because the model saw roughly a sixth of the data [2]. The publication page reports that the filter generalizes across datasets, base models and grader models [3].
Two cautions apply when you copy the recipe. First, Alpaca was itself machine-generated with self-instruct, so it contained many wrong or truncated answers; a filter that removes 80% of a synthetic set may remove far less of a carefully written human set. Second, the grader's scores are not ground truth, which is why the scoring method needs its own validation; see using an LLM judge to score training data.
Does the same pattern hold for multimodal instruction data?
The same less-is-more pattern has been reported for multimodal instruction data, with the same caveats. MM-LIMA applied selection to vision-language instruction data and fine-tuned on 200 selected samples, reporting gains over training on the larger unfiltered set [4]. That supports filtering as a general step, but the selection signals differ: image-text alignment, grounding of the answer in the image, and caption hallucination all need their own checks.
If you buy image-instruction-response sets, treat the 200-sample figure as evidence that selection helps, not as a sizing target. The sourcing side is covered in visual instruction tuning data.
Where "quality is the only factor" goes wrong
Quantity still matters; the evidence says that low-quality examples cost more than they add, not that volume is worthless. Some secondary summaries describe LIMA as proving that quality is virtually the only factor in fine-tuning data [5]. That overstates a paper whose authors flagged limited robustness and expensive curation [1].
Volume and diversity still drive outcomes in at least four situations:
- Coverage of task types. A 1,000-example set cannot cover every ticket category, document type and refusal case a production assistant meets.
- Domain knowledge transfer. When the base model has not seen the domain, more in-domain examples can carry knowledge that style tuning cannot.
- Structured outputs. Teaching a strict JSON schema with dozens of fields usually needs many varied examples per field pattern; see structured-output fine-tuning data.
- Multilingual behavior. Each language and register needs enough native examples, which the non-English instruction data page covers.
The defensible position is that quality sets the ceiling on what each example contributes, and diversity and quantity decide how much of your deployment distribution you reach.
How to score instruction-response pairs before training
Score purchased SFT data in layers: cheap deterministic checks first, then model-based scoring, then human review of a sample. This mirrors the multi-method approach described in recent data-centric work, which combines metric checks for format, duplication, diversity, fluency and factual accuracy with LLM scoring and human review [8]. ISO/IEC 5259-4 offers a process framework for organizing these steps for supervised ML, including labelling [9].
Deterministic checks catch defects that no grader should spend tokens on. Validate that every file is JSON Lines: UTF-8 with no byte order mark, one valid JSON value per line and no blank lines [7]. Then check that each record matches the target trainer's schema, such as the role-and-parts message structure that Vertex AI expects for Gemini SFT [6].
Common low-quality patterns in SFT datasets include:
- Truncated or empty responses, and responses that restate the prompt.
- Answers that contradict the input field, for example a summary that invents an order number.
- Templated boilerplate ("As an AI language model...") or agent signatures copied from support macros.
- Near-duplicate prompts that inflate apparent size; see near-duplicate detection with MinHash and LSH.
- Unredacted personal data in prompts or targets; see scanning a training corpus for PII before fine-tuning.
- Benchmark items leaked into training data, which inflate evaluation scores.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Layer | Check | Typical rule | Action on fail |
|---|---|---|---|
| Format | JSONL validity, encoding, schema | Parse every line; reject BOM, blank lines, missing roles | Reject record |
| Hygiene | Length, truncation, refusal boilerplate | Response under 20 tokens or ends mid-sentence | Reject or flag |
| Duplication | Exact and near-duplicate prompts | Jaccard over 0.8 on 5-gram shingles | Keep one per cluster |
| Grounding | Response consistent with input context | Entities in response present in input | Flag for review |
| Model score | LLM grader, 0-5 rubric for accuracy and helpfulness | Keep at or above a validated threshold | Drop below threshold |
| Human audit | Stratified sample by task type and score band | Reviewer agreement with grader recorded | Recalibrate threshold |
A record that carries its scoring history makes later audits possible:
Illustrative example: invented to show structure; it does not describe an available dataset.
{"id": "sft-000123", "task_type": "support_resolution", "messages": [{"role": "user", "content": "Customer reports duplicate charge on invoice [INVOICE_ID]..."}, {"role": "assistant", "content": "Confirm both charges in the billing ledger, then..."}], "source_system": "helpdesk_export", "redaction": "pii_replaced_v2", "grader_model": "judge-a", "grader_rubric": "acc-help-0to5-v3", "grader_score": 4.5, "human_reviewed": true, "human_label": "accept", "near_dup_cluster": null}
Validate thresholds on your evaluation set, not the grader's scores
Pick a filter threshold by its effect on your own held-out evaluation, because a grader's score distribution says nothing direct about downstream model quality. Train small proxy runs at several cut-offs, for example keeping data scored at or above 3.5, 4.0 and 4.5, and compare them on a fixed evaluation set that reflects your deployment tasks. The best cut-off for a noisy synthetic set will often be stricter than for a set written by domain experts.
Keep the evaluation set separate from anything the grader or supplier has seen, and freeze it before you look at candidate data. Check per-task-type results, not just the average: aggressive filtering often removes hard, long or unusual examples that a grader rates lower, which quietly shrinks long-tail and edge-case coverage. Record the grader model, prompt and rubric version, since changing any of them changes which rows survive.
What this means when you buy SFT data
The buying implication is to contract for quality criteria, scoring evidence and the right to filter, rather than paying for raw row counts. If you expect to discard a large share after scoring, the price, acceptance terms and delivery format should reflect that before signature. These are the terms to raise with any supplier:
Illustrative example: invented to show structure; it does not describe an available dataset.
- Unit of quality. Define what an acceptable pair is: task types, minimum response completeness, grounding in source records and the rubric used.
- Scoring disclosure. Ask for per-record scores, the grader model and rubric version, and the human audit sample with agreement rates; see what a dataset quality report should contain.
- Filtering and rejection rights. Confirm that the license lets you score, filter, drop and modify records, and state how rejected records are treated commercially; acceptance sampling for dataset deliveries gives a sampling method.
- Allowed use. Make sure scoring with a third-party grader model is itself a permitted use, since it sends records to another system; compare fine-tuning-only data licenses.
- Provenance per record. Require a source-system field so you can trace bad clusters back to one export or one annotator pool.
Instruction pairs built from real operational records, such as support resolutions or engineering change notes, often need more grounding checks than synthetic sets, because the "response" was written for a customer, not a model. The conversion steps are covered in turning business records into instruction-response pairs, and definitions of instruction tuning and supervised fine-tuning are in the glossary. For broader context, start at the data quality hub or the AI data guides.
SourceX sources operational datasets, such as support and sales histories, engineering records and documents, from US companies on request, and every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Teams planning domain SFT work can review training data for domain-specific fine-tuning or describe the instruction data they need.
Sourcing instruction-tuning data with quality terms built in
SourceX finds US businesses that hold the operational data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Describe the instruction data and quality criteria you need at SourceX for buyers.
Sources
- Zhou et al., Meta AI and collaborators (arXiv), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- Chen et al. (arXiv), "AlpaGasus: Training A Better Alpaca with Fewer Data" (2023). https://arxiv.org/pdf/2307.08701v1
- Samsung Research America, "AlpaGasus: Training a Better Alpaca with Fewer Data (publication page)". https://sra.samsung.com/publications/alpagasus-training-a-better-alpaca-with-fewer-data
- arXiv, "MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets" (2023). https://arxiv.org/pdf/2308.12067
- Meta Intelligence, "Fine-tuning data: quality over quantity (insight article)". https://meta-intelligence.tech/en/insight-finetuning-data
- Google Cloud, "Prepare supervised fine-tuning data for Gemini models". https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini-supervised-tuning-prepare
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- arXiv, "Towards Next-Generation LLM Training: From the Data-Centric Perspective" (2026). https://arxiv.org/pdf/2603.14712
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality for analytics and machine learning - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.