Skip to content

Fine-tuning and post-training data

Targeted data acquisition: turning model failures into data requests

Quick answer

Targeted data acquisition means buying or collecting only the records that fix specific, measured failures in a fine-tuned model, instead of adding more general data. Cluster production and evaluation errors into a failure taxonomy, translate each cluster into a slice specification (record type, required fields, conditions, volume, label), lock a held-out test slice before any data arrives, then license or collect against that spec. Research calls this data-centric debugging [1]; practitioners call the search step failure mining [2].

By SourceX Editorial · Updated

Why generic data purchases rarely fix a specific failure

A fine-tuned model usually fails on narrow slices, so adding a broad dataset mostly reinforces what it already does well. If your support-ticket assistant mishandles refund disputes involving partial shipments, another 50,000 general tickets may contain a few dozen relevant cases and dilute the training signal for the rest.

Singla et al. showed the alternative on image classifiers: start from a small set of failure samples, retrieve similar examples from a large pool, and train on those. They reported gains on held-out debug sets and noted that new data for rare conditions is expensive to obtain [1]. LIMA makes a related point for post-training: 1,000 carefully curated prompt-response pairs were enough to substantially shape a 65B model, though curation was labor-intensive [3]. For a targeted fix, precision of selection often matters more than raw volume.

For background on how much data a fine-tuning run needs overall, see how much data you need to fine-tune an LLM. This page assumes you already have a model in production or late evaluation and a backlog of failures.

Building a failure taxonomy from production and evaluation errors

The taxonomy is a labeled list of failure clusters, each defined by an observable input condition and a wrong output behavior. Without it, every data request becomes "more of the same domain," which suppliers cannot act on.

Pull failures from at least three streams:

  • Production signals: thumbs-down events, escalations to human agents, regenerations, edits users made to model outputs, and guardrail or schema-validation rejections (for example, JSON that fails your Pydantic or JSON Schema validator).
  • Evaluation runs: items failed in your golden set, regression suites and LLM-as-judge rubrics, with judge rationales saved.
  • Uncertainty and disagreement: low log-probability answers, disagreement between sampled outputs or between your model and a reference model. Failure mining uses uncertainty and disagreement signals like these to decide what to collect, relabel or retrain [2].

Then cluster. Embed the failing inputs, cluster them (HDBSCAN or k-means over sentence embeddings works), and have a domain reviewer name each cluster with a condition, such as "multi-currency invoice with credit note attached," not a symptom such as "wrong total." Tag each cluster with a root-cause hypothesis: missing coverage, wrong label convention, formatting, reasoning depth, or policy ambiguity. Only missing coverage and reasoning depth are reliably fixed by buying data; label convention and policy problems are usually fixed in your own guidelines.

Translating each failure slice into a purchasable data specification

A purchasable spec describes records a real business would hold, not model behavior. Suppliers cannot search for "cases where the model hallucinates a policy"; they can search for "closed chargeback tickets where the resolution cites a written refund policy clause."

For each slice, write down:

  • Record type and source system: for example, Zendesk or Salesforce Service Cloud tickets, Jira issues with linked commits, AP invoices from NetSuite or SAP, or contract redlines from a CLM.
  • Inclusion conditions: the attributes that define the slice (status = resolved, has attachment, language, product line, date range).
  • Required fields: what must be present for the record to become a training example, such as the full message thread, the final resolution, the agent macro used, timestamps.
  • Target output format: SFT pairs, multi-turn chats or preference pairs. Vertex AI and similar platforms define a per-example structure for SFT data, so state the schema you will convert into [6].
  • Volume band and diversity constraints: a range, plus caps per customer, per agent or per month so one source does not dominate.
  • Label or judgment needed: whether raw records suffice or experts must write ideal responses or rank candidates.

Our guide on turning model goals into training data requirements covers requirements for a new capability; this page is narrower, starting from observed errors. For the general mechanics of a request document, the owner guide on how to write a data request for suppliers applies.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldSlice FS-07: partial-shipment refund disputes
Observed failureModel approves full refunds when only part of an order shipped; 31% error on this slice vs 6% overall in the internal eval
Root cause hypothesisMissing coverage: fewer than 40 training examples with split shipments
Record typeResolved e-commerce support tickets with linked order and shipment records
Inclusion conditionsOrder has 2+ shipments; customer requested refund; ticket resolved by a human agent; English; last 24 months
Required fieldsFull thread, order line items, shipment status per line, refund amount issued, policy reference, resolution code
Target formatMulti-turn SFT conversations ending in the agent's resolution; 20% also as preference pairs (agent answer vs model answer)
Volume band1,500 to 4,000 tickets; no single merchant above 15%
Personal data handlingNames, emails, addresses, order numbers replaced with consistent placeholders before delivery
Held-out slice300 internal tickets frozen before purchase; never shared with any supplier
Success metricSlice error below 12% with no more than 1 point regression on the general eval

Choosing SFT, preference data or both for each slice

Match the data type to the failure mode. When the model lacks knowledge of how a case should be handled, demonstrations fix it, so buy SFT records where a competent human produced the correct output. When the model produces plausible answers but picks the worse of two acceptable styles, or ignores a constraint under pressure, preference pairs are the more efficient fix.

Direct Preference Optimization trains on chosen/rejected pairs without a separate reward model [4], which makes "agent's actual answer vs your model's answer on the same input" a practical construction from business records. Details on structure are in our guide to preference datasets for DPO. For slices where correct output is a strict format, such as extracted fields, see structured-output fine-tuning data.

A common failure mode is buying preference data for a coverage gap. If the model has never seen split-shipment logic, ranking two wrong answers teaches little.

Protecting the held-out slice and avoiding tuning to your evaluation set

Freeze a held-out test slice for each failure cluster before you send any request, and never share it. If a supplier sees your failing examples and returns near-duplicates, you will measure memorization, not repair.

Three controls matter:

  1. Describe, do not send, failing items. Give suppliers the slice definition and a few synthetic or heavily altered exemplars. Never send held-out items.
  2. Deduplicate on arrival. Run exact hash and near-duplicate checks (MinHash over n-grams, or embedding cosine above a threshold you set) between delivered records and every eval set.
  3. Separate who builds training data from who builds evaluation data. Research on private evaluation curators describes the conflict when the same party shapes both what a model learns and how it is scored [5].

Also keep a general regression suite. A slice fix that costs two points elsewhere is often a bad trade, and it shows up only if you measure outside the slice.

Sourcing options: mine your own logs, collect new, or license existing records

Check your own data first, then decide between new collection and licensing records another business already holds. Active learning over your unlabeled production logs (rank by uncertainty, send the top items for expert labeling) is the cheapest path when the slice exists in your traffic.

When it does not, the slice often lives in another company's operational systems. Licensing existing records gives real distribution and edge cases; custom collection gives control over conditions but can look staged. Our comparison of custom data collection vs licensing existing records covers the tradeoffs, and prompt-heavy slices are covered in real-world prompt sets for post-training.

Whatever the source, insist on documented provenance and license terms per dataset. The Data Provenance Initiative found license omissions above 70% and license errors above 50% on popular dataset hosting sites [9], so a slice that looks perfect may be unusable commercially.

Accepting the delivery and measuring the fix

Acceptance has two gates: does the delivery match the spec, and does training on it close the slice. Treat them separately so a supplier is judged on what they controlled.

For spec conformance, sample records against inclusion conditions and required fields; our guide on acceptance sampling for dataset deliveries adapts AQL plans to record defects. ISO/IEC 5259-4 offers a process framework for data quality across training and evaluation data, including labelling [7].

For the model outcome, retrain or continue training with the new slice mixed at a fixed ratio, then report:

  • Slice error on the frozen held-out set, before and after
  • General eval delta and any safety or refusal regressions
  • Per-source breakdown if multiple suppliers contributed

Log each cycle (failure cluster, spec version, data received, metric movement) so the taxonomy becomes a running record. The NIST AI RMF Playbook's Measure and Manage functions give a voluntary structure for documenting that loop [8]. Before signing for a larger volume, use our checklist for evaluating a fine-tuning dataset before buying.

Where SourceX fits in failure-driven sourcing

SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records, documents and finance and legal workflows, and manages the licensing process. Data is not held in stock, so a slice spec like the one above is what you bring: you describe the records, not the businesses, and a request does not guarantee a match. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with personal details removed or replaced and the method recorded, though no method is perfect. Teams anywhere can describe a failure slice to SourceX.

Request records for a specific model failure

If your error analysis points to a slice that only exists in another company's operational systems, write it up as a record specification and share it. SourceX looks for US businesses that hold the described data, and nothing is contracted until a supplier agrees and approves the release. Start at the SourceX buyer page.

For the wider cluster, see the fine-tuning and post-training data hub and the AI data guides.

Sources

  1. Singla et al., arXiv, "Data-Centric Debugging: mitigating model failures via targeted data collection" (2022). https://arxiv.org/pdf/2211.09859
  2. Voxel51, "Failure mining". https://voxel51.com/glossary/failure-mining
  3. Zhou et al., arXiv, "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  4. Rafailov et al., arXiv, "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  5. arXiv, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
  6. Google Cloud, "Prepare supervised fine-tuning data for Gemini models". https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini-supervised-tuning-prepare
  7. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality for analytics and ML - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
  8. NIST, "NIST AI RMF Playbook". https://airc.nist.gov/airmf-resources/playbook/
  9. Longpre et al., Nature Machine Intelligence 6 (2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI". https://www.nature.com/articles/s42256-024-00878-8

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data