Procurement, samples and ongoing supply
How Much Training Data Should You Buy? Sizing a Data Purchase
Quick answer
Buy the volume at which your own learning curve flattens, plus enough headroom to survive deduplication, de-identification and quality filtering, and no more in a single commitment. Estimate that point by training on two to four pilot slices of the supplier's data, fitting a curve, and reading off where the next doubling stops paying for itself. Size the evaluation set separately, by the confidence interval you need per slice, and license the rest in tranches with an option to extend.
By SourceX Editorial · Updated
This guide sits in the AI training data procurement hub and gives the ML lead a method to hand procurement. For task-specific rules of thumb on instruction tuning, see how much data you need to fine-tune an LLM; for what suppliers typically see in requests, see how many records AI labs want.
Why a round number is the wrong starting point
A round number such as "100,000 conversations" fixes cost before anyone knows whether example 50,001 improves the model. Published results span orders of magnitude: one 2024 study found most models peaked or nearly peaked at about 60 training samples on a question-answering task [4], while domain adaptation to a broad body of specialized records may keep improving over far larger volumes. The right volume depends on the base model, the task's diversity, the fine-tuning method (full fine-tuning or LoRA) and how far the supplier's distribution sits from your pretraining mix.
Quality also changes the curve. AlpaGasus filtered the 52,000-example Alpaca instruction set down to a much smaller high-scoring subset and reported a model that beat the one trained on the full set [3]. For a buyer, that means two offers with the same row count can sit on very different curves, and the larger one is not automatically the better purchase. The supplier's minimums matter too; what minimum dataset size AI buyers accept covers that side of the negotiation.
Fit a learning curve on pilot slices before committing volume
A learning curve fitted on pilot slices turns "how much data" into a measurable question with a cost per point of improvement. Hestness and colleagues showed across translation, language modeling, vision and speech that generalization error falls as a power law in training-set size, after a small-data region and before an irreducible-error floor [1]. The practical form is error(n) ≈ a · n^(−b) + c, where c is the floor no amount of this data will remove.
Run it like this:
- Get a pilot slice under an evaluation license (see evaluation licenses and NDAs for dataset samples) that is a random draw, not the supplier's showcase records.
- Train on nested subsets at roughly geometric sizes, for example 1k, 2k, 4k and 8k usable examples, holding hyperparameters, base checkpoint and token budget per example fixed.
- Score every run on the same frozen, held-out eval set built from your own task, never from the supplier's pilot.
- Fit a, b and c with nonlinear least squares in log space, and repeat each point with two or three seeds so seed noise does not masquerade as a trend.
- Extrapolate one or two doublings past the largest slice, no further, and compute marginal gain per additional 1,000 usable examples.
The usual failure modes are predictable. Teams fit on two points (any two points make a line), change the learning-rate schedule between slices, evaluate on data that overlaps the training slices, or extrapolate twenty-fold from a curve still in its small-data region. The guide on estimating a dataset's value to your model before you buy covers ablations that complement this curve.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Usable examples | Task error (mean of 3 seeds) | Gain vs previous slice |
|---|---|---|
| 1,000 | 31.0% | n/a |
| 2,000 | 26.4% | 4.6 pts |
| 4,000 | 23.3% | 3.1 pts |
| 8,000 | 21.3% | 2.0 pts |
| 16,000 (extrapolated) | about 20.0% | about 1.3 pts |
| 32,000 (extrapolated) | about 19.2% | about 0.8 pts |
In this invented case the fitted floor c sits near 17.5%, so the second doubling past the pilot buys under a point. If one point of task error is worth less to the product than the price of 16,000 more usable records, the first tranche should stop near 16,000, with an option on more.
Convert usable examples into a raw purchase volume
Buy against usable examples, because raw rows shrink at every preparation step between delivery and training. The loss stack for operational records usually includes exact and near-duplicate removal (MinHash or embedding clustering on support tickets catches templated replies and forwarded threads), records dropped because they cannot be de-identified cleanly, language or format rejects, and quality filtering against your rubric. One acceptance-audit provider claims filtering typically discards 20-30% of a collected dataset; treat that as a self-reported planning figure, not a benchmark [5].
De-identification losses depend on the data. Health records handled under HIPAA must meet Expert Determination or Safe Harbor, and Safe Harbor requires removing 18 listed identifier types [7]. In free-text clinical notes, that removal can leave some records too sparse to use. Measure each loss on the pilot rather than assuming it.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Step | Rate measured on pilot | Records remaining (from 25,000 raw) |
|---|---|---|
| Delivered raw records | n/a | 25,000 |
| Exact and near-duplicate removal | 9% | 22,750 |
| Unusable after de-identification | 4% | 21,840 |
| Wrong language, empty or malformed | 3% | 21,185 |
| Quality filter against task rubric | 22% | 16,524 |
The yield is about 66%, so a 16,000-usable target needs roughly 24,000 to 25,000 raw records. Ask suppliers to quote both raw and expected usable counts, and normalize offers with cost per usable record instead of headline price.
Size the evaluation set by precision, not by training scale
An evaluation set is sized by the smallest difference you need to detect on each slice you report, not as a percentage of training volume. For an accuracy-style metric, the binomial standard error is sqrt(p(1−p)/n): at 80% accuracy, 400 items give roughly ±3.9 points at 95% confidence and 1,000 items give roughly ±2.5 points. Every slice you report separately (product line, language, ticket category) needs that count on its own, which is why stratified evals grow faster than intuition suggests.
Below a few hundred items, the normal-approximation interval is unreliable; an ICML 2025 position paper argues CLT-based error bars on small LLM benchmarks tend to come out too narrow and recommends alternatives [2]. Use Wilson or Bayesian intervals there, and compare models with paired tests on the same items. The eval-set size and statistical power guide works through the power calculation.
Two procurement consequences follow. License the eval slice with terms that keep it out of every training run, and buy it from a different time window or business unit than the training tranche so leakage does not inflate scores; private eval sets versus public benchmarks explains why held-out data matters.
Structure the purchase as tranches with an option
Tranches let the learning curve, not the opening estimate, decide total spend. A typical structure is a paid pilot, a first committed tranche sized to the extrapolated knee, and an option for further tranches at a pre-agreed unit price and the same schema, de-identification method and acceptance criteria. Whether a given supplier will accept an option depends on the deal; it is something to negotiate, not assume.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Tranche | Volume basis | Gate to proceed | What the buyer decides |
|---|---|---|---|
| Pilot | 1,000-5,000 raw, random draw | Schema check, yield measured, curve fitted on nested slices | Whether the data moves the eval at all |
| Eval slice | Per-slice count from the precision target | Held out, disjoint time window | Frozen before any training tranche lands |
| Tranche 1 | Usable target at the curve's knee, divided by measured yield | Acceptance sample passes (n, c) | Refit the curve with real volume |
| Option tranches | Fixed increments at a pre-agreed unit price | Marginal gain per 1,000 usable above threshold | Exercise, defer or stop |
Each tranche needs a delivery acceptance test. A single sampling plan from the NIST/SEMATECH handbook, defined by sample size n and acceptance number c, gives both sides an objective rule: inspect n records and reject the delivery if more than c fail [6]. Define "fail" in the license or schedule (residual PII, broken schema, duplicate of an earlier tranche), and see acceptance criteria for licensed training data for wording. Purchase orders for each tranche can follow the pattern in license purchase orders for AI training.
Sizing checklist to hand to procurement
A one-page sizing memo lets procurement negotiate volume without second-guessing the ML team.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Task and metric: the eval metric, the per-slice precision target and the business value of one point.
- Pilot design: slice sizes, seeds, frozen hyperparameters, and the eval set's provenance.
- Curve fit: fitted a, b, c; extrapolation limit; marginal gain per 1,000 usable examples at each candidate volume.
- Yield: measured dedup, de-identification, format and quality-filter loss rates, and the resulting raw-to-usable ratio.
- Eval purchase: item count per slice, held-out window, and the no-training restriction.
- Tranche plan: committed volume, option increments, unit-price basis, and the (n, c) acceptance plan per delivery.
- Stop rule: the marginal-gain threshold below which no further option is exercised.
For running the pilot itself with a supplier, see how to run a data pilot with a supplier; to request the draw, see how to request a training data sample.
Where operational data from US companies fits
Sizing applies whether the records come from a marketplace, a custom collection or a direct partnership. SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records, documents and finance and legal workflows, and manages the licensing and ongoing purchases. Nothing is held in stock, so each request is matched to businesses that hold the described data, and a request does not guarantee a match. If your curve says you need more of a specific record type, describe the data to SourceX's buyer team.
Sizing a licensed dataset with SourceX
SourceX works with AI teams wherever they are based: you describe the data, and SourceX assesses data and licensing permissions, agrees pricing and allowed uses in a license, and manages later purchases. Every dataset is rights-reviewed, and personal details are removed or replaced before delivery with the method recorded and a sample checked. Start a buyer request at sourcex.si/buyers.
Sources
- Hestness et al., Baidu Research (arXiv:1712.00409), "Deep Learning Scaling is Predictable, Empirically" (2017). https://arxiv.org/abs/1712.00409v1
- ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
- arXiv:2307.08701, "AlpaGasus: Training A Better Alpaca with Fewer Data" (2023). https://arxiv.org/pdf/2307.08701v1
- arXiv:2409.15825, "60 Data Points are Sufficient to Fine-Tune LLMs for Question-Answering" (2024). https://arxiv.org/html/2409.15825v2
- Contra (vendor service listing), "Robot dataset acceptance audit: a verdict before you pay". https://contra.com/s/ee6vvEQU-robot-dataset-acceptance-audit-a-verdict-before-you-pay
- NIST Information Technology Laboratory, "NIST/SEMATECH e-Handbook of Statistical Methods, 6.2.3 Single sampling plans". https://itl.nist.gov/div898/handbook/pmc/section2/pmc23.htm
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.