Skip to content

Procurement, samples and ongoing supply

Estimating a Dataset's Value to Your Model Before You Buy

Quick answer

To evaluate a dataset before buying it, measure what a sample does to your own model rather than judging the data in the abstract. Train or index with the sample against a frozen baseline and an equal-size control at matched compute, and score every arm on a private held-out set the vendor never saw. Repeat at several sample sizes to fit a learning curve, remove contamination, and buy only if the forecast lift at the purchase volume clears a threshold you set before testing.

By SourceX Editorial · Updated

Five kinds of evidence behind a value estimate

A dataset's value to you is the change it produces in a metric you care about, at the volume you would buy, compared with the next-best data you could get. The five kinds of evidence below each estimate one part of that number, at very different cost.

The buyer usually has to generate this evidence: one 2023 analysis found that data marketplaces lack detailed dataset information, transparent pricing and standardized formats [1]. Practitioner guidance on AI procurement makes a similar point: require a proof of concept on your own data rather than relying on vendor presentations [2].

EvidenceQuestion it answersTypical costWhat it cannot tell you
Descriptive profiling (schema, duplicates, label audit, slice counts)Is the data clean, and does it cover your slices?HoursWhether your model improves
Eval-gap analysisHow much headroom sits where the data concentrates?About a day with an existing eval setRealized lift
Matched ablation on the sampleDoes the sample beat equal-size alternative data?A few training or indexing runs per armWhat the full volume will do
Learning curve over nested subsetsHow fast gains shrink, and the forecast at purchase volume4 or 5 sizes, several seeds eachEffects of records the sample lacks
Record-level valuationWhich records or strata carry the gain, and which hurtHighest; often needs a smaller proxy modelA purchase-level forecast on its own

This page covers measured utility. For descriptive quality and coverage checks, see the pre-purchase quality checklist in the training data quality hub; for commercial drivers such as scarcity, rights and exclusivity, SourceX's overviews of how AI companies value datasets and data valuation approaches; for pilot scope and timeline, how to run a data pilot with a supplier.

Eval-gap analysis: size the headroom before training anything

Before any training run, put a ceiling on value. For each slice of your evaluation set, multiply its share of production traffic by the gap between current and target performance, then check how much of the sample falls in each slice. Data concentrated where your model already succeeds cannot add much, however clean it is.

Tag eval items and sample records with one failure taxonomy (escalations, refund policy, billing disputes, non-English requests), labeling the sample with a classifier or embedding nearest-neighbor search. Targeting data at observed failures has research support: in image classification, Singla et al. started from a small set of failure cases, retrieved similar examples from a large pool, and reported a 9.45% improvement on held-out debug sets over baselines that preserved test-set performance [3].

Illustrative example: invented to show structure; it does not describe an available dataset.

Eval sliceTraffic shareCurrent successTargetWeighted headroom (points)Share of vendor sample
Escalations to engineering18%35%60%4.522%
Refund and credit policy12%50%70%2.415%
Billing disputes30%65%75%3.035%
How-to questions40%76%80%1.628%

Overall success here is 62.2%, and total weighted headroom is 11.5 points, 6.9 of them in escalations and refunds, which hold 37% of the sample. A pitch implying more gain than the headroom allows is a red flag before any test runs.

Define lift for the way you will use the data

Lift means different things for fine-tuning, continued pretraining, retrieval and evaluation data. Before the sample arrives, fix the metric, the control arm and the regression check for your use.

Use of the dataValue metric on your private evalControl armTrap that inflates value
Supervised fine-tuning (SFT)Task success or rubric score per sliceEqual supervised tokens of your own or open dataFormat and tone gains saturate early: LIMA changed a 65B model's response style with 1,000 curated examples [4]
Continued pretraining for domain adaptationHeld-out in-domain loss, then downstream task scoresEqual tokens of generic in-domain textLoss falls while downstream scores stay flat, and general skills regress
Retrieval (RAG) corpusRecall@k and grounded-answer accuracy on questions the corpus should answerCurrent index plus an equal number of alternative documentsTest questions written from the sample itself
Evaluation dataSeparation between checkpoints you know differ, and agreement with production outcomesYour existing eval setItems already present in public training data

Protocol detail sits with each data type: the SFT and preference-data ablation protocol, retrieval lift pilots, WER lift for speech data, code dataset samples and agent trajectory samples. In every case, hold the base checkpoint, hyperparameters, training budget and seeds fixed across arms. Score on a private set built from your own records, as in private evaluation sets vs public benchmarks.

Do not let the supplier write or hold that set. Bansal and Maini argue that a curator who sells training data to model developers and also evaluates their models has a conflict of interest, and that its annotators' preferences can tilt results toward models trained on its data [5]. Training test models on a sample also needs explicit permission; see evaluation licenses and NDAs for dataset samples.

From sample to purchase volume: fit a learning curve

A single ablation gives the gain at sample size; a learning curve shows how fast that gain shrinks with more of the same data, which you need to price a purchase 10 or 20 times larger than the sample. Train the candidate arm on nested random subsets (for example 1/8, 1/4, 1/2 and all of the sample), with at least three seeds per size, and plot error against record count on log-log axes.

Hestness et al. found, across machine translation, language modeling, image processing and speech recognition, that generalization error falls as a power law in training-set size [6]. Theory does not yet explain the exponent, so it has to be measured, and model improvements mainly shift the curve rather than change its slope [6]. A common fitting form is error(n) = c + a·n^(−b), where c is the floor that more of this data cannot push below.

  • The floor dominates the forecast. Four or five points pin it down poorly, so report forecasts across plausible floors.
  • Uncertainty grows with distance. A forecast far beyond your largest subset is a hypothesis for a first tranche to confirm.
  • Differences between sizes are small. Score every subset on the same items with paired comparisons. CLT-based intervals are too narrow below a few hundred datapoints, an ICML 2025 position paper argues [7], which matters for per-slice reads; see eval set sizing for statistical power.

Convert the curve into an exchange rate

Fit the same curve for the control arm and read off how many control records reach the gain the vendor sample delivers. Hernandez et al. defined "effective data transferred" as the extra fine-tuning data a same-size model trained from scratch would need to match a pre-trained model's loss, and found it follows a power law in model size and fine-tuning dataset size in the low-data regime [8].

For procurement, vendor records × exchange rate × your fully loaded cost per in-house record is a ceiling on the data's worth relative to producing it yourself. If you could not produce comparable data at all, set the ceiling from the business value of each metric point. Tranche sizing continues in how much training data to buy, and the approver-facing case in the business case for buying training data.

Worked example: forecasting lift for a 120,000-record purchase

A worked forecast shows how a sample result becomes a purchase decision with an honest range.

Illustrative example: invented to show structure; it does not describe an available dataset.

The team behind the eval-gap table fine-tunes an 8B instruct model to resolve B2B software support tickets; baseline error on a 1,500-item private eval is 37.8%. A supplier offers 120,000 resolved ticket threads and provides an 8,000-record random sample under terms that allow test fine-tunes. LoRA settings are fixed on the baseline, each size runs three seeds (spread about ±0.4 points), and arms are matched on supervised tokens.

Records addedVendor arm errorGain vs baselineOwn-records control gain
1,00035.7%+2.1+1.0
2,00034.4%+3.4+1.8
4,00033.2%+4.6+2.7
8,00032.1%+5.7+3.5
  1. Fit. A free fit puts the floor near 19.5% with b ≈ 0.12 and forecasts about +9.2 points at 120,000 records. Floors of 25%, 28% and 30% also fit all four points within seed spread and forecast +8.6, +8.0 and +7.2. The defensible forecast is +7 to +9 points, from a 15-fold extrapolation.
  2. Headroom check. The eval-gap table caps the overall gain near 11.5 points. Billing and how-to questions hold only 4.6 of those points, so at least 2.6 to 4.6 points of the forecast must come from escalations and refunds; slice scores at 8,000 records should show that.
  3. Exchange rate. The control gains about 0.84 points per doubling and, extended on that line, reaches +5.7 only at about 49,000 records. At sample scale, one vendor record is worth roughly six of the team's own for this metric.
  4. Decision. The pre-set rule was a lower forecast of at least +5 points at purchase volume, and a vendor lead over the control of at least 2 points at matched size. Both hold (+7.2; +5.7 vs +3.5). Because of the extrapolation distance, the team licenses a 30,000-record first tranche, forecast at +6.6 to +7.6, with the rest conditional on the measured gain landing in that band.

Payment terms for this structure are in paying for data in milestones tied to acceptance.

Record-level valuation: find the strata that carry the gain

Record-level methods rank records or strata by their contribution to your metric. They decide which slice to license and which records to exclude, but none forecasts full-purchase lift on its own. Each score is relative to one model, metric and set of other records, so it may not carry over to a new base model or eval set.

MethodQuestion it answersComputeProcurement use
Stratum leave-out ablationWhat is lost if one source system, year or record type is dropped?One retrain per stratumChoose and price strata directly
Data Shapley [9]Each record's contribution to performance, averaged over subsetsMany retrains; Monte Carlo and gradient approximationsRank records on a small proxy model
Influence functions [10]Which training points most affect a given test predictionGradients and Hessian-vector products; theory not exact for non-convex modelsTrace a regression back to the supplier records behind it
TracIn [11]How a test loss changed each time a training example was usedSaved checkpoints and first-order gradientsRun on your ablation checkpoints to flag records that lower or raise test loss
Datamodels [12]What the model would output if trained on a given subsetFitted to outcomes of many training runs on subsetsPredict the effect of alternative subset purchases

At language-model scale, the practical pair is stratum leave-out ablations plus TracIn-style tracing on checkpoints you already trained. A finding such as "escalation threads carry most of the gain; pre-2021 records add nothing" is a concrete scope to negotiate.

Rule out leakage and data you already own

Lift from a sample counts only after you remove records that overlap your evaluation set or duplicate data you already train on. Eval overlap inflates measured value; duplication sells you data you already have.

  • Overlap with the eval set. Run exact hashes, then n-gram overlap, which a 2024 survey calls the predominant decontamination technique; thresholds vary, and GPT-3's work used 13-gram matches [13]. Add embedding similarity and human review, because paraphrased or translated test items pass n-gram checks and can inflate scores [14].
  • Overlap with your own training mix. Exact and near-duplicates are common: Lee et al. found one sentence repeated over 60,000 times in C4 [15]. Deduplicate the sample against your corpus with MinHash LSH; what survives is the novel fraction you would pay for.
  • Public benchmarks as value signals. Treat a vendor's public-benchmark gains as weak evidence. In a post dated 23 February 2026, OpenAI said it stopped reporting SWE-bench Verified because gains increasingly reflected training-time exposure [16].
  • Adaptive overfitting. Every change made after reading scores fits the eval set a little more. Keep a final holdout scored once, as in contamination-resistant evaluation design and SourceX's guide to contamination checks for licensed evaluation data.

The value-estimate record approvers should see

A go/no-go is defensible when one record ties the forecast to its evidence and to a threshold set in advance. The broader AI training data procurement lifecycle shows where this record feeds the license and acceptance terms.

Illustrative example: invented to show structure; it does not describe an available dataset.

value_estimate:
  candidate: "resolved B2B support ticket threads, 120,000 records offered"
  owner: ml_research_lead
  eval_set: {version: support-eval-v7, items: 1500, vendor_access: none, final_holdout: 300}
  baseline: {checkpoint: internal-8b-instruct-v3, error_pct: 37.8}
  sample: {records: 8000, draw: simple_random, seed_logged: true, prep_matches_delivery: true}
  arms: {vendor_nested: [1000, 2000, 4000, 8000], own_control_nested: [1000, 2000, 4000, 8000]}
  controls: {seeds_per_size: 3, matched_on: supervised_tokens, hyperparameters_fixed_on: baseline}
  contamination: {eval_overlap_checks: [exact_hash, ngram, embedding_review], dedup_vs_own_mix: minhash}
  curve: {form: "error = c + a*n^-b", floor_range_pct: [19.5, 30.0]}
  forecast_gain_points: {at_30000: [6.6, 7.6], at_120000: [7.2, 9.2]}
  exchange_rate_vs_own_records: "about 6x at 8,000 records"
  pre_set_rule: "lower forecast >= +5.0 at purchase volume; vendor minus control >= +2.0 at 8,000"
  decision: "license 30,000-record tranche; remainder conditional on measured gain in forecast band"

If you source operational records through SourceX, your request can name the weak slices from your gap analysis and the sample size your learning curve needs. SourceX does not publish prices; terms depend on scope, volume, history, rights and exclusivity, so a stratum-level value estimate is useful input when you submit a data request on the SourceX buyer page.

Where pre-purchase estimates overstate value

Most inflated estimates come from a test that differs from the purchase in a way nobody recorded. Check these before sign-off:

  • A curated sample. A "best of" draw inflates the candidate arm and the curve. Require a random or stratified draw with a logged seed, as in requesting a training data sample, and compare it with population statistics using checks for a representative sample.
  • Time mismatch. A latest-year sample says little about older records that dominate the delivery.
  • Compute confound. If the candidate arm trains for more steps because it has more tokens, part of its gain is compute.
  • A changing base model. The curve belongs to one checkpoint and recipe; model improvements shifted the curves in Hestness et al.'s experiments [6], so re-estimate before a renewal or base-model upgrade.
  • Judge bias. An LLM judge that favors longer answers turns style into apparent value; report response length beside scores.

Need operational data you can test against your own evals?

On the SourceX buyer page you can submit a data request that describes the records you need and the slices your evaluation shows are weak, or talk to SourceX first. SourceX looks for US companies that hold matching data, checks the data and the supplier's licensing permissions, and manages the license and delivery. Nothing is contracted until a supplier agrees, and a request does not guarantee a matching dataset. Describe the data you want to test.

Sources

  1. Chen et al., "Data Acquisition: A New Frontier in Data-centric AI" (2023). https://arxiv.org/abs/2311.13712
  2. Amit Kothari, "AI RFP template that tests capability" (practitioner blog; market practice). https://amitkoth.com/ai-rfp-template/
  3. Singla et al., University of Maryland, "Data-Centric Debugging: mitigating model failures via targeted data collection" (2022; WACV 2024 version titled "Mitigating Model Failures via Targeted Image Retrieval"). https://arxiv.org/pdf/2211.09859
  4. Zhou et al., "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  5. Bansal and Maini, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
  6. Hestness et al., Baidu Research, "Deep Learning Scaling is Predictable, Empirically" (2017). https://arxiv.org/abs/1712.00409v1
  7. ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  8. Hernandez, Kaplan, Henighan, McCandlish, "Scaling Laws for Transfer" (2021). https://arxiv.org/abs/2102.01293v1
  9. Ghorbani and Zou, "Data Shapley: Equitable Valuation of Data for Machine Learning" (ICML 2019). https://proceedings.mlr.press/v97/ghorbani19c.html
  10. Koh and Liang, "Understanding Black-box Predictions via Influence Functions" (ICML 2017). https://arxiv.org/pdf/1703.04730
  11. Pruthi, Liu, Kale, Sundararajan, "Estimating Training Data Influence by Tracing Gradient Descent" (NeurIPS 2020). https://arxiv.org/pdf/2002.08484
  12. Ilyas, Park, Engstrom, Leclerc, Madry, "Datamodels: Understanding Predictions with Data and Data with Predictions" (ICML 2022). https://proceedings.mlr.press/v162/ilyas22a.html
  13. arXiv preprint 2406.04244, "Benchmark Data Contamination of Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2406.04244
  14. arXiv preprint 2311.04850, "Rethinking Benchmark and Contamination for Language Models with Rephrased Samples" (2023). https://arxiv.org/pdf/2311.04850v1
  15. Lee, Ippolito, Nystrom, Zhang, Eck, Callison-Burch, Carlini, "Deduplicating Training Data Makes Language Models Better" (2021; ACL 2022). https://arxiv.org/abs/2107.06499v1
  16. OpenAI, "Why SWE-bench Verified no longer measures frontier coding capabilities" (23 February 2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data