Skip to content

Procurement, samples and ongoing supply

How Data Procurement Differs for Pre-training, Fine-tuning, Evaluation and RAG

Quick answer

The training stage changes what you are buying. For pre-training, buy volume, broad rights and clean deduplication. For fine-tuning, buy task fit: labeled inputs, outcomes and expert-quality targets. For evaluation, buy confidentiality and contamination control, since a leaked test set loses its value. For RAG, buy access rights: freshness, update feeds and the right to show retrieved text to users. Write a separate spec, license scope, pricing basis and acceptance test for each stage, even if one supplier serves several.

By SourceX Editorial · Updated

The four profiles below are working hypotheses from practice, not fixed rules. Treat them as a starting point and adjust them to your model, your data and your counsel's reading of the license. For the full purchasing lifecycle, start at the AI training data procurement hub. For the technical differences between the approaches, see what data fine-tuning, RAG and continued pre-training each need.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why one purchase spec fails across training stages

A single spec fails because each stage gets its value from different properties of the same data. A pre-training buyer accepts noise in exchange for tokens. An SFT buyer pays for a small number of examples whose targets are correct. An eval buyer pays mostly for the fact that nobody else has trained on the set.

Teams that reuse a pre-training contract for an eval set often find the license allows the supplier to resell the same records, which ends the set's use as a holdout. Teams that reuse a fine-tuning license for RAG often find it covers training but not showing passages verbatim to end users. Both are scope failures, not quality failures, and they show up after signature. The pattern is common enough that it is one of the causes on our list of why training data purchases fail.

Procurement profile by stage: the comparison table

The table below summarizes how each procurement lever shifts with the stage the data will serve. Use it to set up the RFP sections before you approach suppliers.

Illustrative example: invented to show structure; it does not describe an available dataset.

LeverPre-training / continued pre-trainingFine-tuning (SFT, preference)Evaluation / benchmarksRAG / retrieval corpus
What drives valueToken volume, domain coverage, low duplicationTask fit, label correctness, outcome fieldsSecrecy, representativeness, stable scoringFreshness, coverage of the questions users ask, source authority
Typical unit pricedTokens, documents or GB after dedupAccepted examples or annotated recordsItems or tasks, often with refresh cyclesDocuments plus update feed, sometimes per-query or per-display
Core rights neededTrain, and keep derived weights after the termTrain, often for a named model familyEvaluate only; no training; tight redistribution limitsIndex, retrieve, display or quote to users; cache
Exclusivity questionRarely worth paying forSometimes, for differentiated tasksOften essential to keep the holdout cleanUsually not; access terms matter more
Acceptance testDedup rate, language ID, PII scan, format parse rateLabel audit on a sample, schema validity, inter-annotator agreementContamination check against your training corpora, answer-key auditRetrieval hit rate on your query set, staleness, broken-link rate
Main failure modePaying for duplicates and boilerplateInconsistent or wrong targetsLeakage into training dataDisplay rights missing, stale content

Pre-training: volume, rights breadth and deduplication

Pre-training procurement is a rights-breadth and deduplication problem first, and a quality problem second. At this scale you cannot review records by hand, so the contract and the pipeline checks do the work.

Price on usable tokens after deduplication, not raw files. Lee et al. found many near-duplicate examples and long repeated substrings in common corpora, and showed that deduplication reduced memorized verbatim output and train-test overlap [2]. Ask the supplier to run MinHash or exact-substring dedup against its own delivery and report the rate, then rerun it against your existing corpus before accepting.

Rights are the second exposure. The Data Provenance Initiative audited 1,800+ text datasets and reported widespread missing or wrong license information on popular hosting sites [1]. For licensed operational data, require a provenance record per source: who created it, under what agreement, and which consents cover it. If you place a general-purpose model on the EU market, Article 53 requires a copyright compliance policy and a public summary of training content [7], and the Commission's template for that summary dates from 24 July 2025 [8]. As of October 2026, the AI Office's enforcement powers have applied since 2 August 2026 for models placed on the market after 2 August 2025, so your supplier records must be detailed enough to fill the template. The provenance guide covers the record set.

Key contract terms at this stage: the right to keep model weights after the term ends, treatment of derived artifacts such as tokenizers and filtered subsets, and what happens to retained copies if a source is later withdrawn.

Fine-tuning: task fit, labels and outcome fields

Fine-tuning procurement buys correctness per example, so the acceptance test is a label audit, not a volume count. LIMA fine-tuned a 65B-parameter model on 1,000 curated prompt-response pairs, and its authors noted that this kind of curation takes a lot of work [3]. That cost belongs in the price discussion: you are paying for expert judgment, not storage.

Specify the training method, because it changes the record schema. SFT needs demonstrations: an input and a target output, as in the labeler-written demonstrations used for InstructGPT [4]. RLHF needs ranked comparisons of model outputs [4]. DPO needs preference pairs with a chosen and a rejected response [5]. A supplier holding support tickets with resolution codes, or claims with adjudication outcomes, can often supply outcome fields that work as weak labels. Confirm that those fields were recorded at the time, not reconstructed later.

Accept on a stratified sample: check schema validity, have a domain reviewer score target correctness, and measure inter-annotator agreement where labels were added. ISO/IEC 5259-4 sets out a data quality process framework covering labelling for supervised ML, which helps in defining who checks what [9]. For the detailed pre-purchase check, see how to evaluate a fine-tuning dataset before you buy it.

Evaluation: confidentiality, exclusivity and contamination control

Evaluation data is worth paying for only while it stays out of everyone's training data, so the license and handling terms are as important as the items. In February 2026 OpenAI said it no longer evaluates SWE-bench Verified, citing contamination concerns [6]. A private set that leaks loses its value the same way.

Write the eval license narrowly. Grant evaluation and error analysis, prohibit training and fine-tuning, including by contractors, and restrict who can see answer keys. Ask whether the supplier has sold, or will sell, the same items to anyone else. If they have, negotiate exclusivity or a fresh holdout. Before acceptance, run n-gram or embedding overlap checks between the eval items and your pre-training and SFT corpora. Also ask for the collection date, so you can tell whether the items predate the cutoffs of models you will compare against.

Eval sets also need documentation that training sets can skip: scoring rubric, expected answer format, known ambiguous items and the intended use. Data Cards give a usable structure for this [10]. The difference between the two purchases is explained in training data vs evaluation data, and SourceX describes eval use cases at evaluation datasets from real business work.

RAG: access terms, freshness and display rights

RAG procurement buys the right to retrieve and show content at query time, which most training licenses do not grant. At least one copyright licensing body, Australia's Copyright Agency, lists RAG as an AI licensing activity distinct from fine-tuning and model development [11]. Ask for each right explicitly.

The rights to list are: index and embed; store a cache; quote or display passages to end users; attribute or link to the source; and continue serving cached content after the term ends or after a record is withdrawn. Because the model weights do not change, the supplier can revoke access more easily than in a training deal. Agree the takedown process, the notice period and how deletions propagate to your vector store.

Freshness drives pricing and acceptance. Specify the update cadence, the change feed format (full re-export, incremental diff or event stream) and what metadata each document carries: stable ID, version, effective date and access-control labels. Then test retrieval quality on your own query set. For document-heavy RAG use cases, see RAG evaluation datasets from real company documents.

One request, four specs: a worked example

A single source can feed several stages if you split the purchase into separate scopes. Suppose your team wants a support-operations corpus from a US software company.

Illustrative example: invented to show structure; it does not describe an available dataset.

ScopeRecordsRights requestedAcceptance test
Continued pre-trainingKnowledge base articles and resolved ticket threads, deduplicatedTrain; retain weights after termDuplicate rate below an agreed threshold; PII scan on a sample
SFTTicket and resolution pairs with category and resolution codeTrain for a named model familyReviewer agreement on a 300-record sample
EvaluationHeld-back tickets from a later date range, with gold resolutionsEvaluate only; no training; supplier withholds from other buyersOverlap check against the two scopes above
RAGCurrent knowledge base with weekly incremental updatesIndex, cache and display excerpts with attributionHit rate on 200 internal queries; staleness under the agreed window

Splitting by date keeps the eval scope cleaner than a random split, because later tickets cannot appear in earlier training scopes. Price each scope on its own basis, then compare quotes with the method in comparing data vendor quotes and record pass/fail tests in your acceptance criteria.

How SourceX handles stage-specific requests

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Buyers describe the data they need, not the businesses; SourceX looks for US businesses that hold it, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked; no method is perfect.

Data is not held in stock, and a request does not guarantee a match. You can describe a stage-specific data need to SourceX at any point in your planning.

Buying data for a specific training stage

If you know which stage the data will serve, write it into your request so rights and acceptance tests can be scoped from the start. SourceX works through Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Describe the data for your training stage.

Sources

  1. Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  2. Lee et al. (arXiv), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
  3. Zhou et al. (arXiv; NeurIPS 2023), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  4. Ouyang et al., OpenAI (arXiv), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  5. Rafailov et al. (arXiv; NeurIPS 2023), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  6. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  7. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  8. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  9. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality for analytics and machine learning, Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
  10. Pushkarna, Zaldivar, Kjartansson, Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  11. Copyright Agency (Australia), "AI Licensing Glossary". https://www.copyright.com.au/?p=29938

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data