Procurement, samples and ongoing supply
How Data Procurement Differs for Pre-training, Fine-tuning, Evaluation and RAG
Quick answer
The training stage changes what you are buying. For pre-training, buy volume, broad rights and clean deduplication. For fine-tuning, buy task fit: labeled inputs, outcomes and expert-quality targets. For evaluation, buy confidentiality and contamination control, since a leaked test set loses its value. For RAG, buy access rights: freshness, update feeds and the right to show retrieved text to users. Write a separate spec, license scope, pricing basis and acceptance test for each stage, even if one supplier serves several.
By SourceX Editorial · Updated
The four profiles below are working hypotheses from practice, not fixed rules. Treat them as a starting point and adjust them to your model, your data and your counsel's reading of the license. For the full purchasing lifecycle, start at the AI training data procurement hub. For the technical differences between the approaches, see what data fine-tuning, RAG and continued pre-training each need.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why one purchase spec fails across training stages
A single spec fails because each stage gets its value from different properties of the same data. A pre-training buyer accepts noise in exchange for tokens. An SFT buyer pays for a small number of examples whose targets are correct. An eval buyer pays mostly for the fact that nobody else has trained on the set.
Teams that reuse a pre-training contract for an eval set often find the license allows the supplier to resell the same records, which ends the set's use as a holdout. Teams that reuse a fine-tuning license for RAG often find it covers training but not showing passages verbatim to end users. Both are scope failures, not quality failures, and they show up after signature. The pattern is common enough that it is one of the causes on our list of why training data purchases fail.
Procurement profile by stage: the comparison table
The table below summarizes how each procurement lever shifts with the stage the data will serve. Use it to set up the RFP sections before you approach suppliers.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Lever | Pre-training / continued pre-training | Fine-tuning (SFT, preference) | Evaluation / benchmarks | RAG / retrieval corpus |
|---|---|---|---|---|
| What drives value | Token volume, domain coverage, low duplication | Task fit, label correctness, outcome fields | Secrecy, representativeness, stable scoring | Freshness, coverage of the questions users ask, source authority |
| Typical unit priced | Tokens, documents or GB after dedup | Accepted examples or annotated records | Items or tasks, often with refresh cycles | Documents plus update feed, sometimes per-query or per-display |
| Core rights needed | Train, and keep derived weights after the term | Train, often for a named model family | Evaluate only; no training; tight redistribution limits | Index, retrieve, display or quote to users; cache |
| Exclusivity question | Rarely worth paying for | Sometimes, for differentiated tasks | Often essential to keep the holdout clean | Usually not; access terms matter more |
| Acceptance test | Dedup rate, language ID, PII scan, format parse rate | Label audit on a sample, schema validity, inter-annotator agreement | Contamination check against your training corpora, answer-key audit | Retrieval hit rate on your query set, staleness, broken-link rate |
| Main failure mode | Paying for duplicates and boilerplate | Inconsistent or wrong targets | Leakage into training data | Display rights missing, stale content |
Pre-training: volume, rights breadth and deduplication
Pre-training procurement is a rights-breadth and deduplication problem first, and a quality problem second. At this scale you cannot review records by hand, so the contract and the pipeline checks do the work.
Price on usable tokens after deduplication, not raw files. Lee et al. found many near-duplicate examples and long repeated substrings in common corpora, and showed that deduplication reduced memorized verbatim output and train-test overlap [2]. Ask the supplier to run MinHash or exact-substring dedup against its own delivery and report the rate, then rerun it against your existing corpus before accepting.
Rights are the second exposure. The Data Provenance Initiative audited 1,800+ text datasets and reported widespread missing or wrong license information on popular hosting sites [1]. For licensed operational data, require a provenance record per source: who created it, under what agreement, and which consents cover it. If you place a general-purpose model on the EU market, Article 53 requires a copyright compliance policy and a public summary of training content [7], and the Commission's template for that summary dates from 24 July 2025 [8]. As of October 2026, the AI Office's enforcement powers have applied since 2 August 2026 for models placed on the market after 2 August 2025, so your supplier records must be detailed enough to fill the template. The provenance guide covers the record set.
Key contract terms at this stage: the right to keep model weights after the term ends, treatment of derived artifacts such as tokenizers and filtered subsets, and what happens to retained copies if a source is later withdrawn.
Fine-tuning: task fit, labels and outcome fields
Fine-tuning procurement buys correctness per example, so the acceptance test is a label audit, not a volume count. LIMA fine-tuned a 65B-parameter model on 1,000 curated prompt-response pairs, and its authors noted that this kind of curation takes a lot of work [3]. That cost belongs in the price discussion: you are paying for expert judgment, not storage.
Specify the training method, because it changes the record schema. SFT needs demonstrations: an input and a target output, as in the labeler-written demonstrations used for InstructGPT [4]. RLHF needs ranked comparisons of model outputs [4]. DPO needs preference pairs with a chosen and a rejected response [5]. A supplier holding support tickets with resolution codes, or claims with adjudication outcomes, can often supply outcome fields that work as weak labels. Confirm that those fields were recorded at the time, not reconstructed later.
Accept on a stratified sample: check schema validity, have a domain reviewer score target correctness, and measure inter-annotator agreement where labels were added. ISO/IEC 5259-4 sets out a data quality process framework covering labelling for supervised ML, which helps in defining who checks what [9]. For the detailed pre-purchase check, see how to evaluate a fine-tuning dataset before you buy it.
Evaluation: confidentiality, exclusivity and contamination control
Evaluation data is worth paying for only while it stays out of everyone's training data, so the license and handling terms are as important as the items. In February 2026 OpenAI said it no longer evaluates SWE-bench Verified, citing contamination concerns [6]. A private set that leaks loses its value the same way.
Write the eval license narrowly. Grant evaluation and error analysis, prohibit training and fine-tuning, including by contractors, and restrict who can see answer keys. Ask whether the supplier has sold, or will sell, the same items to anyone else. If they have, negotiate exclusivity or a fresh holdout. Before acceptance, run n-gram or embedding overlap checks between the eval items and your pre-training and SFT corpora. Also ask for the collection date, so you can tell whether the items predate the cutoffs of models you will compare against.
Eval sets also need documentation that training sets can skip: scoring rubric, expected answer format, known ambiguous items and the intended use. Data Cards give a usable structure for this [10]. The difference between the two purchases is explained in training data vs evaluation data, and SourceX describes eval use cases at evaluation datasets from real business work.
RAG: access terms, freshness and display rights
RAG procurement buys the right to retrieve and show content at query time, which most training licenses do not grant. At least one copyright licensing body, Australia's Copyright Agency, lists RAG as an AI licensing activity distinct from fine-tuning and model development [11]. Ask for each right explicitly.
The rights to list are: index and embed; store a cache; quote or display passages to end users; attribute or link to the source; and continue serving cached content after the term ends or after a record is withdrawn. Because the model weights do not change, the supplier can revoke access more easily than in a training deal. Agree the takedown process, the notice period and how deletions propagate to your vector store.
Freshness drives pricing and acceptance. Specify the update cadence, the change feed format (full re-export, incremental diff or event stream) and what metadata each document carries: stable ID, version, effective date and access-control labels. Then test retrieval quality on your own query set. For document-heavy RAG use cases, see RAG evaluation datasets from real company documents.
One request, four specs: a worked example
A single source can feed several stages if you split the purchase into separate scopes. Suppose your team wants a support-operations corpus from a US software company.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Scope | Records | Rights requested | Acceptance test |
|---|---|---|---|
| Continued pre-training | Knowledge base articles and resolved ticket threads, deduplicated | Train; retain weights after term | Duplicate rate below an agreed threshold; PII scan on a sample |
| SFT | Ticket and resolution pairs with category and resolution code | Train for a named model family | Reviewer agreement on a 300-record sample |
| Evaluation | Held-back tickets from a later date range, with gold resolutions | Evaluate only; no training; supplier withholds from other buyers | Overlap check against the two scopes above |
| RAG | Current knowledge base with weekly incremental updates | Index, cache and display excerpts with attribution | Hit rate on 200 internal queries; staleness under the agreed window |
Splitting by date keeps the eval scope cleaner than a random split, because later tickets cannot appear in earlier training scopes. Price each scope on its own basis, then compare quotes with the method in comparing data vendor quotes and record pass/fail tests in your acceptance criteria.
How SourceX handles stage-specific requests
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Buyers describe the data they need, not the businesses; SourceX looks for US businesses that hold it, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked; no method is perfect.
Data is not held in stock, and a request does not guarantee a match. You can describe a stage-specific data need to SourceX at any point in your planning.
Buying data for a specific training stage
If you know which stage the data will serve, write it into your request so rights and acceptance tests can be scoped from the start. SourceX works through Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Describe the data for your training stage.
Sources
- Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Lee et al. (arXiv), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
- Zhou et al. (arXiv; NeurIPS 2023), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- Ouyang et al., OpenAI (arXiv), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
- Rafailov et al. (arXiv; NeurIPS 2023), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality for analytics and machine learning, Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
- Pushkarna, Zaldivar, Kjartansson, Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- Copyright Agency (Australia), "AI Licensing Glossary". https://www.copyright.com.au/?p=29938
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.