Procurement, samples and ongoing supply
Comparing training data vendor quotes: normalizing price to cost per usable record
Quick answer
To compare training data vendor quotes, convert each one, whether priced per record, per hour, per token or as a flat fee, into cost per usable unit: the license fee plus vendor-dependent costs on your side, divided by the units that pass your deduplication, scope, de-identification and label checks on a random sample from that vendor. Rank quotes only at matched rights scope and term, and have vendors requote rather than adjusting for scope by guesswork.
By SourceX Editorial · Updated
Why quotes for the same data request differ
Quotes for one request differ for five reasons: the billing unit, what each vendor counts as a unit, how much of the delivery survives your checks, the rights granted, and the preparation work included. Conversion factors, a fixed counting rule and sample-based yield put the first three on one basis before you negotiate; differences in rights and included preparation need a common scope and added costs instead.
A language-services vendor, comparing in-house and outsourced annotation, argues that hourly rates rarely decide the real cost; ramp-up time, language coverage and QA at scale do [1].
Price drivers such as history depth, outcome coverage and exclusivity are covered in SourceX's guide to what drives the price of licensed enterprise data, deal structures in how AI data deals are priced, and the choice between flat-fee, per-record and per-token licenses in AI data license pricing structures compared. This page is the side-by-side method within the AI training data procurement guide.
Define the usable unit before you open the price sheets
A usable unit is the record, token or hour that passes every check you will run before training or evaluation, written down with each check's measurement method. Fixing it first stops each quote from setting its own denominator.
ISO/IEC 5259-2 sets out data quality measures [2]; define each one by the property and the method used to quantify it. Apply the same discipline here, so "duplicate" means a named algorithm and threshold, not a reviewer's impression.
| Use | Usable unit | Exclusions to state | Measurement |
|---|---|---|---|
| Supervised fine-tuning (SFT) | Prompt-response pair or full conversation | Exact and near-duplicates; wrong language; truncated responses; failed schema; PII left after redaction; below quality cutoff | Dedup threshold, language ID, JSON Schema validator, PII scanner, validated grader cutoff |
| Evaluation set | Item with a verified reference answer or label | Ambiguous items; label errors found in adjudication; overlap with training data or public benchmarks | Double review with adjudication; overlap check against your corpora |
| Pre-training or continued pre-training | Token counted by your tokenizer after filtering | Boilerplate, templates, near-duplicates, off-domain or low-quality text | Your tokenizer and version; your filter pipeline |
| Speech | Speech hour after silence and hold removal, or one call | Hold music, IVR prompts, voicemail, transcripts above your word error rate (WER) threshold | Voice-activity detection; WER on a re-transcribed subsample |
| Operational records (tickets, claims, cases) | Case with required fields and an outcome | Missing outcome fields, system-generated records, dates outside range | Field-completeness and range validators |
See near-duplicate detection with MinHash and LSH and quality filtering for instruction-tuning data for the checks themselves.
Convert per-hour, per-token, per-record and flat-fee prices with measured factors
Every pricing unit converts to your usable unit through a factor, and that factor must come from the vendor's own sample, not a typical value. A per-hour price becomes a per-call price only once you know the average call length in that vendor's population.
| Quoted unit | Factor to measure on the sample | Where it goes wrong |
|---|---|---|
| Per record | Records per usable unit (messages per thread, lines per claim, turns per conversation) | "Record" means a message to one vendor and a thread to another |
| Per hour of audio or video | Recorded minutes per record; speech minutes per recorded minute | File duration includes hold and silence; dual-channel files billed twice |
| Per annotation hour | Items completed per hour at your QA bar, including review passes | A low rate with low throughput or no second review |
| Per token | Vendor tokens per record and your tokens per record | Undisclosed tokenizer; markup, speaker tags and metadata billed as content |
| Flat fee | Record count in the defined population | "All available data" with no count, date range or filter |
Token counts are the least portable unit. Tokens per word, called fertility, depend on the tokenizer's vocabulary and the language: one study reduced fertility by 42% for Hungarian and 73% for Thai by replacing about 10% of a model's vocabulary [3]. Count the sample with your own tokenizer and version, and ask vendors to bill on that count; see per-token training data pricing and estimating token counts before licensing. For audio, state whether a billable hour is file, speech or transcribed duration; how speech data is priced covers the add-ons.
Measure each vendor's yield on a random sample
Yield is the share of billed units that become usable units, and a yield gap alone can reverse a price ranking. Measure it on a random sample drawn from the exact population each vendor quoted, using one pipeline for every vendor.
Typical losses:
- Duplicates. Lee et al. found that standard language-modeling corpora contain many near-duplicates and long repeated substrings, including one sentence repeated more than 60,000 times in C4 [4]. Business exports add overlapping re-exports, macros and auto-replies.
- Label errors. An audit of test sets from 10 widely used datasets estimated an average label error rate of at least 3.3% [5]. Budget re-labeling or rejection for annotated items.
- Quality filtering. In controlled pretraining experiments, quality filters improved downstream performance even though they removed 10% or more of the training data [6]. For instruction data, AlpaGasus kept about 9,000 of the 52,000 Alpaca examples after LLM-graded filtering and reported a better model [7].
- Out-of-scope and emptied records. Wrong language, dates outside the requested range, and records left empty after PII removal.
A robot-dataset audit provider claims that quality filtering typically discards 20 to 30% of a collected dataset [8]; that figure is self-reported and comes from robotics data, so measure your own. With 500 sampled records and an observed yield of 80%, the 95% interval is roughly ±3.5 points (1.96 × √(0.8 × 0.2 ÷ 500)), so convert that band into a cost range for each vendor and treat quotes whose ranges overlap as ties.
How many records to check covers sample sizing, and requesting a training data sample covers getting a random draw rather than a showcase.
Normalize rights scope before you compare price
A quote is comparable only with quotes that grant the same uses, models, term and exclusivity. Do not discount a broad license by a guessed percentage; ask every vendor to price one common base scope plus the same list of priced options.
| Dimension | Base scope to request (example) | Options to price separately |
|---|---|---|
| Permitted use | Fine-tuning and internal evaluation | Pre-training; synthetic data generation |
| Models | One model family | Successor models or all company models (per-model vs enterprise-wide licenses) |
| Term and survival | 3 years; trained models survive termination | Perpetual term (model retention after termination) |
| Exclusivity | Non-exclusive | Time-boxed or field-limited exclusivity (negotiating exclusivity) |
| Users | Licensee and named affiliates | Customer fine-tuning; sublicensing |
| Refresh | None | Periodic refresh at a fixed price per delivery |
An evaluation-only quote buys a different right, not a cheaper training license, so give it its own column (see evaluation-only data license terms). Confirm that each vendor holds the rights it prices: an audit of more than 1,800 text datasets found license omission rates above 70% and error rates above 50% on popular hosting sites [9].
SourceX prepares diligence materials on source, rights, preparation and allowed use per dataset for the buyer's review. It does not publish prices; terms depend on scope, volume, history, rights and exclusivity, so state your base scope and options when you describe the data you need to SourceX.
Add only the costs that change with the vendor
After price and scope are normalized, add the costs on your side that differ between vendors; costs identical for every vendor do not change the ranking and belong in the total cost of ownership model. Price these per vendor:
- Missing preparation. Labels, transcripts, redaction or de-identification left to you, at your internal or third-party rate.
- Schema and format work. Engineering to map a vendor's per-turn JSON or proprietary export into your schema; zero if it delivers to your delivery specification.
- Rework. Re-labeling the error share you measured, unless the vendor replaces failed units at its own cost.
- Transfer. With an Amazon S3 Requester Pays bucket, the requester pays for requests and downloads while the bucket owner pays for storage [10]; see who pays for data egress.
- Review and payment terms. Extra legal and security review for non-standard paper; prepayment versus milestone payments tied to acceptance; minimum commitments and refresh escalators.
Worked example: four quotes for contact-center call recordings
Illustrative example: invented to show structure; it does not describe an available dataset. Figures are in a generic currency and are not market prices.
A buyer wants English support-call recordings with transcripts plus intent and outcome labels, to fine-tune and evaluate a voice support agent. A usable call is unique, has at least 60 seconds of speech, has a transcript that passes a WER spot check, has names and account numbers redacted in audio and text, and carries labels that pass audit. Each vendor supplied a random 500-call sample from the population it quoted.
| Vendor A | Vendor B | Vendor C | Vendor D | |
|---|---|---|---|---|
| Quoted price | 40 per recorded hour | 5.00 per delivered call | 3.60 per 1,000 transcript tokens, vendor tokenizer | 520,000 flat fee |
| Quoted volume | 15,000 recorded hours | 110,000 calls | 175 million tokens | All 2023-2025 calls, about 140,000 |
| Factor from sample | 7.5 recorded minutes per call | None needed | 1,400 vendor tokens per call | Count confirmed from export log |
| Calls billed | 120,000 | 110,000 | 125,000 | 140,000 |
| License fee | 600,000 | 550,000 | 630,000 | 520,000 |
| Usable share of sample | 82% | 80% | 84% | 66% |
| Largest loss | 9% under 60 s of speech | 10% label errors | 6% duplicates from overlapping exports | 14% short calls; 9% transcripts failing WER |
| Usable calls | 98,400 | 88,000 | 105,000 | 92,400 |
| Fee per usable call | 6.10 | 6.25 | 6.00 | 5.63 |
| Vendor-dependent buyer costs | 30,000 audio redaction (only transcripts redacted) | 0 | 15,000 schema mapping | 73,920 outcome labeling at 0.80 per usable call |
| All-in cost per usable call | 6.40 | 6.25 | 6.14 | 6.43 |
| Scope quoted | Fine-tuning and evaluation, one model family, 3 years | Fine-tuning and evaluation, one named model, 2 years | Pre-training, fine-tuning, evaluation, successor models, 5 years | Fine-tuning and evaluation, one model family, 3 years |
What the normalization shows:
- The flat fee only looked cheapest. Vendor D led at 5.63 until the missing outcome labels were priced; it then ranks last.
- The token count is negotiable. Vendor C's tokenizer counted speaker tags and timestamps. The buyer's tokenizer averaged 1,150 transcript tokens per call; billing on that count cuts the fee to 517,500 and the all-in cost to about 5.07.
- Price alone does not separate them. Vendor B's 80% yield from 500 calls carries about ±3.5 points of uncertainty, putting its cost per usable call between about 5.99 and 6.54, a band that contains every other all-in figure.
- The ranking is not yet valid. Vendor C grants far more than Vendor B, so all four must requote one base scope with pre-training and successor models as priced options.
Request quotes on a common price schedule
Much of the normalization work disappears when every vendor fills in the same price schedule. Send it with the request for proposal, and return any quote that leaves a field blank, prices raw volume with no deduplication rule, rests on a vendor-picked sample, describes rights only as "AI use", or holds its unit price only with an exclusivity or minimum commitment it does not price separately.
Illustrative example: invented to show structure; it does not describe an available dataset.
price_schedule_request:
usable_unit: "support call: unique, >= 60 s speech, transcript passes WER spot check, PII redacted in audio and text, labels pass audit"
counting_rules:
record: "one call; both channels count as one call"
hours: "speech duration after voice-activity detection"
tokens: "buyer tokenizer <name, version>, transcript text only"
duplicates: "near-duplicate transcripts (MinHash, Jaccard >= 0.8) billed once"
population:
date_range: "2023-01-01 to 2025-12-31"
estimated_count: "<vendor states count and counting method>"
sample:
size: 500
selection: "random draw from the quoted population; seed recorded"
prepared_like_delivery: true
base_scope: ["fine-tuning", "internal evaluation", "one model family", "3-year term", "trained models survive termination", "non-exclusive"]
priced_options: ["pre-training", "successor models", "12-month field-limited exclusivity", "quarterly refresh"]
price_lines: ["unit price per counted unit", "not-to-exceed total", "price per option", "refresh price and escalator"]
included_work: ["transcripts", "intent and outcome labels", "audio and text redaction", "delivery in buyer schema"]
failed_units: "replaced or credited at unit price after acceptance testing"
payment: "milestones tied to acceptance"
transfer_costs: "party paying egress stated"
quote_validity_days: 90
The AI training data RFP template shows where the schedule sits in a full request for proposal.
From normalized cost to a defensible decision
Use the all-in cost per usable unit at matched scope as the price criterion in your data vendor evaluation scorecard, beside rights evidence, delivery capability and security. Keep the conversion factors, sample seeds, pipeline version and yield results in the decision file so finance can reproduce the ranking.
Then make the sample yield enforceable. Write the usable-unit definition into your acceptance criteria for licensed training data and check each delivery with a sampling plan: ANSI/ASQ Z1.4 prescribes sample sizes and accept/reject numbers from lot size, inspection level and acceptance quality limit (AQL) [11], and acceptance sampling for dataset deliveries adapts that approach to record and label defects. For an independent benchmark before quotes arrive, build a should-cost model for the dataset.
Comparing quotes for licensed operational data?
SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases. Describe the data, your usable-unit definition and the uses you need licensed: SourceX looks for US businesses that hold it, checks the data and the supplier's licensing permissions, agrees pricing and allowed uses in a license, and coordinates delivery and payment. Talk to SourceX about your requirements.
Sources
- Acolad, "Data annotation cost" (vendor page; market practice only). https://www.acolad.com/en/services/data-services/data-annotation-cost
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- Csaki et al., "Efficiently Adapting Pretrained Language Models To New Languages" (2023). https://arxiv.org/pdf/2311.05741
- Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021; ACL 2022). https://arxiv.org/abs/2107.06499v1
- Northcutt, Athalye and Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Longpre et al., "A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity" (2023; NAACL 2024). https://arxiv.org/pdf/2305.13169
- Chen et al., "AlpaGasus: Training A Better Alpaca with Fewer Data" (2023). https://arxiv.org/pdf/2307.08701v1
- Contra, "Robot dataset acceptance audit: a verdict before you pay" (freelance service listing; market practice only). https://contra.com/s/ee6vvEQU-robot-dataset-acceptance-audit-a-verdict-before-you-pay
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version in Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- Amazon Web Services, "Using Requester Pays general purpose buckets for storage transfers and usage" (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/dev/RequesterPaysBuckets.html
- ASQ Quality Press, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables". https://asq.org/quality-press/display-item?item=T1164
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.