Skip to content

Procurement, samples and ongoing supply

Comparing training data vendor quotes: normalizing price to cost per usable record

Quick answer

To compare training data vendor quotes, convert each one, whether priced per record, per hour, per token or as a flat fee, into cost per usable unit: the license fee plus vendor-dependent costs on your side, divided by the units that pass your deduplication, scope, de-identification and label checks on a random sample from that vendor. Rank quotes only at matched rights scope and term, and have vendors requote rather than adjusting for scope by guesswork.

By SourceX Editorial · Updated

Why quotes for the same data request differ

Quotes for one request differ for five reasons: the billing unit, what each vendor counts as a unit, how much of the delivery survives your checks, the rights granted, and the preparation work included. Conversion factors, a fixed counting rule and sample-based yield put the first three on one basis before you negotiate; differences in rights and included preparation need a common scope and added costs instead.

A language-services vendor, comparing in-house and outsourced annotation, argues that hourly rates rarely decide the real cost; ramp-up time, language coverage and QA at scale do [1].

Price drivers such as history depth, outcome coverage and exclusivity are covered in SourceX's guide to what drives the price of licensed enterprise data, deal structures in how AI data deals are priced, and the choice between flat-fee, per-record and per-token licenses in AI data license pricing structures compared. This page is the side-by-side method within the AI training data procurement guide.

Define the usable unit before you open the price sheets

A usable unit is the record, token or hour that passes every check you will run before training or evaluation, written down with each check's measurement method. Fixing it first stops each quote from setting its own denominator.

ISO/IEC 5259-2 sets out data quality measures [2]; define each one by the property and the method used to quantify it. Apply the same discipline here, so "duplicate" means a named algorithm and threshold, not a reviewer's impression.

UseUsable unitExclusions to stateMeasurement
Supervised fine-tuning (SFT)Prompt-response pair or full conversationExact and near-duplicates; wrong language; truncated responses; failed schema; PII left after redaction; below quality cutoffDedup threshold, language ID, JSON Schema validator, PII scanner, validated grader cutoff
Evaluation setItem with a verified reference answer or labelAmbiguous items; label errors found in adjudication; overlap with training data or public benchmarksDouble review with adjudication; overlap check against your corpora
Pre-training or continued pre-trainingToken counted by your tokenizer after filteringBoilerplate, templates, near-duplicates, off-domain or low-quality textYour tokenizer and version; your filter pipeline
SpeechSpeech hour after silence and hold removal, or one callHold music, IVR prompts, voicemail, transcripts above your word error rate (WER) thresholdVoice-activity detection; WER on a re-transcribed subsample
Operational records (tickets, claims, cases)Case with required fields and an outcomeMissing outcome fields, system-generated records, dates outside rangeField-completeness and range validators

See near-duplicate detection with MinHash and LSH and quality filtering for instruction-tuning data for the checks themselves.

Convert per-hour, per-token, per-record and flat-fee prices with measured factors

Every pricing unit converts to your usable unit through a factor, and that factor must come from the vendor's own sample, not a typical value. A per-hour price becomes a per-call price only once you know the average call length in that vendor's population.

Quoted unitFactor to measure on the sampleWhere it goes wrong
Per recordRecords per usable unit (messages per thread, lines per claim, turns per conversation)"Record" means a message to one vendor and a thread to another
Per hour of audio or videoRecorded minutes per record; speech minutes per recorded minuteFile duration includes hold and silence; dual-channel files billed twice
Per annotation hourItems completed per hour at your QA bar, including review passesA low rate with low throughput or no second review
Per tokenVendor tokens per record and your tokens per recordUndisclosed tokenizer; markup, speaker tags and metadata billed as content
Flat feeRecord count in the defined population"All available data" with no count, date range or filter

Token counts are the least portable unit. Tokens per word, called fertility, depend on the tokenizer's vocabulary and the language: one study reduced fertility by 42% for Hungarian and 73% for Thai by replacing about 10% of a model's vocabulary [3]. Count the sample with your own tokenizer and version, and ask vendors to bill on that count; see per-token training data pricing and estimating token counts before licensing. For audio, state whether a billable hour is file, speech or transcribed duration; how speech data is priced covers the add-ons.

Measure each vendor's yield on a random sample

Yield is the share of billed units that become usable units, and a yield gap alone can reverse a price ranking. Measure it on a random sample drawn from the exact population each vendor quoted, using one pipeline for every vendor.

Typical losses:

  • Duplicates. Lee et al. found that standard language-modeling corpora contain many near-duplicates and long repeated substrings, including one sentence repeated more than 60,000 times in C4 [4]. Business exports add overlapping re-exports, macros and auto-replies.
  • Label errors. An audit of test sets from 10 widely used datasets estimated an average label error rate of at least 3.3% [5]. Budget re-labeling or rejection for annotated items.
  • Quality filtering. In controlled pretraining experiments, quality filters improved downstream performance even though they removed 10% or more of the training data [6]. For instruction data, AlpaGasus kept about 9,000 of the 52,000 Alpaca examples after LLM-graded filtering and reported a better model [7].
  • Out-of-scope and emptied records. Wrong language, dates outside the requested range, and records left empty after PII removal.

A robot-dataset audit provider claims that quality filtering typically discards 20 to 30% of a collected dataset [8]; that figure is self-reported and comes from robotics data, so measure your own. With 500 sampled records and an observed yield of 80%, the 95% interval is roughly ±3.5 points (1.96 × √(0.8 × 0.2 ÷ 500)), so convert that band into a cost range for each vendor and treat quotes whose ranges overlap as ties.

How many records to check covers sample sizing, and requesting a training data sample covers getting a random draw rather than a showcase.

Normalize rights scope before you compare price

A quote is comparable only with quotes that grant the same uses, models, term and exclusivity. Do not discount a broad license by a guessed percentage; ask every vendor to price one common base scope plus the same list of priced options.

DimensionBase scope to request (example)Options to price separately
Permitted useFine-tuning and internal evaluationPre-training; synthetic data generation
ModelsOne model familySuccessor models or all company models (per-model vs enterprise-wide licenses)
Term and survival3 years; trained models survive terminationPerpetual term (model retention after termination)
ExclusivityNon-exclusiveTime-boxed or field-limited exclusivity (negotiating exclusivity)
UsersLicensee and named affiliatesCustomer fine-tuning; sublicensing
RefreshNonePeriodic refresh at a fixed price per delivery

An evaluation-only quote buys a different right, not a cheaper training license, so give it its own column (see evaluation-only data license terms). Confirm that each vendor holds the rights it prices: an audit of more than 1,800 text datasets found license omission rates above 70% and error rates above 50% on popular hosting sites [9].

SourceX prepares diligence materials on source, rights, preparation and allowed use per dataset for the buyer's review. It does not publish prices; terms depend on scope, volume, history, rights and exclusivity, so state your base scope and options when you describe the data you need to SourceX.

Add only the costs that change with the vendor

After price and scope are normalized, add the costs on your side that differ between vendors; costs identical for every vendor do not change the ranking and belong in the total cost of ownership model. Price these per vendor:

  1. Missing preparation. Labels, transcripts, redaction or de-identification left to you, at your internal or third-party rate.
  2. Schema and format work. Engineering to map a vendor's per-turn JSON or proprietary export into your schema; zero if it delivers to your delivery specification.
  3. Rework. Re-labeling the error share you measured, unless the vendor replaces failed units at its own cost.
  4. Transfer. With an Amazon S3 Requester Pays bucket, the requester pays for requests and downloads while the bucket owner pays for storage [10]; see who pays for data egress.
  5. Review and payment terms. Extra legal and security review for non-standard paper; prepayment versus milestone payments tied to acceptance; minimum commitments and refresh escalators.

Worked example: four quotes for contact-center call recordings

Illustrative example: invented to show structure; it does not describe an available dataset. Figures are in a generic currency and are not market prices.

A buyer wants English support-call recordings with transcripts plus intent and outcome labels, to fine-tune and evaluate a voice support agent. A usable call is unique, has at least 60 seconds of speech, has a transcript that passes a WER spot check, has names and account numbers redacted in audio and text, and carries labels that pass audit. Each vendor supplied a random 500-call sample from the population it quoted.

Vendor AVendor BVendor CVendor D
Quoted price40 per recorded hour5.00 per delivered call3.60 per 1,000 transcript tokens, vendor tokenizer520,000 flat fee
Quoted volume15,000 recorded hours110,000 calls175 million tokensAll 2023-2025 calls, about 140,000
Factor from sample7.5 recorded minutes per callNone needed1,400 vendor tokens per callCount confirmed from export log
Calls billed120,000110,000125,000140,000
License fee600,000550,000630,000520,000
Usable share of sample82%80%84%66%
Largest loss9% under 60 s of speech10% label errors6% duplicates from overlapping exports14% short calls; 9% transcripts failing WER
Usable calls98,40088,000105,00092,400
Fee per usable call6.106.256.005.63
Vendor-dependent buyer costs30,000 audio redaction (only transcripts redacted)015,000 schema mapping73,920 outcome labeling at 0.80 per usable call
All-in cost per usable call6.406.256.146.43
Scope quotedFine-tuning and evaluation, one model family, 3 yearsFine-tuning and evaluation, one named model, 2 yearsPre-training, fine-tuning, evaluation, successor models, 5 yearsFine-tuning and evaluation, one model family, 3 years

What the normalization shows:

  • The flat fee only looked cheapest. Vendor D led at 5.63 until the missing outcome labels were priced; it then ranks last.
  • The token count is negotiable. Vendor C's tokenizer counted speaker tags and timestamps. The buyer's tokenizer averaged 1,150 transcript tokens per call; billing on that count cuts the fee to 517,500 and the all-in cost to about 5.07.
  • Price alone does not separate them. Vendor B's 80% yield from 500 calls carries about ±3.5 points of uncertainty, putting its cost per usable call between about 5.99 and 6.54, a band that contains every other all-in figure.
  • The ranking is not yet valid. Vendor C grants far more than Vendor B, so all four must requote one base scope with pre-training and successor models as priced options.

Request quotes on a common price schedule

Much of the normalization work disappears when every vendor fills in the same price schedule. Send it with the request for proposal, and return any quote that leaves a field blank, prices raw volume with no deduplication rule, rests on a vendor-picked sample, describes rights only as "AI use", or holds its unit price only with an exclusivity or minimum commitment it does not price separately.

Illustrative example: invented to show structure; it does not describe an available dataset.

price_schedule_request:
  usable_unit: "support call: unique, >= 60 s speech, transcript passes WER spot check, PII redacted in audio and text, labels pass audit"
  counting_rules:
    record: "one call; both channels count as one call"
    hours: "speech duration after voice-activity detection"
    tokens: "buyer tokenizer <name, version>, transcript text only"
    duplicates: "near-duplicate transcripts (MinHash, Jaccard >= 0.8) billed once"
  population:
    date_range: "2023-01-01 to 2025-12-31"
    estimated_count: "<vendor states count and counting method>"
  sample:
    size: 500
    selection: "random draw from the quoted population; seed recorded"
    prepared_like_delivery: true
  base_scope: ["fine-tuning", "internal evaluation", "one model family", "3-year term", "trained models survive termination", "non-exclusive"]
  priced_options: ["pre-training", "successor models", "12-month field-limited exclusivity", "quarterly refresh"]
  price_lines: ["unit price per counted unit", "not-to-exceed total", "price per option", "refresh price and escalator"]
  included_work: ["transcripts", "intent and outcome labels", "audio and text redaction", "delivery in buyer schema"]
  failed_units: "replaced or credited at unit price after acceptance testing"
  payment: "milestones tied to acceptance"
  transfer_costs: "party paying egress stated"
  quote_validity_days: 90

The AI training data RFP template shows where the schedule sits in a full request for proposal.

From normalized cost to a defensible decision

Use the all-in cost per usable unit at matched scope as the price criterion in your data vendor evaluation scorecard, beside rights evidence, delivery capability and security. Keep the conversion factors, sample seeds, pipeline version and yield results in the decision file so finance can reproduce the ranking.

Then make the sample yield enforceable. Write the usable-unit definition into your acceptance criteria for licensed training data and check each delivery with a sampling plan: ANSI/ASQ Z1.4 prescribes sample sizes and accept/reject numbers from lot size, inspection level and acceptance quality limit (AQL) [11], and acceptance sampling for dataset deliveries adapts that approach to record and label defects. For an independent benchmark before quotes arrive, build a should-cost model for the dataset.

Comparing quotes for licensed operational data?

SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases. Describe the data, your usable-unit definition and the uses you need licensed: SourceX looks for US businesses that hold it, checks the data and the supplier's licensing permissions, agrees pricing and allowed uses in a license, and coordinates delivery and payment. Talk to SourceX about your requirements.

Sources

  1. Acolad, "Data annotation cost" (vendor page; market practice only). https://www.acolad.com/en/services/data-services/data-annotation-cost
  2. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  3. Csaki et al., "Efficiently Adapting Pretrained Language Models To New Languages" (2023). https://arxiv.org/pdf/2311.05741
  4. Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021; ACL 2022). https://arxiv.org/abs/2107.06499v1
  5. Northcutt, Athalye and Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  6. Longpre et al., "A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity" (2023; NAACL 2024). https://arxiv.org/pdf/2305.13169
  7. Chen et al., "AlpaGasus: Training A Better Alpaca with Fewer Data" (2023). https://arxiv.org/pdf/2307.08701v1
  8. Contra, "Robot dataset acceptance audit: a verdict before you pay" (freelance service listing; market practice only). https://contra.com/s/ee6vvEQU-robot-dataset-acceptance-audit-a-verdict-before-you-pay
  9. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version in Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
  10. Amazon Web Services, "Using Requester Pays general purpose buckets for storage transfers and usage" (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/dev/RequesterPaysBuckets.html
  11. ASQ Quality Press, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables". https://asq.org/quality-press/display-item?item=T1164

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data