Procurement, samples and ongoing supply
Building an AI training data budget across a model roadmap
Quick answer
An AI training data budget is the annual plan for every cost of getting licensed data into pre-training, fine-tuning, evaluation and retrieval, not just license fees. Build it from the model roadmap: list the datasets each release needs, price each one as a low-high range across samples, license, preparation, review, transfer, integration and refreshes, add a reserve tied to named risks, and release money at evidence gates. Show finance one-time, recurring and multi-year committed spend separately.
By SourceX Editorial · Updated
This page covers the plan across many purchases. For one dataset's full cost, use the total cost of ownership model; to justify one purchase, see the business case for buying training data; to spend as little as possible on a first dataset, see data licensing for AI startups on a budget.
The nine cost lines a data budget carries
A complete data budget carries nine lines per dataset, and most of them never appear on a supplier's quote.
| Line | What drives the cost | How to estimate it | Cost type |
|---|---|---|---|
| Samples and evaluation licenses | Candidate suppliers; tests per sample | Suppliers × test hours, plus sample fees | One-time |
| Paid pilot | Pilot fee; your team's test time | Quoted fee less any license credit | One-time, often creditable |
| License fee | Permitted uses, models, term, exclusivity, volume, history depth | Quotes or a should-cost range | One-time, recurring or minimum commitment |
| Preparation and annotation | Deduplication, PII redaction checks, labeling, expert review, adjudication | Usable units × rate × review passes | Per delivery |
| Rights, privacy and security review | Counsel hours, privacy assessment, security questionnaire, statistical expert for health data | Hours × rate, per supplier | Per supplier; repeats at renewal |
| Transfer and storage | Delivery method, egress, retention period | Volume × cloud rates; payer set by contract | Recurring |
| Integration engineering | Schema mapping, loaders, eval harness, retrieval index builds | Engineer-weeks | One-time, plus per refresh |
| Refreshes and renewals | Cadence, per-delivery price, escalators | Deliveries × price | Recurring |
| Reserve | Failed pilots, re-delivery, re-annotation, late rights findings | Named risks × cost | Released only on approval |
Two costs are often missed. For health data de-identified by HIPAA Expert Determination rather than Safe Harbor, a person with appropriate statistical expertise must find identification risk very small, with no numerical threshold set by HHS [8]: a paid engagement (see Safe Harbor vs Expert Determination). Transfer cost follows the delivery method: with an Amazon S3 Requester Pays bucket the requester pays for requests and downloads [12], while Snowflake Secure Data Sharing copies no data and adds nothing to the consumer's storage charges [13].
Derive the line items from the model roadmap
Start from each release on the model roadmap, because the training stage sets the budget unit, the volume and whether the spend recurs. Record each item as release, dataset, licensed use, usable units needed and the date the data must be ready.
| Stage | Budget unit | What moves the cost | Recurs? |
|---|---|---|---|
| Pre-training or continued pre-training | Tokens after your filters | Volume, then losses to deduplication and quality filters | Per model generation |
| Supervised fine-tuning | Curated example or conversation | Expert curation and review per example | Per release |
| Evaluation | Item with a verified reference answer | Double review, adjudication, contamination checks | When drift or contamination appears |
| Retrieval (RAG) | Documents in the index | Refresh cadence, re-indexing, retrieval and display rights | Every period of the license term |
Raw volume overstates pre-training value: Lee et al. found one sentence repeated more than 60,000 times in C4 [4], so budget on deduplicated counts. Fine-tuning can need few examples at a high cost each; LIMA tuned a 65B-parameter model on 1,000 curated prompt-response pairs, and its authors call that curation labor-intensive [6]. Evaluation sets are small but need careful review: an audit of test sets from 10 widely used datasets estimated an average label error rate of at least 3.3% [5].
How much training data to buy and procurement by training stage cover sizing. Sample and pilot spend lands quarters before the license fee; see a realistic licensing timeline.
Phase spend at gates: sample, pilot, purchase, refresh, renewal
Release each tranche only when the previous gate has produced evidence that the data works for your model. Five gates cover most purchases:
- Sample and evaluation license. The evidence is a representative sample that passes your schema, yield and rights checks. See requesting a training data sample.
- Paid pilot. The evidence is a measured gain on a held-out evaluation set (estimating a dataset's value). Negotiate a license credit and price lock (paid data pilot terms).
- First full purchase. Pay in tranches tied to acceptance criteria, not at signature; milestone payments tied to acceptance shows the structure.
- Refresh deliveries. Recurring spend per period, with volume bands and an escalator. Reddit's 2024 S-1 shows this shape from the seller's side: data licensing arrangements with two-to-three-year terms, delivered through continuous API access and quarterly data transfers [2]. See ongoing data supply agreements.
- Renewal or exit. Budget the decision, not an automatic renewal: renew, renegotiate or let a license lapse.
Estimating lines when suppliers publish no price lists
Carry every external line as a low-high range built from quotes, a should-cost estimate or both, and narrow it as gates produce quotes. A single figure entered before any supplier quotes is false precision.
Wide ranges are normal. For in-house annotation, a language-services vendor argues that workforce, tooling, QA and ramp-up costs compound with scale and matter more than hourly rates [1].
Three rules keep ranges honest. Convert every quote to cost per usable unit (comparing vendor quotes). Ignore headline AI deal values, which bundle rights and scale your purchase will not match (what drives the price of licensed enterprise data). Price rights you will need later, such as successor models or pre-training, as options now (pricing structures compared).
The same applies to data sourced through SourceX, which does not publish prices: terms depend on scope, volume, history, rights and exclusivity and are agreed per deal in writing. Datasets are sourced on request, not held in stock, and a request does not guarantee a match, so describe the datasets on your roadmap to SourceX early and keep the line conditional until a supplier agrees terms.
Worked example: one fiscal year, two model releases
Illustrative example: invented to show structure; it does not describe an available dataset. Figures are in thousands of a generic currency and are not market prices.
Release 1 (Q2) is a support agent fine-tuned on licensed ticket histories and tested on a private evaluation set; Release 2 (Q4) adds retrieval over a licensed documentation corpus refreshed quarterly.
| # | Line | Gate that releases it | Low | High | Cost type |
|---|---|---|---|---|---|
| 1 | Samples and evaluation licenses: 3 ticket suppliers, 2 documentation suppliers (Q1) | Budget approval | 10 | 25 | One-time |
| 2 | Paid pilot on ticket histories, credited if converted (Q1) | Sample passes tests | 30 | 60 | One-time, creditable |
| 3 | Ticket-history license: fine-tuning and evaluation, 2-year term, paid at delivery and acceptance (Q2) | Pilot shows target gain | 250 | 600 | One-time fee, multi-year rights |
| 4 | Ticket preparation: deduplication, redaction checks, schema mapping (Q2) | Delivery received | 40 | 90 | One-time |
| 5 | Private eval set: 2,000 items, two expert reviewers plus adjudication (Q2) | Item spec signed off | 60 | 140 | One-time |
| 6 | Rights, privacy and security review, two suppliers (Q1-Q2) | Shortlist agreed | 35 | 80 | Per supplier |
| 7 | Documentation subscription, 12 months from Q3, quarterly refresh (Q3-Q4 share) | Release 1 shipped | 90 | 200 | Recurring |
| 8 | Refresh ingestion and re-indexing (Q3-Q4) | Each refresh accepted | 20 | 45 | Recurring |
| 9 | Transfer, storage and access control (Q2-Q4) | Not gated | 5 | 15 | Recurring |
| 10 | Supplier records for disclosure duties (Q2-Q4) | Not gated | 10 | 30 | Per dataset |
| Subtotal | 550 | 1,285 | |||
| 11 | Reserve: a second pilot with an already-sampled supplier if the first fails (30-60) plus re-annotation of 10% of the eval set (6-14) | Program owner and finance | 36 | 74 | Contingent |
| Fiscal-year request | 586 | 1,359 |
What the example shows:
- The license is under half the plan. Line 3 is 45% of the low subtotal and 47% of the high one.
- The first quotes narrow the range most. The license accounts for 350 of the 773 difference between the low and high requests.
- Commitments cross the fiscal year. The subscription's last two quarters (90-200) and their refresh work (20-45) belong in next year's plan now; the ticket license needs a renewal decision before month 24.
Separate one-time, recurring and committed spend for finance
Finance approvers need three views of the same lines: cash by quarter, the recurring run-rate after the first purchase, and commitments that bind future fiscal years. List each commitment with its end date, notice date, minimum volume and escalator.
Agree the accounting treatment with your controller before the budget is locked, because it decides whether a fee hits this year's operating expense. Under IFRS, IAS 38 recognizes an intangible asset only if it is identifiable (for example, arising from contractual rights) and controlled, its benefits are probable and its cost is reliably measurable; spend that fails these tests is expensed when incurred [3]. US GAAP has its own rules, and license term and perpetual rights can change the answer.
Record a pilot credit as a conditional reduction of the license line, valid only within the option window. Milestone payments move cash to acceptance dates, so quarterly cash should follow the acceptance schedule. Buyers outside the US should add currency and cross-border terms.
Compliance records that now need budgeted time
Disclosure duties in force or scheduled as of October 2026 make supplier records a budget line: collecting source, rights and personal-data facts at purchase is easier than reconstructing them after a release. Budget one record per dataset covering supplier, source systems, collection period, rights basis, personal-data handling, de-identification method and synthetic share.
- California AB 2013. Developers of generative AI systems made available to Californians (covering systems released on or after 1 January 2022) had to post training-data documentation by 1 January 2026, and again before each later covered system or substantial modification; it covers dataset sources or owners, copyrighted or licensed material, personal information, collection periods and synthetic data [9].
- EU AI Act, Article 53(1)(d). Providers of general-purpose AI models must publish a summary of training content; the Commission published the template on 24 July 2025 [10].
- Colorado SB26-189. Signed 14 May 2026; from 1 January 2027, developers of automated decision-making technology that materially influences consequential decisions must give deployers documentation including training data categories [11].
Open datasets need review time too: an audit of more than 1,800 text datasets found license omission rates above 70% and error rates above 50% on popular hosting sites [7]. SourceX prepares diligence materials on source, rights, preparation and allowed use for each dataset, for the buyer's review; your counsel's hours to review them still belong in the plan. Disclosure requirements compared maps the fields to collect.
Sizing the reserve for failed pilots and re-delivery
Size the reserve from named risks with a cost each, not a flat percentage, and require an approval to draw on it.
| Risk | What it costs if it happens | Basis for the amount |
|---|---|---|
| Pilot fails | Next supplier's sample, pilot and review | Those lines for one more supplier |
| Defects found after acceptance | Re-annotation or re-cleaning | Measured sample error rate, less what contract remedies cover |
| Yield below plan | Extra volume to reach target | Yield gap × unit price |
| Health-data refresh needs a new determination | The expert's re-analysis | HHS: changes in technology and available data can warrant re-examination [8] |
| Scope added mid-term | Price of successor-model or pre-training rights | The option prices negotiated in advance |
Running the budget through the year
Review the data budget at every gate and at least quarterly against four numbers: committed versus spent, cost per usable unit against plan, reserve drawn, and commitments ending in the next two quarters. Reallocate between datasets only through the approver who signed the plan; data procurement KPIs covers tracking.
Checklist before the plan goes to the finance approver:
- Every dataset traces to a roadmap release and a stated licensed use.
- Each external line is a low-high range with its basis: quote, should-cost or prior purchase.
- One-time, recurring and committed spend appear separately, with end and notice dates.
- Pilot credits, milestone payments and refresh escalators are reflected in quarterly cash.
- Legal, privacy and security review hours and a disclosure record per dataset are included, even when another cost center pays.
- The reserve lists named risks, amounts and an approver.
- The controller has agreed capitalization or expense treatment.
Category strategy for AI data spend covers splitting spend by sub-category; the AI training data procurement guide maps every other stage.
Planning data spend that depends on licensed business records?
SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases. Describe the datasets on your roadmap, the uses you need licensed and your delivery needs: SourceX looks for US businesses that hold the data, checks it and the supplier's licensing permissions, agrees pricing and allowed uses in a license, and coordinates delivery and payment. Talk to SourceX about your data requirements.
Sources
- Acolad, "Data annotation cost" (vendor page; market practice only). https://www.acolad.com/en/services/data-services/data-annotation-cost
- U.S. Securities and Exchange Commission (EDGAR), Reddit, Inc., "Form S-1 Registration Statement" (2024). https://www.sec.gov/Archives/edgar/data/1713445/000162828024006294/reddits-1q423.htm
- Deloitte IAS Plus, "IAS 38 - Intangible Assets" (standard summary). https://www.iasplus.com/en/standards/ias/ias38
- Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021; ACL 2022). https://arxiv.org/abs/2107.06499v1
- Northcutt, Athalye and Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Zhou et al., "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version in Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (Chapter 817, Statutes of 2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- European Commission, "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
- Amazon Web Services, "Using Requester Pays general purpose buckets for storage transfers and usage" (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/dev/RequesterPaysBuckets.html
- Snowflake, "About Secure Data Sharing" (documentation). https://docs.snowflake.com/en/user-guide/data-sharing-intro.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.