Skip to content

Fine-tuning and post-training data

What drives the cost of human preference and SFT data

Quick answer

RLHF and SFT data cost is driven less by the headline rate than by five variables: the pricing unit (per task, comparison, hour, token or license), the expertise each item needs, how long each item takes, how many review layers sit on top, and how much output is rejected or reworked. No authoritative public price list exists as of October 2026, so compare quotes on cost per accepted item under a fixed guideline, not on per-task rates.

By SourceX Editorial · Updated

Which pricing units vendors quote for post-training data

Post-training data is quoted in five units, and each one shifts a different risk onto the buyer. Before comparing numbers, convert every quote into the unit your training run consumes: an accepted SFT demonstration, an accepted preference pair, or a licensed record set.

Pricing unitTypical deliverableRisk the buyer carriesWatch for
Per task or per exampleOne SFT demonstration or one rated promptQuality variance; rejected items may still be billedWhether QA failures are re-done at the vendor's cost
Per comparisonOne pairwise or ranked judgmentLow-signal ties and "negligibly better" labelsHow ties and skips are counted
Per hourAnnotator or expert timeProductivity and scope creepTime logging method, minimum blocks, ramp-up hours
Per tokenVolume of written textIncentive toward verbose responsesLength caps in the guideline
Per dataset licenseExisting records or judgmentsFit to your model and distributionAllowed uses, term, refresh terms

Ranking formats change the arithmetic. InstructGPT labelers ranked several model outputs per prompt, and each ranking was expanded into multiple pairwise comparisons for reward-model training [1]. Llama 2 used binary choices between two responses [2]. A per-comparison quote on a ranking task and a per-comparison quote on a binary task are not the same purchase; see pairwise, ranking or rating formats for the trade-offs.

Expertise level is the largest multiplier on human post-training data

Who writes or judges the item often matters more than any other line item, because expert time is scarce and slow to recruit. A generalist can compare two answers about email etiquette; checking a tax treatment, a differential diagnosis or a Kubernetes failure requires someone who can verify correctness, not just style.

Expertise also changes throughput. Experts take longer per item, need calibration sessions, and often work part time, which stretches calendar time and raises coordination overhead. If you need domain specialists, read domain-expert preference data and contracting domain experts before you set the budget.

Ask vendors to separate three costs: recruiting and vetting, paid calibration time, and production time. Bundled hourly rates hide where the money goes.

Task length, turn count and response format set time per item

Time per item is the second driver, and it scales with prompt length, number of turns, and whether the annotator must write from scratch or only judge. Writing an SFT demonstration is slower than comparing two existing responses, and a multi-turn conversation with tool calls is slower than a single-turn answer.

Specific factors that move time per item:

  • Write versus judge. InstructGPT's pipeline needed contractor-written demonstrations for SFT and separate rankings for the reward model [1]; these are distinct tasks with distinct costs.
  • Reading load. Long contexts (contracts, codebases, logs) make the annotator read before they can judge.
  • Rationale fields. Requiring a written justification, a span-level error tag or a corrected rewrite per comparison can multiply time per item.
  • Multi-attribute rating. HelpSteer2 collected ratings on several attributes per response rather than one preference [5]; richer labels cost more per item and can be worth it for reward modeling.
  • Verification. Code that must be executed, math that must be checked, or citations that must be confirmed add tooling and time.

Review layers, agreement targets and rework drive hidden cost

Quality assurance can add substantially to production cost, because every review layer re-reads the item. A common structure is first pass, peer or senior review, then spot audit by the buyer; each layer has its own rate and rejection rate.

Agreement targets are a cost lever. Collecting two or three independent judgments per pair to measure agreement multiplies the number of paid judgments, and adjudicating disagreements adds a senior pass. The preference data quality and noise guide explains how to decide how much redundancy you need.

Rework is the cost most budgets miss. Guideline revisions after a pilot, re-labeling when the policy model changes, and items rejected at audit all consume budget without producing accepted data. Academic work on preference collection notes that annotation has a clear monetized cost while the value delivered per dollar is rarely measured [6].

Turnaround, volume shape and exclusivity change the quote

Schedule and rights terms shift price even when the task is identical. Compressed turnaround forces vendors to recruit and onboard faster, and uneven volume (large bursts followed by idle weeks) is harder to staff than a steady weekly flow.

Volume shape matters for build-versus-buy too. One annotation vendor argues that in-house teams tend to win at low, steady volume in a narrow domain, while outsourcing changes the total cost picture once you count management and tooling [8]. Treat vendor framing as directional.

Exclusivity and allowed uses are often priced terms. Data written only for you, with restrictions on reuse by the vendor, typically costs more than data the vendor can reuse; confirm what the contract says rather than assuming. For a general view of how licensed data is priced, see what drives the price of licensed enterprise data and how AI data deals are priced.

New collection versus licensing existing judgments and records

Licensing existing data and commissioning new collection have different cost structures, and the right choice depends on whether you need on-policy judgments. New preference collection is typically priced on labor (tasks, comparisons or hours) and is tied to your model's outputs; a licensed dataset is priced as a right to use a fixed set of records.

Existing data has limits. Preference pairs judged on another model's outputs are off-policy for yours, and DPO-style training learns directly from whatever pairs you supply [4]. Open sets such as HelpSteer2 show that permissively licensed human preference data exists [5], but coverage of your domain may be thin; check open datasets that allow commercial fine-tuning before paying for collection.

For SFT, operational records can be a cheaper starting point than writing from scratch. Support resolutions, engineering tickets and expert-reviewed documents already contain real requests and real answers, and converting them is often an editing task rather than an authoring task; see turning business records into instruction-response pairs. SourceX sources operational datasets such as support and sales histories, engineering records and finance and legal workflows from US companies on request, and you can describe the records you need; categories are not inventory and a request does not guarantee a match.

When fewer, better examples cut post-training spend

Quantity is not the main lever for SFT, so the cheapest budget is often a smaller, better-curated one. LIMA fine-tuned a 65B model on 1,000 carefully curated prompt-response pairs, while noting that such curation is labor-intensive [3].

For preference data, AI feedback can replace part of the human budget. Strong LLM judges can approximate human preferences on many prompts but show position and verbosity biases [7], so many teams reserve human judgment for hard, high-stakes or disputed items. The trade-off is covered in AI feedback vs human preference data, and splitting spend across stages is covered in post-training budget allocation.

Worksheet: comparing post-training data quotes per accepted item

Normalize every quote to cost per accepted item before you compare vendors. Use the same guideline, sample prompts and acceptance criteria for every bidder.

Illustrative example: invented to show structure; it does not describe an available dataset.

QUOTE NORMALIZATION WORKSHEET (one row per vendor)

Inputs
  unit_quoted            : per_comparison | per_task | per_hour | per_token | license
  rate                   : vendor's quoted rate in that unit
  items_per_hour         : measured in paid pilot (only if per_hour)
  judgments_per_item     : independent labels per pair (e.g., 1, 2 or 3)
  review_layers          : peer review, senior review, adjudication (rate each)
  pilot_acceptance_rate  : share of pilot items passing your audit
  rework_policy          : vendor re-does rejected items at own cost? yes/no
  one_time_costs         : recruiting, calibration, tooling, guideline iteration

Derived
  production_cost_per_item = rate converted to per item x judgments_per_item
  review_cost_per_item     = sum of review layer costs per item
  cost_per_accepted_item   = (production + review) / pilot_acceptance_rate
                             (skip the division if rework is free to you)
  total_budget             = cost_per_accepted_item x target_accepted_items
                             + one_time_costs

Worked example (no prices; time only)
  Vendor A: generalists, 1 judgment, 1 review layer, 70% pilot acceptance
  Vendor B: domain experts, 2 judgments, adjudication, 92% pilot acceptance
  -> B costs more per raw item but may cost less per accepted, agreed item.

Request the pilot as a paid, fixed-size batch with your own audit, and record acceptance rate, inter-annotator agreement and time per item. These three numbers, not the rate card, predict your production spend. For managed-collection versus licensed options, see buying RLHF comparison data and the fine-tuning and post-training data hub.

Budgeting RLHF and SFT data with SourceX

SourceX sources operational datasets from US companies and manages the commercial process, from assessing data and licensing permissions to agreeing pricing and allowed uses in a license. SourceX does not publish prices; terms are agreed per deal, and nothing is contracted until a supplier agrees. Describe the post-training data you need.

Sources

  1. Ouyang et al., arXiv (OpenAI), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  2. Touvron et al., arXiv (Meta), "Llama 2: Open Foundation and Fine-Tuned Chat Models" (2023). https://arxiv.org/pdf/2307.09288
  3. Zhou et al., arXiv (Meta AI), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  4. Rafailov et al., arXiv (Stanford), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  5. Wang et al., arXiv (NVIDIA), "HelpSteer2: Open-source dataset for training top-performing reward models" (2024). https://arxiv.org/pdf/2406.08673
  6. arXiv, "VickreyFeedback: Cost-efficient Data Construction for Reinforcement Learning from Human Feedback" (2024). https://arxiv.org/pdf/2409.18417
  7. Zheng et al., arXiv, "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (2023). https://arxiv.org/html/2306.05685v4
  8. Acolad, "Data annotation cost". https://www.acolad.com/en/services/data-services/data-annotation-cost

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data