Skip to content

Fine-tuning and post-training data

Buying RLHF comparison data: collection services, expert networks or licensed judgments

Quick answer

RLHF data collection services put trained raters in front of responses sampled from your checkpoint. Commissioned raters from a managed service, a dedicated team or an expert network (for comparisons that turn on a domain fact) supply on-policy judgments of your current model's outputs. Licensed judgments, such as supervisor approvals, edits and QA scores recorded in business systems, are off-policy but can serve as reward-model warm starts, rater golds and evaluation sets. The routes can be combined, and the statement of work decides quality more than the vendor's name.

By SourceX Editorial · Updated

For the concept itself, see what RLHF data is and the RLHF glossary entry.

Four ways to buy human comparisons, and what each delivers

The four procurement models differ in who writes the responses, who judges them, how you pay, and whether the comparisons describe your current model.

The judgment is the product in every model. InstructGPT trained a reward model on human rankings of model outputs before reinforcement learning, using contractors hired through Upwork or sourced from Scale AI and selected on screening criteria [1]. Direct Preference Optimization (DPO) fits the policy to chosen and rejected pairs with no separate reward model [2], so a mislabeled pair shapes the policy directly.

ModelResponses fromWho judgesCommon pricing unitOn-policy?FitsMain risk
Managed collection serviceYour checkpointVendor-recruited raters trained on your guidelinesPer comparisonYes, with fresh samples each roundHigh-volume helpfulness and harmlessness roundsRater churn and guideline drift between batches
Dedicated managed teamYour checkpointA named, stable teamPer reviewer hour or managed SLAYesUnreleased models, multi-turn and agentic tasksPaying for idle capacity; know-how stays with the vendor
Expert network or credentialed poolYour checkpointPractitioners recruited per specialtyPer hour or per task at expert ratesYesComparisons decided by clinical, legal, tax or security factsThin supply per specialty; credential checks; employer conflicts
Licensed existing judgmentsEmployees, customers or earlier systemsSupervisors, compliance reviewers and QA graders doing their jobsLicense terms agreed per dealNoWarm starts, rater golds, evaluation sets, rubric seedsBusiness criteria behind the choice; personal data; rights chain

As market practice, Sama describes managed teams producing preference rankings under detailed guidelines with multi-layer QA review [3]. Digital Divide Data describes programs spanning rater recruitment, calibration, rubric design, agreement measurement and adjudication, priced per comparison, per reviewer hour or under a managed SLA [4]. Some providers also sell ready-made preference datasets [5]; check those like open sets, as in open preference datasets that allow commercial fine-tuning.

You can also run raters yourself on a tool such as Label Studio, whose pairwise template shows a prompt with two candidate responses [6]. Prompts and outputs stay in-house, but you recruit, contract, pay and check every rater.

Matching the procurement model to your post-training run

Four questions settle the choice: must the comparisons describe your current checkpoint, does a domain fact decide them, does the signal already exist in recorded decisions, and how confidential are the prompts and outputs?

  • Current checkpoint. On-policy data comes from the policy being trained, and samples from an earlier policy count as off-policy for a later one [7]. Each training round therefore needs fresh comparisons, which a service or team can supply and an archive cannot. The value depends on the task: SimpleMix found on-policy data most effective on objective tasks such as math and coding, off-policy data on open-ended tasks, and a mix better than either [8]. See on-policy vs off-policy preference data.
  • Domain facts. Generalists cannot judge a drug interaction or a filing deadline. Route that slice to experts; see domain-expert preference data and contracting domain experts.
  • Recorded decisions. If reviewers in a business already choose between drafts, versions or proposals, licensing those records avoids recreating the judgment.
  • Confidentiality. Outputs from an unreleased model reveal its capabilities, so prefer a dedicated team or your own raters and ask how many people will see them.

Lee et al. found that training on AI-generated preference labels performed comparably to training on human labels for summarization and helpful dialogue, and better for harmless dialogue [9], so consider reserving paid human work for slices where AI judges prove unreliable. See AI feedback vs human preference data and what drives the cost of human preference data.

What licensed judgments can replace, and what they cannot

Licensed judgments replace commissioned comparisons only when you need a record of how qualified people chose between real alternatives, not a verdict on your own model's outputs.

A QA scorecard grades a support reply item by item, a supervisor's edit pairs a draft with the approved final, a compliance review returns one version and approves another, and code review separates approved changes from requested ones. These records support four uses:

  • Warm starts. One bandit-learning study shows that an offline preference dataset from an expert of unknown competence can warm-start online learning when the method models that competence [10].
  • Golds for vendor raters. A reviewer's recorded decision is a known answer your vendor did not write.
  • Evaluation. Held-out reviewer decisions can check a reward model or an LLM judge; see LLM-as-a-judge calibration sets.
  • Rubric seeds. Reason codes and scorecard items show the criteria practitioners already apply.

The limits: the responses did not come from your model, a rejection can reflect price or negotiation rather than quality, records often keep only the approved final, and both sides can contain personal data. Apply the conversion rules in preference datasets for DPO before any record becomes a training pair.

SourceX works on this licensed route. It sources operational datasets from US companies, including support and sales histories, documents, and finance and legal workflows, on request rather than from stock, so a request does not guarantee a match. Every dataset goes through rights review and is delivered under a license that defines the included records and permitted uses.

Names, emails, phone numbers and account numbers are removed or replaced before delivery, and no de-identification method is perfect. To scope review records as preference data, describe the decisions and fields you need, or see QA scores and corrections.

Statement of work for RLHF data collection: clauses that decide quality

An RLHF statement of work (SOW) should fix response generation, the rater pool, quality measurement, acceptance, deliverables, confidentiality and ownership before the pilot, because each is expensive to renegotiate once raters are trained. The general structure is in writing a data collection statement of work; these rows are specific to comparisons.

SOW sectionWhat to specifyFailure it prevents
Response generationWho samples (you, or the vendor through an endpoint you control), checkpoint IDs, sampling settings, pairwise or k-way rankingComparisons of a stale or unrecorded checkpoint
PromptsSource, ownership, task mix, whether raters may write prompts; see real-world prompt setsOff-distribution or benchmark-copied prompts
Guidelines and rubricBuyer-owned, versioned, change control, who pays to re-rateSilent drift between batches
Judgment formatStrength scale, tie and "both unacceptable" options, rationales, per-criterion verdicts; see pairwise, ranking or rating formatsForced choices between near-identical responses
Rater poolScreening test and pass mark, locales, credentials for flagged slices, share of pilot raters kept in productionA senior team runs the pilot and new hires run production
Golds and calibrationQualification set, gold share in live work, golds written by you, removal rules; see rater calibrationAccuracy never measured independently
Overlap and agreementShare double-rated, statistic per criterion, threshold fixed after the pilotAgreement reported only on the easy overall label
ThroughputBatch size, weekly capacity, turnaround, surge noticeLabels arriving after their training round
Acceptance and reworkAcceptance sample per batch, rejection triggers (gold accuracy, timing, position bias), rework at vendor cost, payment per accepted comparison; see acceptance samplingPaying for noise
DeliverablesJSONL per comparison with every rater's label, pseudonymous rater ID, rubric version, seconds on task, adjudication logMajority votes with no audit trail
AI assistanceWhether raters may use chat assistants; disclosure and detection; see detecting model-generated contentModel-written "human" rationales
ConfidentialityNDAs flowing to every rater, access-controlled tooling, no local copies, rater locations, deletion at term endLeaks of unreleased outputs
SubcontractingNamed sub-vendors and crowd platforms, flow-down of terms, approval before changesAn unknown labor chain
OwnershipAssignment of labels, rationales and rewrites; vendor tools carved out; warranty that each rater signed an assignmentGaps in the rights chain

Three rows need detail. For agreement, Krippendorff's alpha handles any number of raters and items not every rater saw; 1 means perfect reliability and 0 means chance [11]. For deliverables, require every rater's label: standard DPO treats high-disagreement pairs like unanimous ones [12], and you cannot reweight labels you never received.

For ownership, written judgments such as rationales can be protected expression. On 29 September 2026 the Third Circuit held the Westlaw headnotes at issue in Thomson Reuters v. ROSS copyrightable, and ROSS's use of them to train a non-generative legal-research tool not fair use [13]. See IP assignment vs license for commissioned data.

Running the pilot: one batch through every shortlisted service

Send the same pilot batch to each shortlisted service with your own known-answer items mixed in. Decide on agreement, accuracy against your golds and model gain, not on the vendor's QA report.

Summaries of the Secrets of RLHF Part II study report typical inter-annotator agreement of 63-72% on general preference labels [14]. A vendor reporting much higher agreement on subjective comparisons should show how it counted ties, easy pairs and adjudicated items.

Illustrative example: invented to show structure; it does not describe an available dataset.

A team tuning a support assistant samples 400 de-identified production prompts held out from supervised fine-tuning and draws two responses per prompt from its current checkpoint, randomizing order for every rater. Both vendors rate all 400 pairs, double-rate 100 and receive 30 internally adjudicated golds inserted blind. With decision rules fixed in advance, it trains a reward model on each vendor's labels and scores it on 200 separate adjudicated pairs.

Pilot measureVendor AVendor BReading
Gold accuracy27 of 3021 of 30B misses correctness golds
Alpha, correctness criterion0.710.48B's raters disagree on facts
First-position response chosen51%58%B shows position bias
Median seconds per comparison9531B's pace is too fast to read both responses
Tie option used12%Not offeredB's tool forced choices
Reward-model accuracy, internal set68%61%Model gain follows label quality

The team signs with Vendor A, writes the tie option and a minimum time on task into the SOW, and repeats the gold check on every batch. See paid data pilot terms and estimating a dataset's value before purchase.

Vendor claims to verify before you scale

Workforce size, "expert" coverage and multi-layer QA are self-reported, so ask for evidence you can check in your own pilot data.

  • "Thousands of vetted raters." How many worked on your pilot, how many will work in production, and what share stays between batches?
  • "Domain experts." Credential type, jurisdiction, verification method and date per pseudonymous rater; see verifying expert annotator qualifications.
  • "Multi-layer QA." Rejection rates per layer, and both original and reviewed labels when reviewers change them.
  • "No AI use." The written policy, the detection method and past incidents.
  • Rankings. At least one published ranking of RLHF data collection platforms measures how often brands appear in AI assistant answers, not data quality [15].

In production, watch for falling time on task, changing rater IDs, slipping gold accuracy, near-identical rationales and strength labels collapsing to one value. See measuring noise in preference data and provenance for human-annotated data.

Records to require for your own disclosures

Require the service or licensor to deliver the facts you will need to document the data later: where prompts and responses came from, who judged them under which guidelines, and whether any part was synthetic.

  • California AB 2013. As of October 2026, developers of generative AI systems made publicly available to Californians must post training-data documentation with a high-level summary of the datasets. The first deadline was 1 January 2026, and the duty repeats before each later release or substantial modification, which the statute defines to include the results of retraining or fine-tuning [16]. Summaries list items such as data sources and ownership, collection and processing, and synthetic data use [17].
  • EU AI Act Article 10. For high-risk systems, Article 10(2) requires governance of data preparation, including annotation and labelling [18]. As of October 2026, Regulation (EU) 2026/1744 is in the Official Journal, and secondary sources report that Annex III high-risk obligations now apply from 2 December 2027 ; the Regulation also amends Article 10 [19].

See fine-tuning provider duties under the EU AI Act and AB 2013, how to procure enterprise training data and the fine-tuning and post-training data hub.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Considering licensed judgments alongside a collection service?

Describe the decisions you want a model to learn from, such as approvals, edits, QA scores or rejections with reasons, along with the fields, volume and uses the license must cover. SourceX looks for US businesses that hold those records, checks the data and the supplier's licensing permissions, and manages the license and delivery. Nothing is contracted until a supplier agrees. Describe the preference judgments you need.

Sources

  1. Ouyang et al. (OpenAI), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  2. Rafailov et al. (Stanford), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  3. Sama (vendor page), "RLHF Training Data for LLM Alignment". https://info.sama.com/high-quality-human-feedback-for-large-language-models
  4. Digital Divide Data (vendor page), "RLHF services and pricing". https://www.digitaldividedata.com/?p=24655
  5. Argos Multilingual (vendor page), "Reinforcement Learning from Human Feedback (RLHF)". https://data.argosmultilingual.com/llm-training-data/reinforcement-learning-from-human-feedback-rlhf/
  6. Label Studio (product documentation), "Human Preferences collection for RLHF". https://labelstud.io/templates/generative-pairwise-human-preference
  7. arXiv, "Towards a Unified View of Preference Learning for Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2409.02795
  8. Proceedings of Machine Learning Research (ICML 2025), "SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning" (2025). https://proceedings.mlr.press/v267/li25au.html
  9. Lee et al. (ICML 2024), "RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback" (2023). https://arxiv.org/pdf/2309.00267
  10. arXiv, "Online Bandit Learning with Offline Preference Data for Improved RLHF" (2024). https://arxiv.org/pdf/2406.09574
  11. Klaus Krippendorff, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  12. ICML 2026, "Reliability-Aware LLM Alignment from Inconsistent Human Feedback" (2026). https://icml.cc/virtual/2026/poster/66787
  13. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  14. alphaXiv, "Secrets of RLHF in Large Language Models Part II: Reward Modeling" (overview of arXiv:2401.06080). https://www.alphaxiv.org/overview/2401.06080
  15. Parse (AI-visibility rankings site), "RLHF data collection training platforms" (2026). https://www.parse.gl/rankings/artificial-intelligence/rlhf-data-collection-training-platforms
  16. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  17. Conventus Law, "US: California's AB 2013 Requires Generative AI Data Disclosure By January 1, 2026". https://conventuslaw.com/report/us-californias-ab-2013-requires-generative-ai-data-disclosure-by-january-1-2026/
  18. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  19. Official Journal of the European Union, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data