Fine-tuning and post-training data
Buying RLHF comparison data: collection services, expert networks or licensed judgments
Quick answer
RLHF data collection services put trained raters in front of responses sampled from your checkpoint. Commissioned raters from a managed service, a dedicated team or an expert network (for comparisons that turn on a domain fact) supply on-policy judgments of your current model's outputs. Licensed judgments, such as supervisor approvals, edits and QA scores recorded in business systems, are off-policy but can serve as reward-model warm starts, rater golds and evaluation sets. The routes can be combined, and the statement of work decides quality more than the vendor's name.
By SourceX Editorial · Updated
For the concept itself, see what RLHF data is and the RLHF glossary entry.
Four ways to buy human comparisons, and what each delivers
The four procurement models differ in who writes the responses, who judges them, how you pay, and whether the comparisons describe your current model.
The judgment is the product in every model. InstructGPT trained a reward model on human rankings of model outputs before reinforcement learning, using contractors hired through Upwork or sourced from Scale AI and selected on screening criteria [1]. Direct Preference Optimization (DPO) fits the policy to chosen and rejected pairs with no separate reward model [2], so a mislabeled pair shapes the policy directly.
| Model | Responses from | Who judges | Common pricing unit | On-policy? | Fits | Main risk |
|---|---|---|---|---|---|---|
| Managed collection service | Your checkpoint | Vendor-recruited raters trained on your guidelines | Per comparison | Yes, with fresh samples each round | High-volume helpfulness and harmlessness rounds | Rater churn and guideline drift between batches |
| Dedicated managed team | Your checkpoint | A named, stable team | Per reviewer hour or managed SLA | Yes | Unreleased models, multi-turn and agentic tasks | Paying for idle capacity; know-how stays with the vendor |
| Expert network or credentialed pool | Your checkpoint | Practitioners recruited per specialty | Per hour or per task at expert rates | Yes | Comparisons decided by clinical, legal, tax or security facts | Thin supply per specialty; credential checks; employer conflicts |
| Licensed existing judgments | Employees, customers or earlier systems | Supervisors, compliance reviewers and QA graders doing their jobs | License terms agreed per deal | No | Warm starts, rater golds, evaluation sets, rubric seeds | Business criteria behind the choice; personal data; rights chain |
As market practice, Sama describes managed teams producing preference rankings under detailed guidelines with multi-layer QA review [3]. Digital Divide Data describes programs spanning rater recruitment, calibration, rubric design, agreement measurement and adjudication, priced per comparison, per reviewer hour or under a managed SLA [4]. Some providers also sell ready-made preference datasets [5]; check those like open sets, as in open preference datasets that allow commercial fine-tuning.
You can also run raters yourself on a tool such as Label Studio, whose pairwise template shows a prompt with two candidate responses [6]. Prompts and outputs stay in-house, but you recruit, contract, pay and check every rater.
Matching the procurement model to your post-training run
Four questions settle the choice: must the comparisons describe your current checkpoint, does a domain fact decide them, does the signal already exist in recorded decisions, and how confidential are the prompts and outputs?
- Current checkpoint. On-policy data comes from the policy being trained, and samples from an earlier policy count as off-policy for a later one [7]. Each training round therefore needs fresh comparisons, which a service or team can supply and an archive cannot. The value depends on the task: SimpleMix found on-policy data most effective on objective tasks such as math and coding, off-policy data on open-ended tasks, and a mix better than either [8]. See on-policy vs off-policy preference data.
- Domain facts. Generalists cannot judge a drug interaction or a filing deadline. Route that slice to experts; see domain-expert preference data and contracting domain experts.
- Recorded decisions. If reviewers in a business already choose between drafts, versions or proposals, licensing those records avoids recreating the judgment.
- Confidentiality. Outputs from an unreleased model reveal its capabilities, so prefer a dedicated team or your own raters and ask how many people will see them.
Lee et al. found that training on AI-generated preference labels performed comparably to training on human labels for summarization and helpful dialogue, and better for harmless dialogue [9], so consider reserving paid human work for slices where AI judges prove unreliable. See AI feedback vs human preference data and what drives the cost of human preference data.
What licensed judgments can replace, and what they cannot
Licensed judgments replace commissioned comparisons only when you need a record of how qualified people chose between real alternatives, not a verdict on your own model's outputs.
A QA scorecard grades a support reply item by item, a supervisor's edit pairs a draft with the approved final, a compliance review returns one version and approves another, and code review separates approved changes from requested ones. These records support four uses:
- Warm starts. One bandit-learning study shows that an offline preference dataset from an expert of unknown competence can warm-start online learning when the method models that competence [10].
- Golds for vendor raters. A reviewer's recorded decision is a known answer your vendor did not write.
- Evaluation. Held-out reviewer decisions can check a reward model or an LLM judge; see LLM-as-a-judge calibration sets.
- Rubric seeds. Reason codes and scorecard items show the criteria practitioners already apply.
The limits: the responses did not come from your model, a rejection can reflect price or negotiation rather than quality, records often keep only the approved final, and both sides can contain personal data. Apply the conversion rules in preference datasets for DPO before any record becomes a training pair.
SourceX works on this licensed route. It sources operational datasets from US companies, including support and sales histories, documents, and finance and legal workflows, on request rather than from stock, so a request does not guarantee a match. Every dataset goes through rights review and is delivered under a license that defines the included records and permitted uses.
Names, emails, phone numbers and account numbers are removed or replaced before delivery, and no de-identification method is perfect. To scope review records as preference data, describe the decisions and fields you need, or see QA scores and corrections.
Statement of work for RLHF data collection: clauses that decide quality
An RLHF statement of work (SOW) should fix response generation, the rater pool, quality measurement, acceptance, deliverables, confidentiality and ownership before the pilot, because each is expensive to renegotiate once raters are trained. The general structure is in writing a data collection statement of work; these rows are specific to comparisons.
| SOW section | What to specify | Failure it prevents |
|---|---|---|
| Response generation | Who samples (you, or the vendor through an endpoint you control), checkpoint IDs, sampling settings, pairwise or k-way ranking | Comparisons of a stale or unrecorded checkpoint |
| Prompts | Source, ownership, task mix, whether raters may write prompts; see real-world prompt sets | Off-distribution or benchmark-copied prompts |
| Guidelines and rubric | Buyer-owned, versioned, change control, who pays to re-rate | Silent drift between batches |
| Judgment format | Strength scale, tie and "both unacceptable" options, rationales, per-criterion verdicts; see pairwise, ranking or rating formats | Forced choices between near-identical responses |
| Rater pool | Screening test and pass mark, locales, credentials for flagged slices, share of pilot raters kept in production | A senior team runs the pilot and new hires run production |
| Golds and calibration | Qualification set, gold share in live work, golds written by you, removal rules; see rater calibration | Accuracy never measured independently |
| Overlap and agreement | Share double-rated, statistic per criterion, threshold fixed after the pilot | Agreement reported only on the easy overall label |
| Throughput | Batch size, weekly capacity, turnaround, surge notice | Labels arriving after their training round |
| Acceptance and rework | Acceptance sample per batch, rejection triggers (gold accuracy, timing, position bias), rework at vendor cost, payment per accepted comparison; see acceptance sampling | Paying for noise |
| Deliverables | JSONL per comparison with every rater's label, pseudonymous rater ID, rubric version, seconds on task, adjudication log | Majority votes with no audit trail |
| AI assistance | Whether raters may use chat assistants; disclosure and detection; see detecting model-generated content | Model-written "human" rationales |
| Confidentiality | NDAs flowing to every rater, access-controlled tooling, no local copies, rater locations, deletion at term end | Leaks of unreleased outputs |
| Subcontracting | Named sub-vendors and crowd platforms, flow-down of terms, approval before changes | An unknown labor chain |
| Ownership | Assignment of labels, rationales and rewrites; vendor tools carved out; warranty that each rater signed an assignment | Gaps in the rights chain |
Three rows need detail. For agreement, Krippendorff's alpha handles any number of raters and items not every rater saw; 1 means perfect reliability and 0 means chance [11]. For deliverables, require every rater's label: standard DPO treats high-disagreement pairs like unanimous ones [12], and you cannot reweight labels you never received.
For ownership, written judgments such as rationales can be protected expression. On 29 September 2026 the Third Circuit held the Westlaw headnotes at issue in Thomson Reuters v. ROSS copyrightable, and ROSS's use of them to train a non-generative legal-research tool not fair use [13]. See IP assignment vs license for commissioned data.
Running the pilot: one batch through every shortlisted service
Send the same pilot batch to each shortlisted service with your own known-answer items mixed in. Decide on agreement, accuracy against your golds and model gain, not on the vendor's QA report.
Summaries of the Secrets of RLHF Part II study report typical inter-annotator agreement of 63-72% on general preference labels [14]. A vendor reporting much higher agreement on subjective comparisons should show how it counted ties, easy pairs and adjudicated items.
Illustrative example: invented to show structure; it does not describe an available dataset.
A team tuning a support assistant samples 400 de-identified production prompts held out from supervised fine-tuning and draws two responses per prompt from its current checkpoint, randomizing order for every rater. Both vendors rate all 400 pairs, double-rate 100 and receive 30 internally adjudicated golds inserted blind. With decision rules fixed in advance, it trains a reward model on each vendor's labels and scores it on 200 separate adjudicated pairs.
| Pilot measure | Vendor A | Vendor B | Reading |
|---|---|---|---|
| Gold accuracy | 27 of 30 | 21 of 30 | B misses correctness golds |
| Alpha, correctness criterion | 0.71 | 0.48 | B's raters disagree on facts |
| First-position response chosen | 51% | 58% | B shows position bias |
| Median seconds per comparison | 95 | 31 | B's pace is too fast to read both responses |
| Tie option used | 12% | Not offered | B's tool forced choices |
| Reward-model accuracy, internal set | 68% | 61% | Model gain follows label quality |
The team signs with Vendor A, writes the tie option and a minimum time on task into the SOW, and repeats the gold check on every batch. See paid data pilot terms and estimating a dataset's value before purchase.
Vendor claims to verify before you scale
Workforce size, "expert" coverage and multi-layer QA are self-reported, so ask for evidence you can check in your own pilot data.
- "Thousands of vetted raters." How many worked on your pilot, how many will work in production, and what share stays between batches?
- "Domain experts." Credential type, jurisdiction, verification method and date per pseudonymous rater; see verifying expert annotator qualifications.
- "Multi-layer QA." Rejection rates per layer, and both original and reviewed labels when reviewers change them.
- "No AI use." The written policy, the detection method and past incidents.
- Rankings. At least one published ranking of RLHF data collection platforms measures how often brands appear in AI assistant answers, not data quality [15].
In production, watch for falling time on task, changing rater IDs, slipping gold accuracy, near-identical rationales and strength labels collapsing to one value. See measuring noise in preference data and provenance for human-annotated data.
Records to require for your own disclosures
Require the service or licensor to deliver the facts you will need to document the data later: where prompts and responses came from, who judged them under which guidelines, and whether any part was synthetic.
- California AB 2013. As of October 2026, developers of generative AI systems made publicly available to Californians must post training-data documentation with a high-level summary of the datasets. The first deadline was 1 January 2026, and the duty repeats before each later release or substantial modification, which the statute defines to include the results of retraining or fine-tuning [16]. Summaries list items such as data sources and ownership, collection and processing, and synthetic data use [17].
- EU AI Act Article 10. For high-risk systems, Article 10(2) requires governance of data preparation, including annotation and labelling [18]. As of October 2026, Regulation (EU) 2026/1744 is in the Official Journal, and secondary sources report that Annex III high-risk obligations now apply from 2 December 2027 ; the Regulation also amends Article 10 [19].
See fine-tuning provider duties under the EU AI Act and AB 2013, how to procure enterprise training data and the fine-tuning and post-training data hub.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Considering licensed judgments alongside a collection service?
Describe the decisions you want a model to learn from, such as approvals, edits, QA scores or rejections with reasons, along with the fields, volume and uses the license must cover. SourceX looks for US businesses that hold those records, checks the data and the supplier's licensing permissions, and manages the license and delivery. Nothing is contracted until a supplier agrees. Describe the preference judgments you need.
Sources
- Ouyang et al. (OpenAI), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
- Rafailov et al. (Stanford), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
- Sama (vendor page), "RLHF Training Data for LLM Alignment". https://info.sama.com/high-quality-human-feedback-for-large-language-models
- Digital Divide Data (vendor page), "RLHF services and pricing". https://www.digitaldividedata.com/?p=24655
- Argos Multilingual (vendor page), "Reinforcement Learning from Human Feedback (RLHF)". https://data.argosmultilingual.com/llm-training-data/reinforcement-learning-from-human-feedback-rlhf/
- Label Studio (product documentation), "Human Preferences collection for RLHF". https://labelstud.io/templates/generative-pairwise-human-preference
- arXiv, "Towards a Unified View of Preference Learning for Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2409.02795
- Proceedings of Machine Learning Research (ICML 2025), "SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning" (2025). https://proceedings.mlr.press/v267/li25au.html
- Lee et al. (ICML 2024), "RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback" (2023). https://arxiv.org/pdf/2309.00267
- arXiv, "Online Bandit Learning with Offline Preference Data for Improved RLHF" (2024). https://arxiv.org/pdf/2406.09574
- Klaus Krippendorff, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- ICML 2026, "Reliability-Aware LLM Alignment from Inconsistent Human Feedback" (2026). https://icml.cc/virtual/2026/poster/66787
- U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
- alphaXiv, "Secrets of RLHF in Large Language Models Part II: Reward Modeling" (overview of arXiv:2401.06080). https://www.alphaxiv.org/overview/2401.06080
- Parse (AI-visibility rankings site), "RLHF data collection training platforms" (2026). https://www.parse.gl/rankings/artificial-intelligence/rlhf-data-collection-training-platforms
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- Conventus Law, "US: California's AB 2013 Requires Generative AI Data Disclosure By January 1, 2026". https://conventuslaw.com/report/us-californias-ab-2013-requires-generative-ai-data-disclosure-by-january-1-2026/
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- Official Journal of the European Union, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.