Skip to content

Fine-tuning and post-training data

Domain-expert preference data for specialized models

Quick answer

Expert preference data for RLHF is a set of comparisons in which credentialed practitioners, such as attorneys, clinicians, medical coders or accountants, judge which model response is better and record why. Buy it for the comparisons generalist raters cannot judge: those decided by a domain fact, a jurisdiction's rules or a professional standard of care. Specify credentials per prompt type, a rubric that ranks substantive errors above style, double rating with adjudication, and records that keep rationales, ties, "both unacceptable" flags and rater metadata.

By SourceX Editorial · Updated

Where expert judgment changes the preference label

Experts earn their cost on comparisons decided by a domain fact or professional standard that a generalist cannot check, not on comparisons decided by clarity, format or tone. The preference label is the training signal at this stage: InstructGPT-style RLHF learns from human rankings of model outputs [1], and Direct Preference Optimization (DPO) fits the policy directly to chosen and rejected responses without a separate reward model [2]. A confident but wrong preference teaches the model the wrong behavior, and nothing downstream checks the rater's domain knowledge.

Who rates changes the label even on general data. A news summary of one preprint reports that dropping the raters least consistent with themselves flipped the majority harm label for 18.6% of prompts [3]. The effect is likely sharper in specialized domains, because a fluent but wrong answer reads as the better one to a non-specialist. Route each comparison by what decides it:

Comparison typeWhat decides itWho should judge
Format, length or tone of two correct answersReadability for the target userTrained generalists or AI feedback, spot-checked by an expert
A domain fact (drug interaction, filing deadline, tax treatment)Whether a statement is true under current rulesCredentialed practitioner in that specialty
Jurisdiction-dependent advice (state law, payer policy)Which rule applies where the user isPractitioner licensed or admitted in that jurisdiction
Escalation and risk (see a physician, consult counsel)The professional standard of careSenior practitioner, always double-rated
Both responses wrongNothing worth preferringExpert flags it; the item becomes a rewrite task
Generalist and expert labels disagreeUnknown until checkedExpert adjudicates; update the routing rules

A tiered design controls cost: generalists or AI feedback cover the full set, and experts take the slice flagged by domain-critical criteria or disagreement. See AI feedback vs human preference data, what drives the cost of human preference data and, on commissioning versus licensing, buying RLHF comparison data; the fine-tuning and post-training data hub maps the rest.

What an expert comparison record should contain

A usable expert preference record holds the comparison, per-criterion verdicts, an overall preference with strength, a written rationale and pseudonymous rater metadata, so you can audit credentials and agreement without seeing identities. Keep the record richer than the training format: the chosen/rejected pairs described in preference datasets for DPO are an export, not the source of truth.

The fields that matter most for expert preference data:

  • Strength, ties and "both unacceptable". Forcing a winner between two near-identical responses adds noise. A pair of two wrong answers teaches the model which wrong answer to prefer; route it to an expert rewrite as expert demonstration data for SFT.
  • Per-criterion verdicts. They show why the preference went one way and let you retrain on correctness alone if style judgments prove unreliable.
  • Response provenance. Record the checkpoint and sampling settings behind each response; pairs from your own policy and pairs from unrelated models do different jobs (on-policy vs off-policy preference data).
  • Rater attributes, not identities. Credential type, specialty, jurisdiction, verification method and date, a conflict-of-interest attestation and an AI-assistance disclosure.
  • Rubric version and time on task, so relabeling after a rubric change is traceable and implausibly fast judgments stand out.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "pair_id": "anticoag-000412",
  "rubric_version": "med-pref-2.3",
  "prompt": {"text": "I take warfarin. My doctor prescribed a 10-day course of fluconazole. Can I start it today?", "source": "expert_written", "jurisdiction": "US"},
  "responses": [
    {"id": "A", "checkpoint": "policy-sft-0917", "temperature": 0.7},
    {"id": "B", "checkpoint": "policy-sft-0917", "temperature": 0.7}
  ],
  "criteria": {
    "correctness": {"A": "fail", "B": "pass", "note": "A states there is no interaction"},
    "risk_escalation": {"A": "fail", "B": "pass"},
    "completeness": {"A": "partial", "B": "pass"},
    "compliance_scope": {"A": "pass", "B": "pass"},
    "clarity": {"A": "pass", "B": "partial"}
  },
  "overall": {"preferred": "B", "strength": "much_better", "tie": false, "both_unacceptable": false},
  "rationale": "B flags the warfarin-azole interaction and tells the patient to contact the prescriber about INR monitoring before starting; A misses it.",
  "rater": {"rater_id": "R-0417", "credential": "PharmD", "specialty": "anticoagulation", "jurisdiction": "US-OH",
            "verified": {"method": "state board lookup", "date": "2026-08-14"}, "coi": "none declared", "ai_assistance": "none"},
  "seconds_on_task": 388,
  "adjudication": {"status": "agreed", "second_rater": "R-0291", "adjudicator": null}
}

Matching credentials to prompts and screening for conflicts

Verify the specific credential each prompt category needs, in the jurisdiction the prompt assumes, and store how and when it was checked; a generic "domain expert" tag says nothing about whether the rater could see the error. Map prompt categories to credentials before recruiting: Medicare claim edits to certified coders, anticoagulation questions to pharmacists or physicians, US GAAP revenue recognition to CPAs.

  • Verification. Use state licensing-board lookups, the NPPES registry of National Provider Identifiers for US clinicians, and state bar directories for attorneys. Record method and date per rater and re-verify periodically, because licenses lapse.
  • Skill, not just status. Credentials do not prove skill on your task. Annotation platforms separate qualification tasks, which test raters on known-answer items before work starts, from honeypots, which are known-answer items mixed into live work; one platform's guidance keeps honeypots to around 3-5% of items [4]. For expert work, senior practitioners should write the known-answer comparisons. See verifying domain-expert annotators.
  • Conflicts of interest. Exclude raters who wrote one of the responses or the reference answer, who work for a competitor or for the organization whose documents seeded the prompts, or who have financial ties to products named in the content.
  • Confidentiality and AI assistance. Attorneys and clinicians must not import client or patient facts into prompts or rationales. Raters who draft rationales with a chat assistant should disclose it (provenance for human-annotated and preference data).

Engagement models are covered in contracting domain experts for post-training data; for evaluation graders, see sourcing domain-expert raters for LLM evaluation.

A domain rubric that ranks errors before style

Give experts a short rubric with a fixed priority order, so one correctness or safety failure decides the comparison before completeness, compliance or tone are weighed. Without an order, raters trade a dangerous omission against better formatting, and agreement collapses.

CriterionWhat the expert checksLegal failureMedical failureEffect on overall preference
CorrectnessEach substantive statement is true under current rulesRelies on a repealed statute sectionGives an adult dose for a pediatric patientVeto: an incorrect response cannot beat a correct one
Risk and escalationFollowing the advice could cause harm; escalation appears where the standard requires itSays a court deadline can be ignoredMisses red-flag symptoms that need urgent careVeto
CompletenessMaterial issues a practitioner would raiseOmits the limitation periodOmits a contraindicationDecides between responses that pass both vetoes
Compliance and scopeJurisdiction stated, scope of practice respected, no identifiable personal dataGives state-specific advice without naming the stateRepeats a patient's name and date of birthDecides, or vetoes where the rule is mandatory
Clarity and toneFit for the intended readerLegalese to a consumerUnexplained jargon to a patientTiebreaker only

Attach two or three anchored comparisons per criterion and version the rubric. Eliciting criteria is covered in designing evaluation rubrics with domain experts; to train on rubric scores instead of pairs, see rubrics as rewards.

Adjudicating expert disagreement without erasing it

Double-rate a fixed share of items, measure agreement per criterion, and send every disagreement on a veto criterion to a senior adjudicator who decides whether it is a rater error, a rubric gap or legitimate variation in practice. General preference data is noisy: summaries of the Secrets of RLHF Part II paper report typical inter-annotator agreement of 63-72% on preference labels [5][6].

Some disagreement is information. One 2025 paper argues that it can reflect real differences in preference rather than error [7], and a 2026 ICML paper estimates each annotator's reliability to infer latent labels, noting that standard DPO treats high-disagreement pairs like unanimous ones [8]. Ask for every rater's label and the adjudication log, not only the majority.

Measure agreement with Krippendorff's alpha, which handles any number of raters and items that not every rater saw; 1 means perfect reliability and 0 means agreement no better than chance [9]. As market practice, one annotation-quality guide treats alpha of at least 0.800 as reliable and 0.667-0.800 as tentative [10]. Set thresholds per criterion, since subjective criteria such as tone usually agree less than veto criteria. Then act on the adjudication outcome:

  • Rater error: correct the label and count it against that rater's accuracy.
  • Rubric gap: amend the criterion, bump the rubric version and re-rate the affected items.
  • Practice variation: mark the pair contested, then drop it from DPO, downweight it, or keep it for a reward model that represents uncertainty.

Statistic choice is compared in inter-annotator agreement for dataset buyers, noise diagnostics in measuring noise and ambiguity in preference data, and resolution workflows in annotation adjudication and disagreement resolution.

Using professional review records to seed and check expert raters

Records of senior professionals revising, approving or rejecting work are expert judgments made under real accountability, most useful as known-answer items, rubric seeds and a check on your raters. A coding audit pairs the original claim code with the auditor's correction and reason code; a contract redline history shows which drafting a supervising attorney accepted; a QA scorecard grades a reply item by item.

  • Known-answer items. An auditor's correction with a recorded reason makes a qualification or honeypot comparison that your raters did not write.
  • Rubric seeds. Reason codes and scorecard items reveal the criteria practitioners already apply.
  • Rater calibration. Compare your raters' preferences with the recorded reviewer decision; raters who systematically disagree with practicing reviewers need retraining or removal.
  • Direct pairs. Some convert into training pairs; see the rules in preference datasets for DPO.

These records are human data with limits. They are off-policy, the reviewer's credential is often inferable only from role, and they carry client and patient details. Health records need HIPAA de-identification by Safe Harbor, which removes 18 listed identifiers and requires no actual knowledge that the rest could identify someone, or by Expert Determination [11]; compare the two in HIPAA Safe Harbor vs Expert Determination.

SourceX sources operational datasets from US companies, including support histories, documents, and finance and legal workflows, on request rather than from stock. Each dataset goes through rights review and is delivered under a license that defines the included records and their permitted uses; personal details are removed or replaced before delivery, and no de-identification method is perfect. For health records, SourceX requires HIPAA de-identification by Safe Harbor or Expert Determination before anything is considered for a license. If review histories would strengthen your rater program, describe the records and uses you need, and see expert annotations and labels and QA scores and corrections.

Rights, confidentiality and documentation specific to expert judgments

Expert judgments raise three questions that crowd labels rarely do: who owns the written analysis, whether employers or professional duties limit what raters contribute, and how credentials are stored.

  • Written analysis. Expert annotation can be protected expression. On 29 September 2026 the Third Circuit held that the Westlaw headnotes at issue were copyrightable and that ROSS's use of Westlaw material to train a non-generative legal-research tool was not fair use [12]. The case concerned a non-generative tool, and reports of a further appeal were unverified as of October 2026. Obtain an assignment or license of judgments and rationales from each rater.
  • Employers and professional duties. Clinicians and associates who rate in their own time may be bound by employer IP or outside-work policies, and by confidentiality duties to clients and patients.
  • Credential data. License numbers and credential documents are personal data; deliver pseudonymous IDs with attributes.

Dataset documentation standards now cover annotators, and some are machine-readable. Data Statements (schema version 2) include an annotator demographic field [13]; Croissant-RAI makes responsible-AI metadata machine-readable, with data labeling among its use cases [14]; and ISO/IEC 5259-4 covers labelling in its data quality process framework [15]. If the model will sit in a high-risk system under the EU AI Act, Article 10(2) requires governance of data preparation, including annotation and labelling [16]. As of October 2026, secondary sources report that Regulation (EU) 2026/1744 moved high-risk application dates to 2 December 2027 (Annex III) and 2 August 2028 (Annex I) and also amended Article 10's data-governance requirements [17].

Red flags in an expert preference sample

Weak expert preference data usually shows in rater metadata, overlap and timing before you read a label. Renegotiate or reject when you see:

  • "Verified experts" with no credential type, jurisdiction, verification method or date per rater.
  • Agreement given only as raw percent on the overall preference, or no double-rated items at all.
  • Rationales that restate the rubric, or near-identical wording across raters.
  • Median seconds per comparison too low to read both responses at their length.
  • Position bias: response A wins far more often than randomized ordering predicts.
  • Responses from one older model unrelated to your policy, or prompts copied from public benchmarks (decontaminating training data against benchmarks).
  • Raters licensed in a different jurisdiction from your users.
  • No answer on who owns the rationales.

Specifying the rater pool and adjudication in your request

Fixing credentials, rubric, overlap and rights up front lets suppliers quote comparable work and lets you accept a delivery against written criteria. Add these expert-specific fields to your preference data request:

FieldWhat to specify
Credential matrixRequired credential per prompt category, specialty, jurisdictions, minimum years in practice
VerificationMethod (licensing board, NPI registry, bar directory), date per rater, re-verification interval
QualificationKnown-answer comparisons by senior practitioners, pass mark, honeypot share
Conflicts and assistanceExclusion rules, attestation wording, AI-assistance policy
Judgment formatPairwise or ranking, strength scale, ties, both-unacceptable flag (pairwise, ranking or rating formats)
RubricCriteria, priority order and vetoes, anchors, version control
Overlap and adjudicationShare double-rated, adjudicator seniority, outcome codes, alpha threshold per criterion
DeliverablesJSONL records, every rater's label, adjudication log, pseudonymous rater roster, data statement
RightsAssignment of judgments and rationales, employer consents, permitted uses (reward model, DPO, evaluation)

Describe the expert judgments your model needs

If professional review records could seed or supplement your expert preference program, describe the domain, record types, jurisdictions and permitted uses on the buyer page. SourceX looks for US companies that hold that data, checks their licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Start an expert preference data request.

Sources

  1. Ouyang et al. (OpenAI), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  2. Rafailov et al. (Stanford), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  3. AI Weekly, "Study: 18.6% of RLHF Harm Labels Flip When Raters Filtered" (news summary of a preprint). https://aiweekly.co/alerts/study-186-of-rlhf-harm-labels-flip-when-raters-filtered
  4. Dataloop, "Creating Consensus, Honeypot, and Qualification Tasks" (developer documentation). https://developers.dataloop.ai/tutorials/task_workflows/quality_control/chapter
  5. alphaXiv, "Secrets of RLHF in Large Language Models Part II: Reward Modeling" (overview of arXiv:2401.06080). https://www.alphaxiv.org/overview/2401.06080
  6. Hugging Face community (rl-llm-wiki), "source: arxiv:2401.06080 - Secrets of RLHF Part II: Reward Modeling". https://huggingface.co/datasets/rl-llm-wiki/knowledge-base/discussions/160
  7. arXiv, "Maximizing Signal in Human-Model Preference Alignment" (2025). https://arxiv.org/html/2503.04910v1
  8. ICML 2026, "Reliability-Aware LLM Alignment from Inconsistent Human Feedback" (2026). https://icml.cc/virtual/2026/poster/66787
  9. Klaus Krippendorff, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  10. Koji, "Data Annotation Quality: Guidelines, Agreement Metrics, and Gold Tasks That Actually Work" (2026). https://www.koji.so/docs/data-annotation-quality-guide
  11. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  12. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  13. University of Washington Tech Policy Lab, "Data Statements" (schema Version 2, 2021). https://techpolicylab.uw.edu/data-statements/
  14. Jain et al. (MLCommons), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  15. ISO/IEC, "ISO/IEC 5259-4:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
  16. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  17. Official Journal of the European Union, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data