Skip to content

Data quality, coverage and contamination

Preference Data Quality: Measuring Noise, Agreement and Ambiguity in Pairwise Comparisons

Quick answer

Preference data quality comes down to three numbers and one judgment: how often independent raters agree on the same pair, how many pairs are wrong versus genuinely ambiguous, and whether disagreement is random or tied to particular raters or topics. Published work reports agreement on chat preference pairs of roughly 63-72% [1], so a dataset quoting 95% agreement deserves scrutiny. Measure each, then decide per pair whether to keep, relabel, down-weight or drop it before reward modeling or DPO.

By SourceX Editorial · Updated

Why preference noise differs from classification label noise

Preference noise differs because a pairwise label can be wrong, or there can be no single right answer. In classification, a cat mislabeled as a dog is an error with a ground truth. In a comparison between two plausible assistant replies, raters may split because the responses trade off helpfulness against brevity, or because they hold different values. The Secrets of RLHF Part II work separates these into incorrect preferences and ambiguous preferences, and treats them differently in training [1].

That distinction matters for the objective. A Bradley-Terry reward model or the DPO loss [6] treats every chosen/rejected pair as a confident signal. Feed it near-tie pairs labeled as hard wins and it learns spurious margins; feed it flipped pairs and it learns the opposite of the intended preference. Curation studies of preference-optimization datasets find that what goes into the pair set materially shapes DPO results [9][10].

For a general treatment of label error budgets, see how much label noise is acceptable. The rest of this page is specific to comparison data.

Measuring annotator agreement on pairwise comparisons

Measure agreement on a multiply-annotated overlap sample, not on the whole set, and report a chance-corrected statistic alongside raw agreement. With two options and no tie, raw agreement has a 50% chance floor, so 70% raw agreement is a modest signal. Cohen's kappa covers two raters; Krippendorff's alpha handles many raters, missing ratings and ordinal strength scales in one coefficient [7]. Metric choice should follow the annotation design rather than habit [8].

Practical rules for the overlap sample:

  • Size and draw. Draw the overlap set randomly across prompt categories and response sources, not from the easiest slice. Stratify by task type (coding, safety, open-ended writing) because agreement varies by task.
  • Rater count. Three to five independent judgments per pair let you see split votes; two judgments cannot distinguish a 2-1 split from a coin flip.
  • Order control. Randomize left/right position and log it. A position-bias check (share of "A" wins when A and B are swapped) catches raters who default to one side.
  • Strength-aware statistics. If labels use a graded scale, compute ordinal alpha on the scale and binary agreement on the collapsed choice. They answer different questions.

For metric selection in depth, see inter-annotator agreement metrics for dataset buyers. If the supplier reports only a single "accuracy" figure against an internal gold set, ask how the gold labels were produced, since gold sets for preferences inherit the same ambiguity.

Separating incorrect pairs from ambiguous pairs

Separate the two by combining human vote spread with model-estimated preference strength. A pair where three raters unanimously chose B but an ensemble of reward models confidently prefers A is a candidate mislabel for expert review. A pair where raters split 2-1 and the ensemble margin is near zero is likely ambiguous, and no relabeling will make it clean.

The Secrets of RLHF Part II approach trains several reward models, scores each pair, and uses the mean and variance of the predicted margin as a preference-strength estimate [1]. A community summary of that work reports that, on Anthropic's HH-RLHF data, roughly a quarter of pairs looked mislabeled and roughly another quarter looked ambiguous under this scoring [2]. Treat those shares as a secondary report to verify against the paper, and treat any ensemble score as a triage signal, since the ensemble was trained on the same noisy labels.

A workable triage, applied per pair:

  • Strong agreement, positive margin: keep at full weight.
  • Strong human agreement, strongly negative margin: send to expert relabel; if confirmed flipped, swap chosen and rejected.
  • Split votes, near-zero margin: treat as ambiguous; down-weight, convert to a tie, or drop.
  • Low ensemble variance, near-zero margin, unanimous raters: check for near-duplicate responses; the pair may carry no learning signal.

Duplicate responses are common in sampled pairs. A MinHash and LSH near-duplicate pass on chosen versus rejected text removes pairs where the two sides differ only in whitespace or a closing sentence.

Systematic disagreement: rater, topic and value effects

Disagreement is often systematic, so check whether it clusters by rater, topic or demographic before averaging it away. A news digest of a recent preprint reports that filtering out self-inconsistent raters flipped the majority harm label on 18.6% of prompts [4]. The figure comes from a preprint and should be verified, but the mechanism is general: a few unreliable raters can swing majority votes on contested items.

Reliability-aware alignment methods estimate each rater's reliability from inconsistent feedback and weight their labels accordingly during training [3]. Buyers can apply a simpler version during acceptance:

  1. Self-consistency. Re-serve 3-5% of items to the same rater days later, with sides swapped. A rater who flips more than a set threshold on repeats is a weak signal source.
  2. Rater-vs-consensus agreement. Compute each rater's agreement with the leave-one-out majority. Outliers may be careless or may represent a real minority view; read their rationales before discarding.
  3. Slice analysis. Break agreement down by prompt category, language and safety tag. A dataset can average 70% agreement while political or medical prompts sit near chance.
  4. Value disagreement flag. For contested topics, keep the vote distribution rather than collapsing it. Distributional or soft-label training can use it; a hard majority label discards it.

Provenance fields make this analysis possible. See provenance for human-annotated and preference data for the rater-ID, guideline-version and AI-assistance disclosures to request.

Tie options, strength scales and rationales

Collection design determines how much of the noise you can later diagnose. A forced binary choice with no tie converts genuine ambiguity into apparent signal, and you cannot recover it afterward. Llama 2's preference collection paired the binary choice with a graded strength rating, so near-ties could be identified and handled separately in reward modeling [5].

Acceptance asks for any pairwise dataset (a working hypothesis, not a published standard):

  • Tie or "both bad" option offered, with its usage rate reported. A tie rate near zero on open-ended chat suggests raters were discouraged from using it.
  • Strength scale (for example, slightly / moderately / significantly better), stored per judgment, not only aggregated.
  • Free-text rationale captured on at least a sample, so you can audit why ambiguous pairs split.
  • Repeat-item self-consistency per rater, reported as a distribution across the rater pool.
  • Guideline version attached to each judgment, since mid-project guideline changes create discontinuities.

How the format itself trades off cost and signal is covered in pairwise, ranking or rating preference formats.

A preference pair record built for auditing

An auditable record carries the judgments, not just the final chosen/rejected fields. Most public DPO datasets ship as prompt, chosen, rejected; that shape is sufficient for training and insufficient for quality review.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "pair_id": "pp-000412",
  "prompt_id": "pr-1187",
  "prompt_category": "customer_support/refund_policy",
  "response_a": {"text": "...", "source": "model_v3_t0.7"},
  "response_b": {"text": "...", "source": "human_agent_edit"},
  "position_shown": {"r-17": "A_left", "r-22": "B_left", "r-41": "A_left"},
  "judgments": [
    {"rater_id": "r-17", "choice": "B", "strength": 2, "rationale": "B cites the 30-day window", "guideline_ver": "1.4", "ts": "2026-08-02T14:10:00Z"},
    {"rater_id": "r-22", "choice": "B", "strength": 1, "rationale": null, "guideline_ver": "1.4", "ts": "2026-08-02T15:02:00Z"},
    {"rater_id": "r-41", "choice": "tie", "strength": 0, "rationale": "both accurate, A shorter", "guideline_ver": "1.4", "ts": "2026-08-03T09:41:00Z"}
  ],
  "aggregate": {"majority": "B", "vote_split": "2-0-1", "mean_strength": 1.0},
  "audit": {"rm_ensemble_margin_mean": 0.31, "rm_ensemble_margin_std": 0.44, "flag": "ambiguous_review"}
}

The vote_split, strength and rm_ensemble_margin_std fields drive the triage above. Without per-rater IDs, self-consistency and rater weighting are impossible.

Clean, reweight or reject: a decision table

Choose the remediation by error type and by how much of the set it touches. Relabeling is expensive and only fixes incorrect pairs; reweighting suits ambiguity; rejection is right when a whole slice or rater cohort is unreliable.

Illustrative example: invented to show structure; it does not describe an available dataset.

Signal observedLikely causeActionTraining-time handling
Unanimous raters, ensemble strongly disagreesFlipped or mis-keyed labelExpert relabel a sample; swap if confirmedFull weight after fix
Split votes, near-zero marginGenuine ambiguityKeep vote distributionSoft label or reduced weight; or drop for DPO
One rater cohort far from consensusUnreliable raters or minority valuesRead rationales, then filter or keep as sliceReliability weighting [3]
Agreement near chance in one categoryUnderspecified guidelinesReject slice or request recollectionExclude until fixed
High "A" win rate regardless of contentPosition biasCheck randomization logsDrop affected rater batches
Chosen and rejected near-identicalLow-information pairDeduplicateDrop

Record what you removed and why in the acceptance report, using the structure in what a dataset quality report should contain. ISO/IEC 5259-1 gives shared terminology for writing those findings up across the data life cycle [11].

What to request from a preference data supplier

Request the raw judgments and the measurement method, not a headline agreement number. A supplier who can provide per-rater, per-judgment records lets you run every check above; one who can provide only majority labels leaves you trusting their cleaning.

The minimum request: overlap-sample size and selection method, raw and chance-corrected agreement by category, tie and strength-scale distributions, rater pool size and self-consistency rates, guideline versions, how response pairs were sampled, and any automated filtering already applied. For buying decisions beyond quality, see buying RLHF comparison data and the data quality cluster hub. Background definitions are at preference data, what is RLHF data, what is human-feedback data and training data for reward models.

Some teams pair purchased comparisons with real operational records, such as support conversations with agent edits, that implicitly encode preferred responses. If that is your plan, you can describe the records you need to SourceX, which sources operational datasets from US companies on request.

Sourcing operational records for preference training

SourceX sources operational datasets from US companies on request, including support and sales histories and new recordings of hands-on work, and manages licensing through to ongoing purchases. Every dataset is rights-reviewed, delivered under a license that defines records, uses, term and delivery, and comes with per-dataset diligence materials; a request does not guarantee a match. Tell SourceX what data you need.

Frequently asked questions

What agreement rate should a preference dataset have?

There is no universal threshold. Published chat preference work reports roughly 63-72% inter-annotator agreement [1], so much higher figures on open-ended prompts warrant a question about how agreement was measured. Compare by category and against a chance-corrected statistic.

Should ambiguous pairs be dropped for DPO?

Often, or down-weighted. DPO optimizes the margin between chosen and rejected directly [6], so a near-tie labeled as a clear win pushes the policy toward noise. Keep them if you train with soft labels or a tie-aware objective.

Can a reward-model ensemble replace human relabeling?

No. The ensemble learned from the same labels, so it is a triage filter that ranks pairs for review [1]. Confirm flips with expert relabeling on a sample before swapping at scale.

Sources

  1. alphaXiv, "Secrets of RLHF in Large Language Models Part II: Reward Modeling (overview, arXiv:2401.06080)" (2024). https://www.alphaxiv.org/overview/2401.06080
  2. Hugging Face (rl-llm-wiki knowledge base), "source: arxiv:2401.06080 - Secrets of RLHF Part II: Reward Modeling (community discussion)". https://huggingface.co/datasets/rl-llm-wiki/knowledge-base/discussions/160
  3. ICML 2026, "Reliability-Aware LLM Alignment from Inconsistent Human Feedback" (2026). https://icml.cc/virtual/2026/poster/66787
  4. AI Weekly, "Study: 18.6% of RLHF harm labels flip when raters filtered" (2026). https://aiweekly.co/alerts/study-186-of-rlhf-harm-labels-flip-when-raters-filtered
  5. Meta AI, "Llama 2: Open Foundation and Fine-Tuned Chat Models" (2023). https://arxiv.org/pdf/2307.09288
  6. Stanford University, "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  7. University of Pennsylvania, Annenberg School for Communication, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  8. arXiv, "Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation" (2026). https://arxiv.org/pdf/2603.06865
  9. arXiv, "When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets" (2025). https://arxiv.org/pdf/2511.10985
  10. arXiv, "What Matters in Data for DPO?" (2025). https://arxiv.org/pdf/2508.18312
  11. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-1:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 1: Overview, terminology, and examples" (2024). https://www.iso.org/standard/81088.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data