Evaluation and benchmarking datasets
Human preference evaluation: sourcing side-by-side judgments for model comparison
Quick answer
Human preference evaluation compares two models by showing raters the same prompt with two blinded, order-randomized responses and recording which one they prefer, or a tie. To make the result decision-grade, buyers need three things sourced deliberately: a prompt set drawn from the real domain, a rater panel matched to the task with measured agreement, and enough paired comparisons to separate a true win-rate difference from noise. Keep these judgments separate from preference pairs used for training.
By SourceX Editorial · Updated
What a side-by-side evaluation actually measures
A side-by-side evaluation measures relative preference between two systems on a fixed prompt distribution, not absolute quality. The output is a win rate (wins, losses and ties for model B against model A) plus a confidence interval, and it is only as general as the prompts and raters behind it. Public crowdsourced leaderboards popularized this design: a user sees answers from two anonymous models, votes, and only then learns which models they were. Teams that judge models with LLMs face the same position effects humans do, which is why the protocol carries much of the rigor [1].
For a product team comparing a candidate release against production, a public leaderboard is a weak proxy. Start from your map of LLM evaluation datasets to decide where side-by-side judgments fit alongside private eval sets. Arena prompts reflect whoever visits the site, while your launch decision depends on your users' tasks, your policies and your failure costs. The rest of this guide treats the evaluation as a dataset you specify and buy: prompts, response pairs, judgments and rater metadata.
Prompt sourcing: why real domain prompts beat generic pools
Real domain prompts produce win rates that predict production behavior; generic prompt pools mostly measure general chat preference. A model that wins on open-ended trivia and creative writing can still lose on refund-policy questions, multi-step troubleshooting or contract clause questions, and a generic pool will not surface that. The practical hypothesis most teams adopt is that the closer the prompt set is to logged production traffic, the more the win rate transfers.
Good prompt sources include de-identified support tickets, sales and pre-sales emails, internal knowledge-base questions, engineering incident threads and document-review requests. Each prompt should carry slice tags (intent, difficulty, risk tier, channel, language) so you can report win rates per slice rather than one blended number. Building prompts from operational records is covered in more depth in building a golden evaluation dataset from real business records, and allocating prompts to rare and high-risk cases matters here too.
Protect the prompt set from contamination. If prompts leak into training data or into public benchmarks, later comparisons inflate. Hold the set out, version it, and rotate a share of prompts each cycle, following the patterns in contamination-resistant evaluation design.
Blinding, order randomization and interface design
Blinding and randomized presentation order are the two controls that make pairwise judgments trustworthy. Raters must never see model names, version labels, latency hints or formatting tells that identify a system. Position effects are well documented in pairwise judging, so assign which response appears on the left (or first) at random per item and log that assignment [1]. Response length is a second common confound: longer answers tend to look more thorough, so normalize formatting and track length per pair.
Interface choices change the data you get:
- Choice scale. A binary choice forces a decision; a 3-way scale (A, B, tie) or a 5- or 7-point scale (A much better to B much better) captures strength and lets you analyze ties explicitly.
- Rationale field. A short required free-text reason helps adjudication and catches raters clicking through without reading.
- Rubric visibility. Show the same written criteria (correctness, policy adherence, completeness, tone) to every rater, and record the rubric version on each judgment.
- Multi-turn context. For conversational products, show the full prior conversation and evaluate only the final turn, or both responses will be judged on different context.
- Attention checks. Seed known-answer pairs, where one response is clearly wrong, to measure rater diligence.
Who rates: crowd panels, trained generalists and domain experts
The right rater pool is the cheapest one whose judgments agree with experts on your task. Crowd agreement with experts on general chat does not transfer automatically to medical triage, tax workflows or code review, so test it on a stratified sample before trusting it. For specialized prompts, use domain-expert raters for LLM evaluation, or a mixed design where generalists rate everything and experts adjudicate a stratified sample.
Record rater metadata on every judgment: a pseudonymous rater ID, qualification level, training date, rubric version and time spent. Without these fields you cannot detect a single rater dominating a slice, rater drift over a multi-week collection, or AI-assisted rating. Ask your collection provider to disclose whether raters may use AI tools, consistent with the practices in human annotation provenance.
Agreement ceilings and tie handling
Human raters disagree on a meaningful share of pairs, and that disagreement sets the ceiling on how sharp any comparison can be. Practitioner guides report that strong LLM judges reach about 80% agreement with humans, close to human-to-human agreement, and use roughly 80% as the realistic ceiling [3]. Reviews of RLHF reward modeling report inter-annotator agreement on preference data in roughly the 63-72% range [6]. A panel that reports 95% agreement on open-ended prompts is more likely measuring easy pairs or shared shortcuts than unusual rigor.
Measure agreement with a chance-corrected statistic. Krippendorff's alpha handles any number of raters, missing ratings and ordinal scales [7], which suits a 5-point preference scale with uneven overlap; the trade-offs are summarized in inter-annotator agreement metrics for dataset buyers. Our page on preference data quality, noise and ambiguity covers diagnosing low-agreement pairs.
Decide tie handling before collecting data, then report it. Common conventions are to count a tie as half a win, to exclude ties and report the non-tie win rate, or to report wins, ties and losses as three numbers. LIMA's human study reported its results as "equivalent or preferred" rather than wins alone [8], which shows how a tie convention changes the headline. A high tie rate on a slice is itself a finding: the two models may be indistinguishable there, or the rubric may be too coarse.
Sample size for win-rate differences
Paired designs, where every prompt is answered by both models and judged once or more, are far more sensitive than comparing two separately scored samples [2]. Because each prompt acts as its own control, prompt difficulty cancels out, and the analysis reduces to testing whether the win rate differs from 50%.
A useful planning approximation: with n non-tie judgments and a true win rate near 50%, the standard error is about the square root of 0.25/n. At 400 non-tie judgments the 95% interval is roughly plus or minus 4.9 percentage points; at 1,600 it narrows to about 2.5 points. Detecting a 55% versus 45% split with about 80% power takes roughly 800 non-tie judgments per slice; detecting 52% versus 48% needs about 5,000. Bootstrap over prompts, not individual judgments, when several raters score the same prompt, because judgments on one prompt are correlated.
For comparisons across more than two models, fit a Bradley-Terry model (a logistic model of pairwise wins) on all outcomes rather than chaining head-to-head win rates. Report a confidence interval for each model's score and treat overlapping intervals as ties in the ranking.
Keep evaluation judgments separate from training preference pairs
Evaluation judgments and training preference pairs look alike but must stay in separate stores with separate licenses. The same pairwise format, a prompt and two responses with a human choice, is what trains reward models: Llama 2 collected annotator choices between two outputs to train its reward model [5], and InstructGPT used human rankings both to train and to compare models [4]. If eval judgments leak into reward-model or DPO training, the next comparison will overstate the gain.
Practical controls include separate buckets and access groups, a dataset-level use: evaluation_only flag checked by the training pipeline, hash lists of eval prompts that training deduplication excludes, and license language that matches the intended use. If you are buying data to train rather than to measure, see buying RLHF comparison data and the SourceX overview of training data for reward models. Terms to negotiate for test-only use are covered in evaluation-only data licenses.
Calibrating an LLM judge against the human panel
Most teams use the human side-by-side set to validate an automated judge, then run the judge at scale between human rounds. Calibration typically uses a stratified set of a few hundred real inputs, scored by two or three humans with the exact rubric text the judge receives, followed by agreement measurement and rubric revision [1]. Run the judge in both presentation orders and treat inconsistent verdicts as ties, since position bias affects LLM judges strongly [1]. See LLM-as-a-judge calibration sets for the label design.
Specification for a side-by-side evaluation dataset
Use the template and record schema below to brief a collection provider, an internal team or SourceX's buyer desk; adapt fields to your rubric. The broader evaluation dataset specification guide covers acceptance criteria.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Element | What to specify | Example decision |
|---|---|---|
| Prompt source | Origin, de-identification method, date range | De-identified support tickets, last 12 months |
| Slices | Tags and minimum judgments per slice | Intent x risk tier, 300 non-tie judgments each |
| Systems | Model IDs, decoding settings, system prompt | Prod v4.2 vs candidate v4.3, temperature 0.2 |
| Blinding | What raters cannot see | Names, latency, markdown differences normalized |
| Order | Randomization and logging | Per-item random left/right, seed recorded |
| Scale | Choice options and tie rule | 5-point; ties counted as half a win and reported |
| Raters | Pool, qualification, overlap | Trained generalists, 20% double-rated, expert adjudication |
| Agreement | Statistic and floor | Krippendorff's alpha, ordinal, reported per slice |
| Analysis | Test and interval | Paired bootstrap over prompts, 95% CI |
| Use | License and storage | Evaluation only; excluded from training dedup lists |
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"judgment_id": "sbs-000412-r07",
"prompt_id": "p-000412",
"prompt_slice": {"intent": "billing_dispute", "risk_tier": "high", "turns": 3},
"response_left": {"system": "blinded_B", "response_id": "r-88121"},
"response_right": {"system": "blinded_A", "response_id": "r-88120"},
"order_seed": 59113,
"preference": "left_slightly_better",
"tie": false,
"rationale": "Right response cites a refund window that is not in policy.",
"rubric_version": "support-sbs-v3",
"rater_id": "rt-0193",
"rater_tier": "trained_generalist",
"ai_assistance_disclosed": false,
"seconds_on_item": 142,
"use_restriction": "evaluation_only"
}
Sourcing real prompts for side-by-side model comparison
SourceX sources operational datasets, such as support and sales histories, engineering records and documents, from US companies on request, and every release is approved by the supplying company and delivered under a license defining records, uses, term and delivery. Personal details are removed or replaced before delivery, with the method recorded and a sample checked. Describe the prompt data your comparison needs at SourceX for AI data buyers.
Frequently asked questions
Can an LLM judge replace human side-by-side raters?
Partly. Strong judges approach human-human agreement on general prompts [3], but they inherit position and verbosity biases and need recalibration whenever the rubric, domain or candidate models change. Keep a human panel for launch decisions and high-risk slices.
Should raters see a 5-point scale or a binary choice?
Use a graded scale when you need to distinguish "slightly better" from "much better" or want explicit ties; collapse it to wins, ties and losses for the headline. Binary choices are faster but hide indifference and inflate apparent differences on ambiguous prompts.
How often should the comparison set be refreshed?
Refresh when production traffic shifts, when a new product surface launches, or when you suspect prompts have leaked. Many teams keep a stable core for trend lines and rotate a fresh share each release cycle.
Sources
- OneUptime, "How to Calibrate an LLM-as-a-Judge Against Human Labels with Cohen's Kappa" (2026). https://oneuptime.com/blog/post/2026-08-31-calibrate-llm-judge-cohens-kappa/markdown
- Python Package Index, "abeval". https://pypi.org/project/abeval/
- LangChain, "How to Calibrate LLM-as-a-Judge". https://www.langchain.com/articles/llm-as-a-judge
- Ouyang et al., arXiv, "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
- Touvron et al., arXiv, "Llama 2: Open Foundation and Fine-Tuned Chat Models" (2023). https://arxiv.org/pdf/2307.09288
- alphaXiv, "Secrets of RLHF Part II: Reward Modeling (overview)" (2024). https://www.alphaxiv.org/overview/2401.06080
- Klaus Krippendorff, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- Zhou et al., arXiv (NeurIPS 2023), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.