Evaluation and benchmarking datasets
LLM-as-a-judge calibration sets: the human labels that validate your judge
Quick answer
An LLM-as-a-judge calibration dataset is a set of real inputs and model outputs that qualified people have scored with the exact rubric your automated judge uses, so you can measure how often the judge agrees with them before you trust its scores. Practitioner guides work with roughly 100 to 300 real items, two or three human raters per item, chance-corrected agreement such as Cohen's kappa or Krippendorff's alpha, and a test split that is scored only once [1][2].
By SourceX Editorial · Updated
What a calibration set measures, and what it does not
A calibration set measures the judge, not the model under test: each item pairs an output with a trusted human verdict, and the judge's verdicts are compared against those. It differs from the eval set the judge later scores, from preference data used for training, and from the checks that keep human raters consistent.
Rater consistency is a precondition, covered in rater calibration for rubric-graded data, and the rubric comes first, as in designing evaluation rubrics with domain experts. For the wider evaluation stack, see the buyer's map of LLM evaluation datasets and the model evaluation glossary entry; retrieval relevance labels are covered in LLM relevance labels vs human assessors.
Statistically, the judge is one more rater: Krippendorff defines alpha as a reliability coefficient for observers, coders, judges, annotators or measuring instruments that assign values to the same units [3]. Not every task needs a judge: LiveBench scores answers automatically against objective ground truth because, its authors argue, LLM judges can introduce biases and break down on hard questions [4]. If an exact match, a unit test or a recorded business outcome can decide correctness, use that (see outcome-labeled evaluation data) and spend human labels on criteria that need judgment, such as policy compliance, tone or groundedness.
How many human labels a judge needs
Plan on a few hundred labeled items for a single-criterion judge, with enough of the rare verdict class (usually the failures the judge exists to catch) in every split to measure it. One guide recommends 150 to 300 stratified real inputs scored by two or three people using the same rubric text the judge receives [1]; another works with about 100 pass/fail traces divided into train, dev and test splits [2].
The binding constraint is the rare class, not the total. If 10% of production outputs fail, a random 200-item sample holds about 20 failures, and the judge's catch rate rests on those 20. Sample failures and boundary cases at a higher rate than production, then reweight when you report. Small sets also need honest intervals: an ICML 2025 position paper argues that central-limit-theorem confidence intervals come out too narrow on evaluations with fewer than a few hundred items and recommends other methods there [5].
| Split | What it is for | Rule that keeps it honest |
|---|---|---|
| Few-shot pool | Labeled examples placed in the judge prompt to show where criteria boundaries sit | Never reported as accuracy; its items never appear in dev or test |
| Dev | Measuring the judge while you edit its prompt, rubric wording or model | Assume you will overfit it; re-measure after every change |
| Test | One final measurement of the frozen judge | Used once [2]; kept out of prompts; replaced after it informs a decision |
Budget in human judgments, not items: 250 items with three raters each is 750 judgments before adjudication, and clinical, legal or financial criteria need domain-expert raters. Sizing the eval set itself is covered in how many examples an LLM eval set needs, and slice allocation in stratified evaluation sets for rare and high-risk cases.
Agreement numbers to require from raters and judge
Require two sets of numbers: human-human agreement on the calibration items, which sets the ceiling, and judge-human agreement broken out by class, which shows whether the judge can stand in for the raters.
Use Cohen's kappa for two raters and Krippendorff's alpha for three or more (see choosing an inter-annotator agreement statistic); one guide treats human agreement below roughly 0.7 as a sign the rubric needs work before any judge is tested [1]. Alpha also handles missing ratings, uneven rater counts and ordinal or interval scales [3][6]. Raw agreement looks impressive when one label dominates, so read kappa alongside the confusion matrix, class prevalence and the disagreements themselves [1]. For binary judges, one practitioner workflow tracks the judge's true positive and true negative rates on the dev split and iterates until both exceed 90% [2].
LangChain cites a benchmark in which strong LLM judges reach about 80% agreement with human evaluators, close to human-human agreement, and notes that results depend on how well the judge prompt captures what "good" means for the use case [7]. That is raw agreement, not a target for your rubric.
| Judge output | Primary statistic | Also report |
|---|---|---|
| Pass/fail per criterion | Cohen's kappa (two raters) or alpha (three or more) | True positive and true negative rates, confusion matrix, prevalence |
| Ordinal score, such as 1 to 5 | Alpha with an ordinal metric, or weighted kappa | Spearman correlation, size of disagreements, score distribution per rater |
| Pairwise preference (A, B or tie) | Agreement with the adjudicated preference | Verdict consistency after swapping order, tie rate |
Worked example: raw agreement, kappa and corrected pass rates
The numbers are invented to show the arithmetic. Humans label a 200-item test split 180 pass and 20 fail. A judge that passes everything agrees on 90% of items, but chance agreement is also 0.90 (0.9 × 1.0 + 0.1 × 0), so kappa is 0 and the judge catches none of the failures.
A better judge passes 171 of the 180 human passes and fails 14 of the 20 human fails. Raw agreement is 92.5%, kappa is about 0.61, the true positive rate (TPR, pass on human-pass items) is 0.95 and the true negative rate (TNR, fail on human-fail items) is 0.70. With only 20 failures, a normal-approximation 95% interval on that TNR runs from about 0.50 to 0.90, and the ICML 2025 warning about such intervals on small samples applies here [5].
Those rates also correct the judge's production numbers. If the true pass rate is θ, the observed pass rate is p = TPR × θ + (1 − TNR) × (1 − θ), so θ = (p + TNR − 1) / (TPR + TNR − 1). When the judge passes 88% of 10,000 unlabeled outputs, the corrected estimate is (0.88 + 0.70 − 1) / (0.95 + 0.70 − 1), about 0.89; if TNR is really 0.50 or 0.90, it moves to about 0.84 or 0.92. The paper "How to Correctly Report LLM-as-a-Judge Evaluations", as summarized in one review, addresses how judge-based results should be reported, including small-sample interval adjustments such as Agresti–Coull [8].
Oversampling failures to 100 items narrows the TNR interval to roughly ±0.09. TPR and TNR are conditioned on the human label, so oversampling one class does not distort them; kappa and raw agreement must be recomputed at production prevalence, and oversampled boundary cases need per-stratum reweighting.
Where calibration labels come from
Calibration labels come from three places: your own experts, contracted raters labeling fresh items, or existing records in which people already graded work against a standard. Most teams combine them, because each source fails differently.
| Label source | Strengths | Failure modes | What to check |
|---|---|---|---|
| In-house experts | Know the product and the rubric's intent | Scarce hours; rubric authors grade to intent, not text | Agreement on a blind subset; rubric version |
| Contracted or vendor raters | Scale, blinding, controlled rubric versions | Generalists miss domain errors; rater pools drift between batches | Qualifications, accuracy on gold items, per-rater agreement, adjudication log |
| Existing graded records: contact-center QA scorecards, code review approvals, claims or coding audits, editorial sign-offs | Real judgments under operational stakes, often with years of history | Grade human work under an internal form, not model outputs under your rubric; contain employee-performance data | Mapping from form to your criteria, scorer and agent identities removed, date range, form version |
A support QA scorecard is a supervisor's rubric-based verdict on a real customer interaction, the same shape of judgment a support-reply judge makes; see human feedback datasets from QA scores and corrections and QA scorecards for AI. Model outputs fail differently from human agents, with fluent fabrications and over-long answers, so add a freshly labeled slice of real model outputs before trusting mapped labels. Inherited labels also carry errors: Northcutt et al. estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets [9], so run a gold-label audit on a sample before treating inherited labels as truth.
SourceX sources operational datasets from US companies, including support and sales histories and finance and legal workflows, and manages the licensing and any ongoing purchases. Datasets are sourced on request rather than held in stock, so a request does not guarantee a match. If graded business records would anchor your judge, you can describe the records and criteria you need to SourceX.
Slices that expose judge errors
Build the set from real traffic and deliberately add the cases judges get wrong, because a purely random sample under-represents the failures the judge is meant to catch.
- Representative traffic: real inputs and outputs from the graded system, across product surfaces, languages and user segments.
- Known failures: outputs that broke policy, invented facts or missed the task, drawn from incident reviews and escalations.
- Boundary cases: items near rubric thresholds where raters disagree; keep the split votes as information rather than forcing one gold label.
- Rare high-risk cases: low-frequency, high-consequence situations, each large enough to measure alone.
- Order-swapped pairs: for pairwise judges, every pair in both orders; one guide recommends averaging verdicts across both orders to neutralize position bias [1].
- Length-controlled pairs: pairs where the better answer is shorter; the same guide recommends stripping length cues so the judge does not reward verbosity [1].
- Outputs from the judge's own model family: research on self-bias examines how LLM evaluators score output from their own models [10], so include such outputs and compare.
A calibration record that survives judge changes
Store each item with its rubric version, every individual rating, the adjudicated label and each judge verdict by judge version, so agreement can be recomputed whenever any of them changes. Per-rater labels are what a single "gold" column loses.
Illustrative example: invented to show structure; it does not describe an available dataset.
item_id: cal-support-000317
split: test # fewshot | dev | test
stratum: boundary # traffic | known_failure | boundary | rare_high_risk
input_source:
origin: licensed_business_record # production_log | licensed_business_record | commissioned
source_system: helpdesk_tickets
record_date: "2025-11-04"
deidentification: agent_and_customer_names_replaced
license_scope: evaluation_only
input: "Customer requests a refund 45 days after purchase; policy window is 30 days."
candidate_output: "I've processed a full refund for you today."
candidate_type: model_output # model_output | human_agent_reply
rubric: { id: support-policy, version: "3.2", criterion: follows_refund_policy }
ratings:
- { rater: r-07, label: fail, rationale: "Refund outside window, no exception code" }
- { rater: r-12, label: fail, rationale: "No approval reference" }
- { rater: r-19, label: pass, rationale: "Assumed VIP exception applies" }
adjudication: { label: fail, method: senior_reviewer, note: "Account has no VIP flag" }
judge_runs:
- { judge_model: judge-a-2026-09, prompt_version: p14, label: pass }
Croissant-RAI extends the Croissant metadata vocabulary with responsible-AI fields whose use cases include data labeling [11], and NeurIPS 2026 sets Croissant-RAI-based metadata requirements for its Evaluations and Datasets Track [12]. Ask for rater pools, labeling procedure and known limitations documented in that form.
Keeping the judge calibrated after launch
Measured agreement holds only for the rubric, judge model and product it was measured on, and LangChain advises recalibrating whenever any of them changes [7].
- Pin versions. Record the judge model snapshot and prompt version with every agreement figure, because hosted model aliases can change.
- Use a fresh test split after each judge, prompt or rubric change. The old split has informed a decision, so it joins the dev pool.
- Sample production for drift. Send a small, regular sample of judged production items to human raters; new features create failure types the original set lacks.
- Feed disagreements back. Tooling reflects this loop: LangChain describes calibrating judges with human corrections [13], and the Potato annotation tool documents a judge-alignment workflow [14]. Both are evidence of practice, not standards.
- Protect the test split. Every judge run sends items to the judge model's provider, so check retention and training-on-inputs terms; keeping a private eval set private covers the controls.
License and privacy questions specific to calibration data
Calibration data is often used in ways a plain evaluation license does not name, so map each use to a license term before labels are produced.
- Few-shot examples in production. Items pasted into the judge prompt run in production and reach the judge model's provider on every call.
- Fine-tuning a judge. Training a smaller judge on the labels is training use, which an evaluation-only data license may exclude.
- Ownership of new labels. Contracts with raters or labeling vendors should assign rights in labels and rationales and keep items confidential.
- Rater and employee data. Rater IDs, and agent and supervisor names on historical scorecards, are personal data; pseudonymize them. Rationales can quote customer details, so de-identify them with the inputs.
- Retention for re-audit. You need items and labels every time the judge changes, so match retention to the judge's service life.
When SourceX sources a dataset, it goes through rights review and is delivered under a license defining the included records, allowed uses and term. Names, account numbers and similar personal details are removed or replaced before delivery and the method is recorded, though no de-identification method is perfect.
Calibration set request checklist
Use this list to specify a calibration set to an internal team, labeling vendor or data supplier; each line can become an acceptance criterion.
- Judge task, output type (pass/fail, ordinal or pairwise), criteria, rubric ID and version
- Item sources: production sample, licensed records, or model outputs (which models and versions)
- Total items, plus a minimum count per verdict class and per stratum
- Raters per item, qualifications, blinding, adjudication rule, and how split votes are kept
- Agreement reporting: kappa or alpha with intervals, confusion matrix, prevalence, per-stratum results
- Splits (few-shot, dev, test), who holds the test split, and the single-use rule
- Record fields: per-rater labels and rationales, adjudicated label, provenance, de-identification method
- Permitted uses: evaluation, few-shot prompting, judge fine-tuning, third-party model APIs, retention period
- Recalibration triggers and refresh volume
Need human-graded records to calibrate a judge?
Describe the judge, its rubric and the graded records you need. SourceX looks for US businesses that hold that data, checks the data and the supplier's licensing permissions, and manages the license and delivery; every release is approved by the supplying company, and nothing is contracted until a supplier agrees. Describe your evaluation data needs.
Sources
- OneUptime blog, "How to Calibrate an LLM-as-a-Judge Against Human Labels with Cohen's Kappa" (2026). https://oneuptime.com/blog/post/2026-08-31-calibrate-llm-judge-cohens-kappa/markdown
- Kortix marketplace (hamelsmu/evals-skills), "validate-evaluator (evals-skills)". https://kortix.com/marketplace/hamelsmu--evals-skills/validate-evaluator
- Klaus Krippendorff, University of Pennsylvania Annenberg School for Communication, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- White et al., arXiv (2406.19314); ICLR 2025, "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
- ICML 2025 (position paper poster page), "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
- arXiv (2603.06865), "Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation" (2026). https://arxiv.org/pdf/2603.06865
- LangChain, "How to Calibrate LLM-as-a-Judge". https://www.langchain.com/articles/llm-as-a-judge
- The Moonlight (literature review site; German-language page), "How to Correctly Report LLM-as-a-Judge Evaluations (paper review)". https://www.themoonlight.io/de/review/how-to-correctly-report-llm-as-a-judge-evaluations
- Northcutt, Athalye, Mueller (arXiv 2103.14749; NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv (2509.26600), "When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation" (2025). https://arxiv.org/pdf/2509.26600
- Jain et al. (MLCommons Croissant RAI task force), arXiv (2407.16883), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- NeurIPS blog, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
- LangChain, "How to Calibrate LLM-as-Judge with Human Corrections". https://www.langchain.com/resources/llm-as-a-judge
- Potato annotation tool documentation (Read the Docs), "Judge alignment". https://potatoannotator.readthedocs.io/en/latest/agent-evaluation/judge_alignment/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.