Skip to content

Evaluation and benchmarking datasets

Designing evaluation rubrics with domain experts

Quick answer

Good LLM evaluation rubric design turns expert judgment into criteria that each decide one observable thing about one output: a required fact, a forbidden claim, a policy step, a disqualifying error, or a few graded qualities with written anchors. Build the criteria from real failures and the rules experts already apply, pilot the rubric with two or three experts on the same outputs, rewrite every criterion they split on, and version the final text so human raters and an LLM judge apply identical words.

By SourceX Editorial · Updated

What a rubric must do for human raters and an LLM judge

A rubric is the written decision procedure that lets different graders reach the same verdict on the same output, so its job is consistency, not completeness. One calibration guide states the operating rule: human raters should grade with the exact rubric text the judge will receive, and if the humans cannot agree, the rubric itself is the problem [1]. A judge model will not repair an ambiguity that experts could not resolve.

Keep four artifacts separate, because teams often merge them:

  • Criteria: the individual checks, each with a definition and a rule for when it applies.
  • Anchors: pass, fail and borderline examples (or one example per level on a graded scale) that fix where each boundary sits.
  • Gold labels: expert verdicts on specific outputs, stored per criterion with the rubric version used.
  • Judge prompt: the rubric text plus output-format instructions. It adds formatting, never new criteria.

The rubric sits upstream of the rest of an evaluation program. Calibrating an LLM judge and calibrating human raters both measure agreement with a rubric, and neither can succeed if the criteria are vague. For the wider map of evaluation data, start at the LLM evaluation datasets hub.

Eliciting criteria: failures and decision rules, not adjectives

Start from what experts reject and why, not from quality words like "accurate", "helpful" or "professional". Put a few dozen real outputs from your agent in front of two or three practitioners, ask each to mark every output acceptable or not and write the reason, then cluster the reasons. Each recurring reason becomes a candidate criterion; one-off reasons wait in a backlog until they recur.

Experts also apply rules that failure review will not surface until the agent hits them. Mine the artifacts their operation already uses to judge human work:

  • QA scorecards from contact centers and back offices, which already split quality into scored items (see human feedback datasets with QA scores and corrections).
  • Review checklists and policy manuals, such as a payables desk's exception procedure or an underwriting referral guide.
  • Reviewer corrections and redlines, where the edit shows which rule was broken.
  • Escalation and denial reason codes, which are failure taxonomies someone already maintains.

Public benchmarks follow the same pattern. τ-bench builds its customer-service tasks on domain-specific policy documents, with ground-truth annotations per user scenario, and scores an agent by comparing the final database state with an annotated goal state [2]. LegalBench was assembled from tasks designed and hand-crafted by legal professionals, 162 in all across six types of legal reasoning [3]. In both, what the benchmark encodes is a written rule and a verdict, not a general impression.

Then rewrite every adjective as an observable condition. "Accurate" becomes "states the variance amount shown on the match report". "Compliant" becomes "does not recommend payment when the remit-to bank account changed without a recorded callback". If two experts cannot say what evidence would make a criterion fail, it is not ready.

Licensed operational records can supply both rubric seeds and graded examples, because QA scorecards, approval decisions and reviewer corrections are expert verdicts recorded under a house rubric. SourceX sources operational datasets from US companies, including support histories, documents, and finance and legal workflows, on request rather than from stock; each dataset goes through rights review and is delivered under a license that defines the records included and their allowed uses. If such records would seed your rubric, describe the records and uses you need, and see licensing QA scorecards for AI. Expect personal details to be removed or replaced before delivery, and confirm that de-identification keeps the fields your criteria depend on (de-identifying evaluation data without breaking the test).

Pass/fail items, required and forbidden claims, or graded scales

Default to binary pass/fail criteria, use required and forbidden claims for factual content, and reserve graded scales for qualities experts genuinely rank, with every level anchored. A binary item gives raters one boundary to agree on and gives a judge a verdict that can be checked directly; one published evaluator-validation workflow labels about 100 traces Pass/Fail and splits them into train, dev and test sets [4].

Scoring formUse it whenWhat the grader recordsAgreement statisticTypical failure
Gating criterionOne error makes the output unusable: wrong decision, unsafe action, policy breachPass/fail plus an evidence quoteCohen's kappa (two raters) or Krippendorff's alphaNot marked as gating, so a high average hides a disqualifying error
Binary criterionOne observable property: field present, step taken, correct routingPass, fail or not applicableKappa or alpha (nominal)Compound wording ("correct and complete") that raters split on
Required and forbidden claimsFactual content where wording can varyEach required claim present or absent; any forbidden claim presentAgreement per claimClaims written from one reference answer, so they encode one phrasing
Ordinal scale with anchors (3 to 4 levels)Experts reliably rank a quality, such as how actionable a note isLevel plus the anchor it matchedKrippendorff's alpha (ordinal)Unanchored middle levels that each rater reads differently
Pairwise preferenceComparing two systems on overall qualityPreferred output, optional tieAgreement on the preferenceUsed where an absolute pass line is needed

Required and forbidden claims suit knowledge-heavy domains. A guide on RAG ground truth recommends "source first, answer second": the annotator locates the authoritative evidence, decides whether the corpus can answer the question, and only then writes required and forbidden claims, so different phrasings of a correct answer can pass [5]. TREC's 2025 RAGTIME track works the same way at scale: its human-curated nuggets are question-answer pairs, a nugget counts only when the report sentence is grounded in its citations, and nuggets credit any one of several acceptable answers [6].

When a graded scale is warranted, name each level and give it a meaning. The TREC 2022 NeuCLIR track had assessors judge as if gathering material for a report, choosing among four named categories from Very Valuable to Not Relevant, which the released judgments map to a three-level graded scale [7]. For head-to-head model comparison, use a human preference evaluation design rather than forcing a preference question into an absolute rubric.

Anchors, applicability rules and the borderline cases that decide agreement

Rater disagreement tends to cluster at a criterion's boundary, so each criterion needs a pass example, a fail example, at least one ruled borderline case, and a rule for when it does not apply. Without a "not applicable" option, raters guess, and the guesses show up as disagreement.

Four construction rules prevent most rework:

  1. One decision per criterion. Split "identifies the variance and routes it correctly" into two criteria.
  2. Gates before scores. List disqualifying errors first; any gate failure fails the output regardless of other criteria.
  3. Evidence on every verdict. Require the grader to quote the output span, and the source span for factual claims, so adjudicators can see the reasoning.
  4. Fixed context. State exactly what the grader sees: input, retrieved documents, tool results, policy excerpt. A rater who cannot see the purchase order cannot judge a claim about it.

Illustrative example: invented to show structure; it does not describe an available dataset.

rubric_id: ap-exception-note
version: 1.3.0
task: >
  Agent reviews an invoice that failed three-way match and writes an exception
  note recommending pay, hold or route_to_buyer
grader_sees: [invoice_pdf, purchase_order, goods_receipt, match_report,
              vendor_master_change_log, ap_policy_excerpt]
criteria:
  - id: G1
    type: gate
    text: Recommends hold when the remit-to bank account differs from the vendor
          master and no callback verification is recorded.
    pass_example: "Hold: bank details changed 2 days ago; no callback logged."
    fail_example: "Pay: amounts match PO."
    borderline_ruling: Callback made to a phone number supplied in the change
          request counts as not verified.
  - id: B1
    type: binary
    text: Names the variance type shown on the match report (price, quantity,
          missing receipt).
  - id: B2
    type: binary
    text: Routes a price variance above policy tolerance to the buyer named on the PO.
    not_applicable_when: match report shows no price variance
  - id: C1
    type: required_claims
    claims: [PO number, invoice number, variance amount, tolerance applied]
  - id: F1
    type: forbidden_claims
    claims: ["any PO, receipt or approval absent from the provided documents"]
  - id: O1
    type: ordinal
    text: Actionability of the note for the AP specialist
    anchors:
      3: Next step, owner and required document all explicit
      2: Next step explicit; owner or document missing
      1: Restates the exception without a next step
scoring: Output fails if G1 or F1 fails; otherwise report each criterion separately.
changelog:
  - "1.3.0: split routing out of B1 into B2 after pilot disagreement"

Note the scoring line: report criterion-level results, not one weighted total. A total lets a fluent note with a disqualifying error outscore a terse note that made the right decision. Teams testing payables or reconciliation agents can pair this structure with the test-set design in accounting agent evaluation datasets.

Piloting the rubric: agreement per criterion before scale

Before anyone grades at volume, have a small group of experts independently grade the same real outputs with the final rubric text, compute agreement for each criterion, and rewrite the criteria that fall below a threshold you set in advance. Agreement on the overall verdict can look acceptable while one criterion is close to random.

Pick the statistic for the scale. Krippendorff's alpha measures agreement among any number of raters, tolerates missing ratings and handles nominal, ordinal and interval data; 1 means perfect reliability and 0 means agreement no better than chance [8]. One 2026 vendor guide on annotation quality applies Krippendorff's conventions: alpha of at least 0.800 for reliable data, with 0.667 to 0.800 supporting only tentative conclusions [9]. Raw percent agreement can look impressive when one label dominates, as with a gate that almost every output passes, which is why calibration guides prefer chance-corrected measures [1].

Treat pilot numbers as directional. A position paper at ICML 2025 argues that confidence intervals based on the central limit theorem are unreliable for LLM evaluations with fewer than a few hundred items [10], and a rubric pilot is usually much smaller. Use the pilot to find broken criteria, not to certify the rubric.

Diagnose each disagreement before changing the text:

What the graders' evidence showsLikely causeFix
Both quoted the same span but reached different verdictsCriterion wording admits two readingsRewrite; add the case as a ruled borderline example
One rater cited a document the other did not seeInconsistent grading contextFix what the grader sees; regrade
Raters cite different internal policies or practicesGenuine expert disagreement about the ruleEscalate to the policy owner; record the ruling in the rubric
One rater diverges from the others across criteriaRater error or training gapRetrain or replace; see rater calibration
Disagreement concentrated in one input typeCriterion does not fit that sliceAdd an applicability rule or a slice-specific criterion

Adjudicate disagreements rather than taking a majority vote, because a majority label hides the ambiguity the pilot exists to find. Who the experts are matters as much as how many; for screening and contracting them, see sourcing domain-expert raters for LLM evaluation.

Handing the same rubric to an LLM judge

Give an LLM judge the identical criterion text the experts used, ask for one criterion's verdict per call with a quoted span as evidence, and trust it only after checking its verdicts against expert labels it never saw while you tuned the prompt. LangChain, an evaluation tooling vendor, describes the loop as collect human labels, run the judge, inspect disagreements, refine the rubric and repeat, and recommends adding labeled examples of correct judgments to the prompt so the judge learns where criterion boundaries sit [11]. The same article cites a benchmark in which strong LLM judges reached about 80% agreement with human evaluators, similar to human-human agreement, while noting that your rate depends on how well the evaluator prompt captures what good means for your use case [11].

Public evaluations already divide the work this way. In TREC 2025 RAGTIME, long topics were scored by Auto-ARGUE, an automatic process built on the human-curated nuggets, with a Llama 3 70B Instruct model as the judge [6]. Humans wrote the rubric; the model applied it.

Three rules follow. Assign gates and forbidden-claim checks to the judge only if it sees the evidence those criteria need. Re-validate the judge whenever the rubric version, the judge model or the product changes [11]. Keep the labels used to tune the judge prompt apart from the labels used to report its accuracy [4]; building that held-out label set is covered in LLM-as-a-judge calibration sets.

Versioning the rubric and accepting graded data from others

Treat the rubric as a versioned artifact: every grade records the rubric version, criterion-level verdicts, rater, evidence and any adjudication, and a rubric change triggers regrading of the items it affects. Mixed rubric versions make scores incomparable across runs.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "item_id": "ap-0412",
  "output_id": "agent-v7-run3-ap-0412",
  "rubric_id": "ap-exception-note",
  "rubric_version": "1.3.0",
  "rater_id": "R-07",
  "rater_role": "accounts payable specialist",
  "verdicts": [
    {"criterion": "G1", "value": "fail",
     "evidence": "Output: 'Recommend pay'. Change log: bank account updated, no callback."},
    {"criterion": "B2", "value": "not_applicable",
     "evidence": "Match report: quantity variance only"},
    {"criterion": "O1", "value": 2, "evidence": "Next step stated; no owner named"}
  ],
  "overall": "fail",
  "adjudication": {"status": "none_required"},
  "ai_assistance_used": false
}

The same record is your acceptance standard when a vendor or supplier grades outputs or delivers expert-labeled items. Ask for:

  • The rubric text, version history and ruled borderline cases, not only final labels.
  • Criterion-level labels with evidence, plus per-criterion agreement on a double-graded subset.
  • Adjudication records and the qualifications behind each rater role.
  • Disclosure of any LLM drafting or pre-labeling; provenance for human-annotated data covers what to require.
  • An audit sample you regrade yourself, following gold-label audits for delivered eval sets.

The audit is not a formality. Northcutt and colleagues estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets and showed that such errors can change model rankings [12]. Data quality standards and documentation formats both treat labeling as part of data quality: ISO/IEC 5259-4 includes data labelling and evaluation within its process framework for ML data quality [13], and Croissant-RAI lists data labeling among the uses of its machine-readable dataset documentation [14].

For high-risk systems under the EU AI Act, Article 10(2) requires data governance practices covering data preparation, including annotation and labelling, for training, validation and testing data sets [15]. As of October 2026, the Act has been amended by Regulation (EU) 2026/1744, and secondary sources report that the Annex III application date moved to 2 December 2027, while that regulation also amends Article 10 [16]. If your agent may be classified as high-risk, keep the rubric and its change log with that data governance record.

Rubric mistakes that surface only after scaling

  • A single 1-to-10 holistic score with no anchors, which no two experts read the same way.
  • Criteria that depend on information the grader is not shown.
  • A rubric written by the data vendor, so the evaluation measures its definition of quality (commissioning a custom evaluation dataset covers who should own it).
  • Changing the rubric mid-collection without bumping the version and regrading.
  • An expert pool drawn from one firm, which encodes one house style as the standard.
  • Reusing an evaluation rubric as a training reward without separate checks; rubrics as rewards treats that use on its own terms.

Need expert-graded business records for your rubric?

Describe the task your agent performs, the outputs your experts grade, and the operational records that could seed criteria or serve as graded examples, such as QA scorecards, reviewer corrections or approval decisions. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Describe your evaluation data needs.

Sources

  1. OneUptime (engineering blog), "How to Calibrate an LLM-as-a-Judge Against Human Labels with Cohen's Kappa" (2026). https://oneuptime.com/blog/post/2026-08-31-calibrate-llm-judge-cohens-kappa/markdown
  2. Yao et al., Sierra (arXiv:2406.12045), "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  3. Guha et al. (NeurIPS 2023 Datasets and Benchmarks), "LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models" (2023). https://proceedings.neurips.cc/paper_files/paper/2023/hash/89e44582fd28ddfea1ea4dcb0ebbf4b0-Abstract.html
  4. Kortix marketplace, hamelsmu/evals-skills (practitioner guide), "validate-evaluator". https://kortix.com/marketplace/hamelsmu--evals-skills/validate-evaluator
  5. OneUptime (engineering blog), "How to Build Ground Truth for RAG Evaluation When No Reference Answers Exist" (2026). https://oneuptime.com/blog/post/2026-08-31-build-rag-ground-truth-without-reference-answers/markdown
  6. TREC RAGTIME track organizers (arXiv:2602.10024), "Overview of the TREC 2025 RAGTIME Track" (2026). https://arxiv.org/pdf/2602.10024
  7. TREC NeuCLIR track organizers (arXiv:2304.12367), "Overview of the TREC 2022 NeuCLIR Track" (2023). https://arxiv.org/pdf/2304.12367
  8. Klaus Krippendorff, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  9. Koji (vendor documentation, market practice), "Data Annotation Quality: Guidelines, Agreement Metrics, and Gold Tasks That Actually Work" (2026). https://www.koji.so/docs/data-annotation-quality-guide
  10. ICML 2025 (poster), "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  11. LangChain (vendor article, market practice), "How to Calibrate LLM-as-a-Judge". https://www.langchain.com/articles/llm-as-a-judge
  12. Northcutt, Athalye, Mueller (arXiv:2103.14749; NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  13. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
  14. Jain et al., MLCommons Croissant RAI task force (arXiv:2407.16883), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  15. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  16. Official Journal of the European Union (EUR-Lex), "Regulation (EU) 2026/1744 (Digital Omnibus on AI) amending Regulation (EU) 2024/1689" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data