Skip to content

Retrieval, RAG and grounding data

LLM relevance labels vs human assessors for retrieval data

Quick answer

Use LLM relevance labels where volume matters more than the last point of precision: relabeling training negatives, filling pool gaps, and screening candidate passages. Pay for human or domain-expert judgments where the labels define the target: the held-out test qrels that pick your production system, the calibration set that validates the LLM judge, and the nuggets that anchor RAG scoring. Treat any automatic label as unproven on your domain until it has been measured against human judgments on your own queries.

By SourceX Editorial · Updated

What the evidence says about LLM relevance assessors

LLM assessors are now used inside official evaluation designs, but even those designs keep humans at the point where ground truth is defined. TREC's 2025 RAGTIME track judged short topics manually and scored long report topics with an automatic judge that relied on human-curated nuggets [2]. That split is a useful template for buyers: automate the scoring, not the definition of what counts as correct.

Studies of LLM relevance assessors on TREC-style collections often report strong agreement on average system rankings, though researchers note that a system tuned exclusively to an LLM judge risks performing worse for human assessors. Both sides of that debate agree on one operational point: model choice, prompt wording and grade definitions change the labels, so a judge validated once is not validated forever. Buyers should verify any specific agreement figure in the original paper before quoting it in a budget memo.

General LLM-judge research adds known biases to the risk list: preference for the first-presented candidate (position bias), for longer answers (verbosity bias) and for outputs resembling the judge's own (self-enhancement bias) [5]. In relevance labeling, position bias shows up when passages are judged in batches, and self-enhancement shows up when the same model family generates queries, retrieves, and judges. That circularity is the strongest reason to keep human labels on anything that decides a launch.

Where LLM labels are acceptable in a retrieval data budget

LLM labels are acceptable when errors are cheap, diluted across many examples, and checked against a human-labeled sample. The clearest case is training data for dense retrievers and rerankers, where mined hard negatives frequently contain the answer and teach the model the wrong boundary. One 2026 study uses open-source LLMs to relabel answer-containing hard negatives as positives and discard ambiguous ones [1]; see reranker training data for how those groups feed listwise and distillation objectives.

Other workable uses include pre-screening passages before human pooling, labeling depth beyond the human pool (for example ranks 21 to 100) to reduce unjudged-document bias, and generating development qrels for fast iteration that never decide a launch. Embedding-model and ranking-model judges are cheaper complements that can flag disagreement for human review rather than replacing LLM scoring outright [3]. In each case, keep the automatic label in a separate field so it can be filtered out later.

Where to pay for human or expert relevance judgments

Pay for human judgments wherever a label decides something irreversible or becomes a reference for other labels. That includes the final test qrels for a private collection, any set used to compare your top two or three candidate systems, the calibration set for your LLM judge, and RAG nuggets. The RAGTIME hybrid described above is the pattern to copy: humans write and validate nuggets, and automation matches answers to them [2].

Domain expertise is the second trigger. Relevance in contract review, clinical documentation, tax guidance or incident runbooks depends on knowing which superseded clause, outdated procedure or near-duplicate draft is wrong, and general-purpose models are least reliable on exactly those distinctions. The guide to relevance assessment guidelines and graded scales covers how to write a scale experts can apply consistently, and buying relevance judgments (qrels) covers vendor and pooling choices.

Do not assume human labels are clean either. Audits of widely used benchmarks estimate average test-set label error of at least 3.3% [6], so budget for adjudication of disagreements, not just a single pass.

Illustrative example: invented to show structure; it does not describe an available dataset.

Label useDefault labelerHuman roleAccept automatic labels when
Hard-negative relabeling for trainingLLM (open-weight acceptable)Audit a stratified sample per batchSampled precision on flipped labels meets your threshold
Pool depth beyond human cutoffLLMNone per item; periodic auditJudged-rate and metric stability improve without ranking flips
Development qrels for iterationLLMSpot checksNever used for launch or vendor decisions
Final test qrelsHuman domain expertsPrimary labeler plus adjudicatorNot applicable
Top-system comparisonHumanBlind side-by-side or graded judgingNot applicable; judge-gaming and circularity risk
LLM-judge calibration setHuman, 2-3 ratersDefines the referenceNot applicable
RAG nugget creationHuman expertsWrites and validates nuggetsAutomatic matching of answers to nuggets is acceptable after calibration

How to calibrate automatic qrels before trusting them

Calibrate on a stratified, human-labeled sample drawn from your own corpus and queries, then accept the LLM only where its agreement approaches human-human agreement. Market practice suggests using a representative, stratified sample rated by multiple humans using the exact rubric the judge receives, and treats human-human agreement as the practical ceiling for the LLM judge [4]. Krippendorff's alpha suits this because it handles multiple raters, missing ratings and ordinal grades, and it can compare a model against human coders as another measuring instrument [8]; the inter-annotator agreement guide compares the metric choices.

Stratify by query type, grade, document length and source system, and oversample the hard cases: near-duplicates, superseded versions, tables and partial answers. Report agreement per stratum, not one global number, because a judge that agrees well on head queries can still fail on tail queries. Also measure ranking stability directly: rescore your candidate systems with human and LLM qrels and check whether the order of the top systems changes.

A practical sequence for a search relevance lead:

  1. Freeze the rubric and prompt, including grade definitions (for example 0 to 3) and tie-breaking rules.
  2. Label the calibration set with humans first, blind to model output.
  3. Run the judge with fixed model version, temperature and passage order; repeat with shuffled order to measure position sensitivity [5].
  4. Compute weighted kappa or alpha per stratum and confusion between adjacent grades.
  5. Rescore candidate systems under both label sets and compare top-k orderings.
  6. Re-run calibration whenever the judge model, prompt, corpus or query mix changes, since small prompt changes can shift labels.

The companion page on LLM-as-a-judge calibration sets covers calibration for generation quality; this page is limited to relevance labels.

Recording label provenance so mixed qrels stay auditable

Every relevance label should carry its origin, because mixed human and automatic qrels are only reusable if you can separate them later. Standard TREC qrels files carry a topic, an iteration field, a document identifier and a grade [7], which leaves no room for who or what produced the grade. Extend the record rather than overloading the iteration column.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "query_id": "q-0412",
  "doc_id": "kb-7781#p3",
  "grade": 2,
  "label_source": "llm",
  "judge_model": "open-weight-70b@2026-09",
  "prompt_version": "rel-rubric-v4",
  "calibration_run": "cal-2026-09-a",
  "human_grade": null,
  "adjudicated": false,
  "label_use": "train_only"
}

A label_use field such as train_only, dev or test prevents automatic labels from leaking into launch decisions. When you buy human-labeled data, require annotators' AI-assistance disclosure; see provenance for human-annotated data and detecting model-generated content in purchased human data.

What the source data behind relevance labels has to look like

Labels are only as useful as the query and document pairs they grade, and the most valuable human signal often already exists in operational records. Support tickets linked to the knowledge articles agents used, resolved engineering issues linked to runbooks, and legal or finance workflows with cited documents can supply naturally occurring relevance evidence; see support tickets linked to knowledge articles and the broader retrieval and RAG data guide. These links still need expert review and a grade scale, but they ground both human and LLM labeling in real usage.

SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages licensing and ongoing purchases. Data is sourced on request rather than held in stock, every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and a request does not guarantee a match. If you need a private corpus to build expert test qrels on, describe the data you need on the buyers page. For how evaluation data fits more broadly, see AI evaluation data and the data annotation glossary entry.

Sourcing data for human-grounded relevance labels

SourceX looks for US businesses holding the operational data you describe, with each release approved by the supplying company and personal details removed or replaced before delivery. Nothing is contracted until a supplier agrees, and terms are agreed per deal. Start by describing your retrieval corpus and labeling needs at https://sourcex.si/buyers.

Sources

  1. arXiv, "ARHN: Answer-Centric Relabeling of Hard Negatives with Open-Source LLMs for Dense Retrieval" (2026). https://arxiv.org/pdf/2604.11092
  2. arXiv (TREC RAGTIME organizers), "Overview of the TREC 2025 RAGTIME Track" (2026). https://arxiv.org/pdf/2602.10024
  3. ACL Anthology (KnowledgeNLP workshop 2025), "EKRAG: Benchmark RAG for Enterprise Knowledge Question Answering" (2025). https://aclanthology.org/2025.knowledgenlp-1.13.pdf
  4. OneUptime blog, "How to Calibrate an LLM-as-a-Judge Against Human Labels with Cohen's Kappa" (2026). https://oneuptime.com/blog/post/2026-08-31-calibrate-llm-judge-cohens-kappa/markdown
  5. arXiv (Zheng et al.; NeurIPS 2023), "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (2023). https://arxiv.org/html/2306.05685v4
  6. arXiv (Northcutt, Athalye, Mueller; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  7. NIST TREC, "Data - English Relevance Judgements". https://trec.nist.gov/data/reljudge_eng.html
  8. University of Pennsylvania, Annenberg School for Communication (Krippendorff), "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data