Skip to content

Data quality, coverage and contamination

Inter-Annotator Agreement for Dataset Buyers: Choosing Cohen's Kappa, Fleiss' Kappa or Krippendorff's Alpha

Quick answer

Inter-annotator agreement (IAA) measures how consistently independent annotators label the same items, corrected for agreement expected by chance. Use Cohen's kappa for exactly two raters on every item, Fleiss' kappa for a fixed number of raters per item on nominal labels, and Krippendorff's alpha when ratings are missing, raters vary, or labels are ordinal or interval. Treat alpha of 0.800 or higher as a common reliability convention, not a guarantee, and always pair agreement with gold-set accuracy.

By SourceX Editorial · Updated

Agreement is one input to a broader acceptance decision covered in the training data quality hub. This page focuses on a narrower question that evaluation leads face when a supplier's data card says "IAA = 0.71": which statistic should have been used, and is the number good enough for SFT, reward modeling or evaluation data.

Which agreement statistic fits which labeling design

The right statistic follows from three facts about the labeling job: how many raters saw each item, whether every item has the same raters, and what scale the label lives on [2][3]. Get those three facts from the supplier before you look at any number. A statistic computed on the wrong design is not conservative or optimistic; it is simply answering a different question.

Illustrative example: invented to show structure; it does not describe an available dataset.

Labeling designLabel scaleMissing ratings?Statistic to requestCommon misuse to catch
Two fixed annotators label every itemNominalNoCohen's kappaReporting percent agreement only
Two annotators, ordered labels (e.g., 1-5 helpfulness)OrdinalNoWeighted Cohen's kappa (linear or quadratic)Unweighted kappa treats 4 vs 5 the same as 1 vs 5
Fixed group of k raters per item, raters may differ across itemsNominalNoFleiss' kappaApplying it to ordinal rubric scores
Variable number of raters per item, crowd pools, overlap subsetsAnyYesKrippendorff's alpha with the matching distance metricDropping incomplete items to force a kappa
Pairwise preference (A vs B, with or without tie)Nominal or ordinalOftenKrippendorff's alpha, plus raw agreement and tie rateAveraging Cohen's kappa over arbitrary rater pairs
Continuous scores (e.g., 0-100 quality ratings)Interval or ratioEitherKrippendorff's alpha (interval) or an ICCPearson correlation, which ignores systematic offset

Cohen's kappa compares observed agreement between two raters with the agreement their individual label distributions would produce by chance [3]. Fleiss' kappa extends chance-corrected agreement (strictly, Scott's pi rather than Cohen's kappa) to a fixed number of raters per item on nominal categories, but it assumes each item has the same count of ratings [3]. Krippendorff's alpha was built to handle any number of observers, missing values, and nominal, ordinal, interval or ratio metrics in one framework, which is why it is the default for messy production labeling [1].

Why a kappa of 0.29 can sit beside 92% raw agreement

Chance-corrected statistics drop sharply when one class dominates, so high raw agreement on imbalanced labels can produce a low kappa. This is the most common reason a supplier's two numbers appear to contradict each other. It is a property of the math, not evidence that one figure is wrong.

Illustrative example: invented to show structure; it does not describe an available dataset.

A worked example: two annotators label 100 chatbot responses as safe or unsafe.

  • Both say safe: 90 items. Both say unsafe: 2 items. They disagree on 8 items (4 each way).
  • Observed agreement p_o = 92 / 100 = 0.920.
  • Each annotator labeled 94 items safe and 6 unsafe, so chance agreement p_e = 0.94 x 0.94 + 0.06 x 0.06 = 0.887.
  • Cohen's kappa = (0.920 - 0.887) / (1 - 0.887) = 0.29.
  • Krippendorff's alpha (nominal) on the same 200 values = 1 - (199 x 16) / (2 x 188 x 12) = 0.29.

The 92% figure is mostly the two raters agreeing that common items are safe. On the 10 items where at least one rater said unsafe, they agreed on only 2, and those are exactly the items a safety classifier or reward model needs. For imbalanced tasks, ask for agreement on the minority class (positive specific agreement, or per-class alpha) alongside the overall coefficient.

What alpha or kappa value is acceptable for training data

The most cited convention, from content-analysis practice, treats Krippendorff's alpha at or above 0.800 as reliable and 0.667 to 0.800 as acceptable only for tentative conclusions [4]. Krippendorff frames these as judgments about the cost of drawing wrong conclusions, so the cutoff should move with the use of the data [1]. No single threshold is right for every dataset.

A practical way to set thresholds by use:

  • Evaluation and test sets. Hold the highest bar. Disagreement in a test set becomes noise in every model comparison, and widely used benchmark test sets have been shown to contain label errors that change model rankings [5].
  • Reward-model and preference data. Expect lower agreement because many comparisons are close calls. Low agreement here can carry signal about genuine ambiguity, so ask whether ties and confidence were captured rather than forcing a binary choice. See preference data noise and ambiguity.
  • SFT and classification training data. Moderate agreement can be tolerable if disagreements are adjudicated and the adjudicated label, not a single rater's label, is what you receive.
  • Model ceiling estimates. Agreement on a re-annotated subset tells you roughly how well a model can be expected to match humans. DocLayNet, for example, annotated a subset of pages two or three times to measure agreement and compared baseline detectors against it [6].

Write the threshold into the acceptance criteria before delivery, along with the statistic, metric and sample, so the number cannot be reinterpreted after the fact. The annotation quality audit guide covers how to run the re-annotation check on your side.

Agreement measures consistency, not correctness

High agreement only shows that annotators applied the same rule; it says nothing about whether the rule or the shared answer was right [4]. Two raters trained on the same flawed guideline can agree at 0.9 and both be wrong. Pair every agreement figure with accuracy against a gold set adjudicated by a domain expert [4].

The failure modes buyers see most often:

  • Shared bias. Annotators recruited from one pool converge on the same mistakes, inflating agreement. Verify expertise through domain-expert annotator qualification checks.
  • Pre-annotation anchoring. When a model proposes labels and annotators accept them, agreement measures acceptance of the model, not independent judgment. Ask whether IAA items were labeled blind to model suggestions.
  • Collusion or copied work. Unusually high agreement between specific rater pairs is a signal worth checking against timestamps.
  • Easy-item sampling. Agreement computed only on a convenience subset that excludes hard items overstates reliability for the full set.

Low agreement usually points to the guideline, not the annotators

When agreement is low, inspect the confusion between specific label pairs before blaming the workforce. Persistent confusion between two adjacent classes usually means the class boundary is undefined in the guideline or the taxonomy merges two real concepts. Random scattered disagreement across all classes is more consistent with rater fatigue, poor screening or unclear task instructions.

Ask the supplier for a confusion matrix between raters and for the guideline version used for each batch. If guideline revisions happened mid-project, agreement should be reported per guideline version, because pooling across versions hides both the improvement and the earlier, weaker labels. Rubric-scored and human-feedback data benefit from explicit rater calibration rounds before production labeling.

How disagreements were resolved also changes what you receive. Majority vote, expert adjudication and soft labels produce different training targets from the same raw ratings; the trade-offs are covered in resolving annotator disagreement.

How suppliers should report agreement

A single overall coefficient is not enough to accept a dataset; request agreement broken down by class, by subgroup and by batch. Subgroup-level agreement matters because a guideline can work well for one dialect, product line or customer segment and fail on another, which links agreement to the representation questions in a dataset bias audit. For regulated deployments, documented measurement of data quality also supports the MEASURE function of the NIST AI RMF [7] and, for high-risk systems in the EU, the training, validation and testing data quality criteria of AI Act Article 10 [8].

Illustrative example: invented to show structure; it does not describe an available dataset.

Agreement reporting request (paste into an RFP or data card template)

agreement_report:
  statistic: krippendorff_alpha        # or cohen_kappa / fleiss_kappa / icc
  distance_metric: ordinal             # nominal | ordinal | interval | ratio
  implementation: "library name and version used to compute it"
  overlap_sample:
    items_multiply_annotated: 1200
    selection: stratified_random       # not convenience or easy-item subset
    raters_per_item: "2-5 (variable)"
    blind_to_model_prelabels: true
  results:
    overall: { value: 0.78, ci_95: [0.74, 0.81] }
    per_class:
      - { label: "refusal_appropriate", value: 0.71 }
      - { label: "refusal_overcautious", value: 0.58 }
    per_subgroup:
      - { subgroup: "es-US conversations", value: 0.69 }
    per_guideline_version:
      - { version: "v2.1", value: 0.66 }
      - { version: "v2.3", value: 0.81 }
  raw_agreement: 0.86
  gold_set_accuracy: { items: 300, adjudicator: "domain expert", accuracy: 0.93 }
  disagreement_resolution: expert_adjudication  # majority_vote | soft_labels
  rater_level_ids: pseudonymous                 # enables per-rater analysis

Ask for a confidence interval as well. Alpha computed on 60 overlap items can swing widely, so a point estimate near a threshold should not settle acceptance on its own. Where you can, request the raw per-rater labels for the overlap sample with pseudonymous rater IDs so you can recompute the statistic yourself.

Applying these checks to operational records

Many datasets buyers license are not purpose-built annotation projects but operational records whose labels came from business processes: support ticket dispositions, QA scorecards, claim outcomes, document classifications. These usually have one label per record, so IAA has to be created by re-annotating a sample rather than read from the supplier's file. The resulting agreement between the historical label and fresh expert labels is often the most informative quality number for this kind of data; verifying outcome labels in operational records walks through that check.

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents and finance and legal workflows, and manages the licensing process for buyers. Datasets are not held in stock, so a request does not guarantee a match. Buyers describing human feedback QA scores and corrections or other annotated data can set out their agreement requirements in the request through the SourceX buyer intake, and diligence materials covering source, rights, preparation and allowed use are prepared per dataset.

Specify agreement requirements for inter-annotator agreement in your data request

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, and nothing is contracted until a supplier agrees. Describe the labeled data and the agreement evidence you need at https://sourcex.si/buyers.

Frequently asked questions

Can I average Cohen's kappa across rater pairs instead of using Fleiss' kappa or alpha?

You can, and some teams report it as "mean pairwise kappa," but the result depends on which pairs happened to overlap and how many items each pair shared. With variable raters per item, Krippendorff's alpha uses all available ratings in one estimate [1], which makes it easier to compare across batches and suppliers.

Should I use Krippendorff's alpha for LLM-as-a-judge agreement with humans?

Treat the judge as one more rater and compute the same statistic you use between humans, on the same items and rubric. If judge-to-human agreement is close to human-to-human agreement, the judge is performing near the human ceiling; if it is far above, check whether the humans saw the judge's output first.

Is percent agreement ever enough?

Percent agreement is worth reporting next to a chance-corrected statistic because it helps diagnose imbalance, as in the worked example above. On its own it overstates reliability whenever one label dominates, so it should not be the acceptance metric.

Sources

  1. University of Pennsylvania, Annenberg School for Communication (Klaus Krippendorff), "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  2. arXiv, "Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation" (2026). https://arxiv.org/pdf/2603.06865
  3. aman.ai, "Primers: Inter-Annotator Agreement". https://aman.ai/primers/ai/inter-annotator-agreement/
  4. Koji, "Data Annotation Quality Guide" (2026). https://www.koji.so/docs/data-annotation-quality-guide
  5. arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  6. arXiv (IBM Research), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
  7. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
  8. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data