On this page
The short answer
Human evaluation agreement measures consistency between reviewers, not automatic correctness. Inspect the cases and labels behind the number, distinguish ambiguity from mistakes, and use adjudication to improve the rubric while preserving the original judgments.
What to take away
- High agreement can coexist with a shared misunderstanding.
- Overall agreement can hide poor performance on rare labels.
- Preserve disagreements as evidence about the task specification.
Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.
Recent annotation research highlights codebook ambiguity
The 2026 observational-equivalence preprint examines human and model annotation on political-science text tasks and reports that clarifying coding rules can reduce disagreement. It does not establish that models and humans are interchangeable for every domain. The useful lesson is to investigate the measurement instructions themselves.
Evidence: Observational Equivalence of LLM and Human Annotation [1]
Collect independent judgments first
Rubrex recommends blind, independent labeling for the calibration subset. If reviewers see one another’s answers first, the apparent agreement may reflect conformity rather than a clear rubric. Supply the same evidence and record whether any reviewer lacked context needed to score the case.
Specify the unit of annotation and allowed labels. A sentence-level factuality task differs from an overall response-quality task. For ordinal scales, define what each level means and whether adjacent disagreements are less consequential than distant ones. Do not choose an agreement statistic before defining the measurement.
Read the label distribution beside the agreement score
Raw agreement is easy to understand, but it can be high when almost every example belongs to one class. Chance-adjusted statistics add another perspective and bring their own assumptions. Report per-label counts, confusion patterns, and the prevalence of ambiguous cases alongside the summary statistic.
A small rare-failure subset may matter more to the release decision than a high overall score. Inspect whether reviewers consistently identify that subset. Where multiple interpretations are legitimate, retain an ambiguity label or a distribution of judgments rather than manufacturing one certain reference.
Illustrative example: easy negatives dominate
Imagine a review set containing many obvious nonviolations and a few subtle policy violations. Reviewers agree on nearly all easy cases but disagree on most consequential ones. The aggregate agreement can look reassuring while the release gate remains poorly calibrated.
Review the disputed cases, identify the missing rule or evidence, and add anchored examples. Then evaluate a fresh subset with enough relevant boundary cases to assess the revised instruction. Keep targeted calibration results separate from a traffic-weighted estimate.
Make adjudication improve the instrument
Record the original labels, the final resolution, the evidence, and the rubric revision if one was required. A lead reviewer’s decision can establish an operational reference, but it should not be described as infallible truth. Track disagreement trends after changes in task mix, reviewer group, or source material.
- Blind reviewers to model identity when it is irrelevant.
- Inspect disagreement by criterion and task slice.
- Retain not-assessable cases and their reasons.
- Distinguish operational adjudication from an objective ground truth.
Limits of the evidence
Agreement statistics depend on label prevalence, sampling, reviewer expertise, and the task definition. The cited preprint’s findings concern a particular domain. They do not remove the need for qualified review in specialized or consequential settings.
Common questions
Does 90% agreement prove the labels are correct?
No. Reviewers can share a mistaken rule, and the remaining disagreement may contain the most important cases.
Should disagreements always be forced into one answer?
No. Some reflect legitimate ambiguity. Record that ambiguity and define how it affects the downstream decision.
Sources & further reading
Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.
- Observational Equivalence of LLM and Human Annotation arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.
Questions or corrections? Write to Rubrex. Read our editorial approach.