On this page

The short answer

An effective LLM evaluation rubric defines observable requirements, anchored scoring levels, and the evidence needed for each judgment. Separate criteria that measure different failures, and keep critical violations outside an average that could hide them.

What to take away

  • Replace vague quality adjectives with observable behavior.
  • Avoid scoring the same defect under several overlapping criteria.
  • Test the rubric on disagreement cases before scaling review.

Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.

Ambiguity can be a measurement problem

A recent annotation preprint studies text-classification tasks using shared codebooks and reports that ambiguity helps explain disagreement between and within human and model annotators. Its domain is limited, but it provides a useful reason to inspect coding rules before treating disagreement as reviewer incompetence.

Evidence: Observational Equivalence of LLM and Human Annotation [1]

Start with the decision the output must support

Rubrex recommends listing the required facts or behaviors before choosing a score scale. For a support answer, useful criteria might include factual support, task completion, and appropriate uncertainty. Politeness may matter, but it should not compensate for inventing an entitlement.

Define the evidence available to the reviewer. If a factuality judgment requires a policy document, supply the applicable version. A reviewer who sees only the response cannot reliably decide whether its claims are supported by private source material. Include a not-assessable outcome where necessary.

Use anchors that describe observable differences

The following template is a starting point for a source-grounded answer. Adapt it to the actual product and validate it with examples. A three-level scale can be easier to calibrate than a finely divided scale with indistinguishable boundaries. The number of levels is not itself a measure of rigor.

Illustrative rubric for a source-grounded support answer
CriterionPassNeeds revisionFail
Evidence supportMaterial claims follow from applicable sourcesA minor claim lacks clear supportA material claim is invented or contradicted
Task completionAddresses the request and required next stepA useful detail is missingThe main request is unanswered or misdirected
UncertaintyClarifies or abstains when evidence is insufficientUncertainty is present but unclearUnsupported certainty changes the user’s decision

Illustrative example: one defect, several labels

An answer promises an unavailable refund. That may be a factual-support failure and a policy-boundary violation. If the overall rubric subtracts points for factuality, accuracy, correctness, and truthfulness, it may count the same issue four times without adding useful information.

Keep a primary observable failure and record a separate critical-rule violation if the release decision requires it. Then test another answer that is factually correct but does not tell the user how to proceed. The rubric should distinguish those cases because their fixes differ.

Run a rubric trial before large-scale scoring

Ask reviewers to score the same varied examples independently. Discuss the evidence behind disagreements, revise unclear anchors, and preserve a small set of adjudicated examples. Re-score a fresh subset after the revision. Record rubric versions so a score change caused by a new definition is not mistaken for model improvement.

  • Specify required evidence and acceptable variation.
  • Add positive and negative examples at each important boundary.
  • Separate critical violations from weighted quality scores.
  • Keep an explicit route for ambiguous or unscorable cases.

Limits of the evidence

This is an illustrative rubric, not a validated instrument or client deliverable. The cited annotation study concerns particular text-classification tasks. Rubric quality must be checked against the task and the people or models applying it.

Common questions

Should every criterion have equal weight?

Only if that reflects the decision. Some requirements should be mandatory gates instead of weighted dimensions.

Can a model write the rubric?

It can help draft one, but the task owner must verify the requirements, remove overlap, and calibrate the anchors against examples.

Sources & further reading

Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.

  1. Observational Equivalence of LLM and Human Annotation arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.

Questions or corrections? Write to Rubrex. Read our editorial approach.

From reading to a repeatable evaluation

Put a quality bar around your use case.

The AI Reliability Sprint covers one use case, a baseline, failure analysis, and one improvement experiment in 10 business days.

Explore the Sprint