On this page
The short answer
An effective LLM evaluation rubric defines observable requirements, anchored scoring levels, and the evidence needed for each judgment. Separate criteria that measure different failures, and keep critical violations outside an average that could hide them.
What to take away
- Replace vague quality adjectives with observable behavior.
- Avoid scoring the same defect under several overlapping criteria.
- Test the rubric on disagreement cases before scaling review.
Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.
Ambiguity can be a measurement problem
A recent annotation preprint studies text-classification tasks using shared codebooks and reports that ambiguity helps explain disagreement between and within human and model annotators. Its domain is limited, but it provides a useful reason to inspect coding rules before treating disagreement as reviewer incompetence.
Evidence: Observational Equivalence of LLM and Human Annotation [1]
Start with the decision the output must support
Rubrex recommends listing the required facts or behaviors before choosing a score scale. For a support answer, useful criteria might include factual support, task completion, and appropriate uncertainty. Politeness may matter, but it should not compensate for inventing an entitlement.
Define the evidence available to the reviewer. If a factuality judgment requires a policy document, supply the applicable version. A reviewer who sees only the response cannot reliably decide whether its claims are supported by private source material. Include a not-assessable outcome where necessary.
Use anchors that describe observable differences
The following template is a starting point for a source-grounded answer. Adapt it to the actual product and validate it with examples. A three-level scale can be easier to calibrate than a finely divided scale with indistinguishable boundaries. The number of levels is not itself a measure of rigor.
| Criterion | Pass | Needs revision | Fail |
|---|---|---|---|
| Evidence support | Material claims follow from applicable sources | A minor claim lacks clear support | A material claim is invented or contradicted |
| Task completion | Addresses the request and required next step | A useful detail is missing | The main request is unanswered or misdirected |
| Uncertainty | Clarifies or abstains when evidence is insufficient | Uncertainty is present but unclear | Unsupported certainty changes the user’s decision |
Illustrative example: one defect, several labels
An answer promises an unavailable refund. That may be a factual-support failure and a policy-boundary violation. If the overall rubric subtracts points for factuality, accuracy, correctness, and truthfulness, it may count the same issue four times without adding useful information.
Keep a primary observable failure and record a separate critical-rule violation if the release decision requires it. Then test another answer that is factually correct but does not tell the user how to proceed. The rubric should distinguish those cases because their fixes differ.
Run a rubric trial before large-scale scoring
Ask reviewers to score the same varied examples independently. Discuss the evidence behind disagreements, revise unclear anchors, and preserve a small set of adjudicated examples. Re-score a fresh subset after the revision. Record rubric versions so a score change caused by a new definition is not mistaken for model improvement.
- Specify required evidence and acceptable variation.
- Add positive and negative examples at each important boundary.
- Separate critical violations from weighted quality scores.
- Keep an explicit route for ambiguous or unscorable cases.
Limits of the evidence
This is an illustrative rubric, not a validated instrument or client deliverable. The cited annotation study concerns particular text-classification tasks. Rubric quality must be checked against the task and the people or models applying it.
Common questions
Should every criterion have equal weight?
Only if that reflects the decision. Some requirements should be mandatory gates instead of weighted dimensions.
Can a model write the rubric?
It can help draft one, but the task owner must verify the requirements, remove overlap, and calibrate the anchors against examples.
Sources & further reading
Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.
- Observational Equivalence of LLM and Human Annotation arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.
Questions or corrections? Write to Rubrex. Read our editorial approach.