On this page
The short answer
Evaluate an LLM application against a defined user task, representative examples, and explicit acceptance criteria. Combine mechanical checks with calibrated judgment, compare changes on the same cases, and inspect important failures before making a release decision.
What to take away
- Evaluate the complete workflow that the user experiences.
- Keep critical failures separate from average quality.
- Preserve the dataset, configuration, outputs, and scoring decisions.
Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.
What current evidence adds
Anthropic’s January 2026 engineering guidance treats agent evaluation as a combination of tasks, trials, graders, and the environment. Its practical distinction between an agent’s transcript and the outcome it produces helps clarify what should actually be scored. This is vendor engineering guidance, not independent proof that one evaluation recipe works everywhere.
Evidence: Demystifying evals for AI agents [1]
Start with a release decision, not a metric
Write a sentence describing the decision the evaluation must support. For example: can the support assistant answer account-policy questions without inventing an entitlement? This identifies a user, a task, and a failure that matters. A broad goal such as improving helpfulness leaves too much room for a fluent answer to receive an undeserved pass.
Next, describe the unit being evaluated. One generated response is appropriate for a summarizer. A complete session may be necessary for an assistant that asks follow-up questions. An agent that updates a record should be judged against the resulting record as well as its explanation. Record the model, prompt, retrieval configuration, and tool permissions together. Changing any of them creates a different system.
Build a small evaluation you can inspect
As a Rubrex starting method, collect routine examples, boundary cases, and known failures into separate groups. Give every example an expected behavior and a reason for inclusion. Reserve cases that the team will not repeatedly use while tuning. The first dataset should expose ambiguity in the specification before it is used to justify a launch.
Use code for conditions with an exact answer, such as required fields or allowed identifiers. Use a rubric for meaning, completeness, or appropriate uncertainty. Have reviewers independently score a calibration subset before delegating large volumes to an automated judge. Keep unscorable cases visible; silently excluding them changes the question the evaluation answers.
- Freeze a baseline and retain individual outputs, including failed requests.
- Run the candidate against the same inputs and operating constraints.
- Review paired wins and losses, not only a change in average score.
- Assign an owner to every release-blocking failure.
Illustrative example: a policy assistant
Suppose a policy assistant must cite the current policy, identify exceptions, and avoid promising approval. An answer can satisfy two conditions and still fail overall by promising something it cannot authorize. Score each condition for diagnosis, but also maintain an all-required-conditions pass result.
If a prompt change improves citation formatting while increasing unsupported promises, the evaluation has discovered a tradeoff, not a clean improvement. The next experiment should target the promise behavior and rerun the earlier cases. These are proposed test conditions, not observed Rubrex client results.
What a useful evaluation handoff contains
Deliver the examples, grader definitions, version identifiers, baseline outputs, failure taxonomy, and an explanation of unresolved uncertainty. Include instructions for rerunning the evaluation and adding a new production failure. The handoff should let another engineer reproduce the decision without reconstructing context from a presentation. If the evidence is insufficient, record the next data needed rather than converting uncertainty into a pass.
Limits of the evidence
A pre-release dataset samples possible behavior. It cannot establish that an application is reliable on every future input. This workflow is a Rubrex recommendation informed by current guidance, not a validated universal release standard.
Common questions
Do all outputs need human review?
No. Exact checks can run automatically, while ambiguous or consequential judgments need a calibrated review process. The right balance depends on the task and the cost of missed failures.
Does a good benchmark score make an application ready?
No. Public benchmark performance does not test your specific data, workflow, permissions, and acceptance criteria.
Sources & further reading
Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.
- Demystifying evals for AI agents Anthropic · 2026 · Engineering guidanceReviewed: technical article. Accessed September 25, 2026.
Questions or corrections? Write to Rubrex. Read our editorial approach.