On this page

The short answer

An AI failure taxonomy is a shared classification of what went wrong and where to investigate. Separate the observed symptom from the proposed cause, retain the evidence, and assign a next experiment. A label is useful when it changes what the team does next.

What to take away

  • Do not confuse a visible symptom with a proven cause.
  • Allow multiple contributing failures without double-counting incidents.
  • Connect each category to an investigation and an owner.

Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.

Recent benchmarks emphasize stage-level diagnosis

The 2026 ChemCost preprint separates errors in grounding, retrieval, procurement selection, and arithmetic within a scientific tool task. A separate enterprise RAG preprint proposes a multidimensional diagnostic framework. Both illustrate why an end-to-end score alone can be insufficient for locating a failure. Their domains and assumptions differ from general business assistants.

Evidence: Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning [1]Overcoming the "Impracticality" of RAG: Proposing a Real-World Benchmark and Multi-Dimensional Diagnostic Framework [2]

Keep three levels of description

Rubrex recommends recording a symptom, an investigation area, and a confidence level for the causal hypothesis. For example: the answer cites an outdated policy; investigate retrieval freshness; cause not yet confirmed. This prevents a reviewer from assigning a prompt bug simply because the final response is wrong.

Use a small top-level vocabulary that different reviewers can apply consistently. Expand a category when it contains failures that need materially different remedies. A taxonomy with dozens of subtle labels can create the appearance of precision while increasing disagreement and review time.

Classify before proposing the fix

Review the input, retrieved context, tool results, and output together. Ask whether the necessary evidence was available, whether it reached the model, whether the model used it, and whether the downstream system accepted the result. Several failures may occur in one trace; identify the primary decision-relevant failure while retaining contributing factors.

Track incident identity separately from labels. If one failed response receives both retrieval and citation tags, it is still one failed response in the denominator. Category counts may overlap, but that overlap should be explicit.

A practical investigation map
Observed failureNext investigation
Missing required factCheck source availability and retrieval coverage
Wrong identifier sent to a toolInspect argument construction and entity resolution
Correct output rejected downstreamInspect schema and integration contracts
Repeated action after timeoutInspect state reconciliation and retry behavior

Illustrative example: a duplicate update

An agent updates a customer record twice after a timeout. The symptom is a duplicate state change. Possible contributors include missing idempotency, an ambiguous tool response, and a retry policy that assumes failure means no side effect. Labeling the event as hallucination would obscure all three.

Test the tool boundary with controlled timeout responses and inspect the resulting state. Then compare a retry-policy change under the same conditions. The experiment is more informative than asking whether a new system prompt sounds more cautious.

Review the taxonomy when it stops guiding work

Hold a periodic review of unclassified cases and frequent disagreements. Merge labels that lead to the same investigation, and split labels that hide different interventions. Preserve the taxonomy version with evaluation runs so changes in classification do not appear to be changes in system behavior.

Limits of the evidence

The taxonomy and examples here are proposed engineering tools. Causal attribution requires investigation; traces can be incomplete, and a plausible explanation may be wrong. Findings from specialized benchmarks should not be generalized into production error rates.

Common questions

Can one failure have multiple labels?

Yes. Keep a primary category for prioritization and secondary contributors for investigation, while counting each failed case once in the overall failure rate.

Should every failure lead to a prompt change?

No. Some require better source data, tool contracts, permissions, validation, or retry handling.

Sources & further reading

Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.

  1. Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.
  2. Overcoming the "Impracticality" of RAG: Proposing a Real-World Benchmark and Multi-Dimensional Diagnostic Framework arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.

Questions or corrections? Write to Rubrex. Read our editorial approach.

From reading to a repeatable evaluation

Put a quality bar around your use case.

The AI Reliability Sprint covers one use case, a baseline, failure analysis, and one improvement experiment in 10 business days.

Explore the Sprint