On this page
The short answer
An AI failure taxonomy is a shared classification of what went wrong and where to investigate. Separate the observed symptom from the proposed cause, retain the evidence, and assign a next experiment. A label is useful when it changes what the team does next.
What to take away
- Do not confuse a visible symptom with a proven cause.
- Allow multiple contributing failures without double-counting incidents.
- Connect each category to an investigation and an owner.
Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.
Recent benchmarks emphasize stage-level diagnosis
The 2026 ChemCost preprint separates errors in grounding, retrieval, procurement selection, and arithmetic within a scientific tool task. A separate enterprise RAG preprint proposes a multidimensional diagnostic framework. Both illustrate why an end-to-end score alone can be insufficient for locating a failure. Their domains and assumptions differ from general business assistants.
Evidence: Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning [1]Overcoming the "Impracticality" of RAG: Proposing a Real-World Benchmark and Multi-Dimensional Diagnostic Framework [2]
Keep three levels of description
Rubrex recommends recording a symptom, an investigation area, and a confidence level for the causal hypothesis. For example: the answer cites an outdated policy; investigate retrieval freshness; cause not yet confirmed. This prevents a reviewer from assigning a prompt bug simply because the final response is wrong.
Use a small top-level vocabulary that different reviewers can apply consistently. Expand a category when it contains failures that need materially different remedies. A taxonomy with dozens of subtle labels can create the appearance of precision while increasing disagreement and review time.
Classify before proposing the fix
Review the input, retrieved context, tool results, and output together. Ask whether the necessary evidence was available, whether it reached the model, whether the model used it, and whether the downstream system accepted the result. Several failures may occur in one trace; identify the primary decision-relevant failure while retaining contributing factors.
Track incident identity separately from labels. If one failed response receives both retrieval and citation tags, it is still one failed response in the denominator. Category counts may overlap, but that overlap should be explicit.
| Observed failure | Next investigation |
|---|---|
| Missing required fact | Check source availability and retrieval coverage |
| Wrong identifier sent to a tool | Inspect argument construction and entity resolution |
| Correct output rejected downstream | Inspect schema and integration contracts |
| Repeated action after timeout | Inspect state reconciliation and retry behavior |
Illustrative example: a duplicate update
An agent updates a customer record twice after a timeout. The symptom is a duplicate state change. Possible contributors include missing idempotency, an ambiguous tool response, and a retry policy that assumes failure means no side effect. Labeling the event as hallucination would obscure all three.
Test the tool boundary with controlled timeout responses and inspect the resulting state. Then compare a retry-policy change under the same conditions. The experiment is more informative than asking whether a new system prompt sounds more cautious.
Review the taxonomy when it stops guiding work
Hold a periodic review of unclassified cases and frequent disagreements. Merge labels that lead to the same investigation, and split labels that hide different interventions. Preserve the taxonomy version with evaluation runs so changes in classification do not appear to be changes in system behavior.
Limits of the evidence
The taxonomy and examples here are proposed engineering tools. Causal attribution requires investigation; traces can be incomplete, and a plausible explanation may be wrong. Findings from specialized benchmarks should not be generalized into production error rates.
Common questions
Can one failure have multiple labels?
Yes. Keep a primary category for prioritization and secondary contributors for investigation, while counting each failed case once in the overall failure rate.
Should every failure lead to a prompt change?
No. Some require better source data, tool contracts, permissions, validation, or retry handling.
Sources & further reading
Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.
- Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.
- Overcoming the "Impracticality" of RAG: Proposing a Real-World Benchmark and Multi-Dimensional Diagnostic Framework arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.
Questions or corrections? Write to Rubrex. Read our editorial approach.