On this page

The short answer

A RAG assistant should abstain or clarify when the available evidence cannot support the requested conclusion. Evaluate unsupported answering and unnecessary refusal together, using answerable and unanswerable cases with clearly defined expected behavior.

What to take away

  • Confidence language is not a calibrated probability.
  • A refusal-only system can appear safe while being unusable.
  • Judge abstention against evidence availability and user intent.

Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.

Uncertainty deserves its own evaluation

URAG, a 2026 preprint, studies uncertainty quantification across RAG methods and domains using accuracy and prediction-set size. It treats uncertainty as a measurable property rather than a stylistic hedge. Its research framework does not supply a universal threshold for answering business questions.

Evidence: URAG: A Benchmark for Uncertainty Quantification in Retrieval-Augmented Large Language Models [1]

Distinguish answer, clarify, and abstain

An answer is appropriate when the evidence and task definition are sufficient. Clarification is appropriate when a missing user detail could resolve ambiguity. Abstention is appropriate when the required evidence is absent or the requested action is outside scope. These outcomes should have separate labels in the evaluation.

Avoid treating phrases such as probably or likely as a substitute for a decision rule. A model can hedge an unsupported assertion and still mislead the user. Review the actual claim and the next step it suggests, not only its tone.

Pair answerable and unanswerable cases

Rubrex recommends creating pairs that differ in one decisive evidence condition. One case includes an approved policy; another removes it. A third includes two contradictory policies. A fourth has enough evidence but a deliberately unfamiliar phrasing. This helps distinguish appropriate abstention from a blanket refusal strategy.

Track false answers on unanswerable cases and unnecessary refusals on answerable cases. Preserve the denominator for each. If an automated confidence signal routes cases to review, validate that signal on representative held-out examples and record what happens near the threshold.

Illustrative example: an unknown integration limit

A customer asks for the maximum size accepted by an integration, but the indexed documentation omits the limit. An unsupported numeric answer should fail even if it sounds plausible. A useful response states that the available documentation does not establish the limit and identifies a relevant support or verification step.

Now add an authoritative document with the exact limit. The assistant should answer from that source rather than continuing to refuse. This paired test makes usefulness part of the quality bar and prevents abstention from becoming an easy way to avoid all factual errors.

Choose thresholds from consequences

Set routing and release criteria around the cost of an unsupported answer, the cost of review, and the cost of a missed opportunity to help. Different tasks may need different thresholds. Document the chosen tradeoff and reevaluate it when traffic or source coverage changes.

  • Keep ambiguous cases visible instead of forcing binary labels.
  • Record whether clarification actually resolves the task.
  • Measure the usefulness of the fallback response.
  • Recheck thresholds after model, retrieval, or corpus changes.

Limits of the evidence

Uncertainty estimates depend on the task, labels, and calibration method. A verbal expression of confidence is not automatically calibrated. The proposed examples are illustrative, and URAG’s results are limited to the evaluated research configurations.

Common questions

Should a RAG system always answer from retrieved text?

Only when the retrieved material supports the task. Irrelevant or incomplete retrieval should not be treated as permission to guess.

How much refusal is acceptable?

There is no universal rate. Compare necessary abstention with unnecessary refusal and the consequences of each for the product.

Sources & further reading

Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.

  1. URAG: A Benchmark for Uncertainty Quantification in Retrieval-Augmented Large Language Models arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.

Questions or corrections? Write to Rubrex. Read our editorial approach.

From reading to a repeatable evaluation

Put a quality bar around your use case.

The AI Reliability Sprint covers one use case, a baseline, failure analysis, and one improvement experiment in 10 business days.

Explore the Sprint