On this page

The short answer

Test retrieval recall against the evidence needed to answer each question, not against a list of documents that merely share its topic. Record partial coverage and missing exceptions, then assess whether additional retrieved context helps the final answer.

What to take away

  • Label supporting evidence at the level the decision requires.
  • Include questions that need more than one passage.
  • Keep retrieval budget and corpus version fixed during comparisons.

Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.

Question granularity changes what a benchmark tests

HieraRAG studies how synthetic question granularity affects discrimination between RAG configurations. Ragas separately identifies context recall and context precision as evaluation dimensions. Together they provide a reason to examine both evidence coverage and the usefulness of the returned material.

Evidence: How Fine-Grained Should a RAG Benchmark Be? A Hierarchical Framework for Synthetic Question Generation [1]List of available metrics [2]

Create an evidence requirement for each question

For each case, write the facts that must be recoverable from the retrieved context. Map each fact to one or more acceptable passages. Where several sources can support the same fact, do not require one arbitrary document identifier. Where a conclusion requires two documents, record both dependencies.

Use explicit labels for evidence that is necessary, supportive but optional, contradictory, or irrelevant. This gives reviewers a way to explain why a retrieval result fails. A general semantic-similarity score may be a useful diagnostic, but it does not establish that the required exception or qualifying condition was found.

Compare retrieval under a controlled budget

Rubrex recommends holding the question set, corpus snapshot, access policy, and generator constant when investigating a retriever change. Record the passage limit, token budget, reranking configuration, and any truncation. A comparison between ten short passages and ten entire documents does not control the same resource.

Inspect retrieval misses by category: wrong vocabulary, missing source, poor segmentation, permissions, stale indexing, or ranking. Some are not retriever-model problems. If the necessary policy never entered the index, changing embeddings will not recover it.

  • Separate corpus coverage from retrieval success within the corpus.
  • Count required evidence units, with the labeling rule documented.
  • Retain the actual context sent downstream, not only search results.
  • Measure answer behavior after any recall-oriented change.

Illustrative example: the exception lives elsewhere

A customer asks whether an annual plan can be canceled after a renewal. The general cancellation policy is retrieved, but an account-specific exception sits in a separate authorized document. A topic-relevance grader may give the result a high score while the evidence needed for the answer remains incomplete.

The reference should specify both the general rule and the applicable exception. If the exception document is inaccessible to this user, the expected answer must also respect that constraint. A retrieval evaluation should never reward exposing a document that the requesting user is not entitled to see.

Turn misses into a retrieval experiment

Select an intervention tied to the observed failure: improve indexing coverage, adjust chunk boundaries, preserve headings, add metadata filters, or change reranking. Measure the paired result and examine newly introduced failures. Keep the original examples in the regression suite even after the immediate miss is fixed.

Limits of the evidence

Recall depends on the completeness of the evidence labels. A reference set can itself omit valid supporting passages. The methods here are proposed evaluation practices; research on synthetic question granularity does not establish the optimal retrieval settings for another corpus.

Common questions

Is a larger top-k always better?

No. It can improve coverage while introducing distractors, stale content, or truncation. Measure the final answer as well as the retrieved set.

Can relevance judgments be automated?

They can assist review, but calibrate them against labeled examples and inspect disagreement on consequential evidence requirements.

Sources & further reading

Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.

  1. How Fine-Grained Should a RAG Benchmark Be? A Hierarchical Framework for Synthetic Question Generation arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.
  2. List of available metrics Ragas · Living reference · Maintained documentationReviewed: metric taxonomy. Accessed September 25, 2026.

Questions or corrections? Write to Rubrex. Read our editorial approach.

From reading to a repeatable evaluation

Put a quality bar around your use case.

The AI Reliability Sprint covers one use case, a baseline, failure analysis, and one improvement experiment in 10 business days.

Explore the Sprint