On this page

The short answer

A useful evaluation dataset records the input, relevant context, expected behavior, provenance, and failure category for each case. Combine authorized real examples with carefully reviewed synthetic cases, and preserve a held-out set for decisions that tuning should not influence.

What to take away

  • Store why a case exists, not only its prompt and answer.
  • Synthetic examples expand coverage but need independent validation.
  • Keep dataset changes versioned and reviewable.

Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.

Recent work makes dataset structure a design choice

HieraRAG, a June 2026 preprint, investigates the granularity of synthetic questions in RAG benchmarks. It frames granularity as something to assess within a specific retrieval configuration. This supports asking which distinctions a dataset exposes rather than assuming that more questions automatically create a better benchmark.

Evidence: How Fine-Grained Should a RAG Benchmark Be? A Hierarchical Framework for Synthetic Question Generation [1]

Define the record before collecting examples

Rubrex recommends a case record with an immutable identifier, input, allowed context, expected behavior, source category, collection date, and reviewer status. For a tool-using workflow, include initial state and the authorized end state. For open-ended writing, use an anchored rubric instead of requiring one exact sentence.

Record whether the case came from production, a support escalation, a domain specialist, or synthetic generation. Preserve a public-safe description of its origin even when the raw interaction cannot be retained. A dataset should not quietly acquire sensitive information just because a log export is convenient.

Build coverage deliberately

Map the product into meaningful slices before writing prompts. Examples include task type, document age, input length, evidence availability, language, and user permission. Avoid the full Cartesian product when it produces combinations that cannot occur. Prioritize combinations with real exposure or a specific failure hypothesis.

For synthetic cases, specify the behavior being stressed and have a separate review step verify that the input is plausible and the reference is supported. A model can generate both a flawed question and a matching flawed answer. Agreement between those two artifacts is not independent validation.

Suggested dataset fields
FieldPurpose
case_idStable identity across runs and revisions
expected_behaviorThe observable requirement, including acceptable uncertainty
provenanceOrigin, permission to use, and collection context
slice_tagsThe task and failure conditions represented
review_statusWho or what validated the reference and whether ambiguity remains

Illustrative example: versioned policy questions

A benefits assistant receives questions about a current policy and an archived policy. Create distinct cases for current questions, explicitly historical questions, and questions with no date. Expected behavior differs: use the current rule, answer historically, or clarify the applicable date. Merely changing names in one easy example would miss this distinction.

Keep these cases linked to policy versions so a real policy update does not appear to be a model regression. Retire invalid expectations through review, with an explanation, rather than deleting difficult cases from the dataset.

Treat the dataset as a maintained product artifact

Separate a development set from a decision set and record access to each. Add new failure cases through a lightweight review queue. When reporting a score, identify the dataset version and explain material changes. Comparing last month’s result on one mix of tasks with this month’s result on another can confuse traffic changes with model improvement.

Limits of the evidence

A dataset is a constructed sample, not a complete model of future traffic. HieraRAG’s findings concern a particular research setup. The record design and policy example here are Rubrex recommendations, not a reproduction of that study.

Common questions

Can the dataset be entirely synthetic?

It can be useful during early design, but its realism and coverage need validation. Keep its synthetic origin explicit and compare it with authorized real use when available.

Should references be exact answers?

Use exact answers for exact tasks. For open-ended outputs, define required facts, unacceptable behavior, and acceptable variations.

Sources & further reading

Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.

  1. How Fine-Grained Should a RAG Benchmark Be? A Hierarchical Framework for Synthetic Question Generation arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.

Questions or corrections? Write to Rubrex. Read our editorial approach.

From reading to a repeatable evaluation

Put a quality bar around your use case.

The AI Reliability Sprint covers one use case, a baseline, failure analysis, and one improvement experiment in 10 business days.

Explore the Sprint