On this page

The short answer

There is no universal evaluation sample size. Choose the number of independent cases from the precision you need, the kinds of failures you must cover, and the difference you want to detect. Repeated generations measure variability but do not replace diverse tasks.

What to take away

  • Coverage and statistical precision answer different questions.
  • Treat repeated outputs from the same task as related observations.
  • Report uncertainty and the sample construction process with the score.

Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.

What the research says about limited samples

A 2025 preprint studies Bayesian capability assessment when samples are limited. A 2026 active-evaluation paper studies how choosing which tasks to evaluate affects ranking efficiency. These address different settings; neither supplies a universal number of examples for a business application.

Evidence: Confidence in Large Language Model Evaluation: A Bayesian Approach to Limited-Sample Challenges [1]Active Evaluation of General Agents: Problem Definition and Comparison of Baseline Algorithms [2]

Separate coverage from precision

Coverage asks whether the important situations are represented. Precision asks how uncertain an estimated rate is under a sampling model. A large set of nearly identical easy questions can produce a narrow interval for an unhelpful population. A smaller, deliberately varied set can reveal design problems without supporting a precise production failure rate.

Begin with a map of tasks and failure consequences. For a document assistant, that might include answerable questions, missing evidence, conflicting versions, restricted documents, and long inputs. Label targeted challenge cases separately from randomly sampled traffic. Combining them without disclosure makes the overall percentage hard to interpret.

A planning calculation, with assumptions

For an independent binary outcome sampled from a stable population, a rough normal-approximation planning formula is n ≈ 1.96² × p × (1 − p) / e² for a 95% interval with margin e. Using p = 0.5 and e = 0.05 gives about 385 independent cases. This arithmetic is an illustration, not a recommended default.

Real evaluation sets often violate these assumptions. Multiple questions from one document, repeated agent trials, and synthetic variations can be correlated. Rare severe failures require a different design from estimating a common success rate. For small counts or extreme rates, use an appropriate interval rather than trusting this approximation. Model comparisons also need paired analysis, not two unrelated margins.

Illustrative example: ten retries are not ten customers

Imagine 40 support scenarios, each generated ten times. There are 400 outputs, but only 40 distinct scenarios. The retries show whether the model behaves consistently on those tasks. They do not add coverage of 360 new customer situations.

Keep a scenario-level summary and a trial-level summary. If the release decision concerns generalization to new requests, add independent scenarios. If it concerns unstable behavior on a known critical request, allocate more repeated trials there. Document which uncertainty the additional spending is intended to reduce.

Set the stopping rule before looking at the winner

Decide in advance what improvement would matter, which regressions block release, and what uncertainty is acceptable. A pilot helps estimate failure prevalence and disagreement, after which a larger study can be scoped. Avoid repeatedly checking results and stopping only when the candidate looks better; that makes the apparent confidence misleading.

  • Record independent task count and total generation count separately.
  • Keep challenge-set scores separate from traffic-weighted estimates.
  • Include per-slice counts beside percentages.
  • Escalate consequential statistical decisions to an appropriate specialist.

Limits of the evidence

The numerical example assumes independent binary observations and is only a planning illustration. The cited research uses its own models, assumptions, and evaluation populations. It does not validate a particular sample size for your deployment.

Common questions

Are 50 examples enough?

They can uncover specification gaps and obvious failures. They generally cannot justify a precise claim about rare production errors without stronger assumptions and additional evidence.

Should every category have the same sample size?

Not necessarily. Allocate cases according to exposure, uncertainty, and consequences, while explaining how any overall estimate is weighted.

Sources & further reading

Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.

  1. Confidence in Large Language Model Evaluation: A Bayesian Approach to Limited-Sample Challenges arXiv · 2025 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.
  2. Active Evaluation of General Agents: Problem Definition and Comparison of Baseline Algorithms arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.

Questions or corrections? Write to Rubrex. Read our editorial approach.

From reading to a repeatable evaluation

Put a quality bar around your use case.

The AI Reliability Sprint covers one use case, a baseline, failure analysis, and one improvement experiment in 10 business days.

Explore the Sprint