On this page

The short answer

Use retrieval when the system needs access to relevant information, and consider training changes when consistent behavior remains a problem after a strong prompting baseline. RAG and fine-tuning can be combined, but neither automatically fixes missing permissions, poor source documents, or an unclear definition of correctness.

What to take away

  • Test whether the necessary evidence reached the model.
  • Separate knowledge access from output behavior.
  • Check current provider support before planning a training pipeline.

Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.

Identify the failure you can observe

Start with failed examples and ask whether the system had the evidence required to answer. If the policy document was absent, changing the model may hide a retrieval problem rather than solve it. If the evidence was present but the output repeatedly ignored a required format, investigate instructions and behavior. Our recommendation is to keep these failure classes separate. They lead to different experiments and make it easier to attribute any improvement to the change that produced it.

Establish a practical baseline

Try clear instructions, representative examples, and relevant context before adding a training pipeline. OpenAI describes an iterative evaluation and optimization process, but its current documentation also warns that its fine-tuning platform is winding down and is unavailable to new users. Treat fine-tuning as an architectural option whose availability depends on the provider and model. Do not design a project around an old tutorial without checking the supported surface and the lifecycle of its base model.

Evidence: OpenAI model optimization [1]

Compare the right metrics

Test retrieval changes on evidence availability and answer support. Test behavior changes on the exact consistency problem you observed. A shared held-out set helps compare the final application, while component measures show why performance changed. Keep training examples out of the final evaluation and document how near duplicates are handled. The table proposes measurement definitions; it does not claim that one technique achieves a particular accuracy or that any single threshold applies to every product.

Suggested measurement plan; examples are illustrative, not industry benchmarks.
MetricHow to calculate or checkDecision it supports
Evidence recallQuestions with required evidence retrieved / answerable questionsIs retrieval supplying the facts?
Supported answer rateAnswers supported by the allowed sources / assessed answersDoes the response use evidence faithfully?
Behavior complianceOutputs satisfying the task rubric / assessed outputsIs the behavior consistent?
Update effortHours and cost to incorporate a controlled policy changeCan the system stay current?

Run an update experiment

Choose a controlled source change, such as a revised product limit, and follow it through the system. Check which documents, indexes, prompts, examples, and caches need updating. Confirm that an older policy no longer overrides the new one and that users without access cannot retrieve it. This experiment reveals operational work that a static benchmark misses. Select the approach whose maintenance requirements the team can own, then repeat the evaluation after each material source or model change.

Limits of the evidence

This article offers a diagnostic method, not a guarantee about retrieval or training performance. Provider support changes; the OpenAI availability warning reflects documentation checked September 25, 2026 and does not describe other providers.

Common questions

Can fine-tuning replace a knowledge base?

Do not assume it provides current, attributable facts. Test the need for retrieval separately, especially when information changes or has access restrictions.

Can RAG and fine-tuning work together?

Yes. Retrieval can supply evidence while a supported training approach shapes behavior, but the combined system still needs an independent evaluation.

Sources & further reading

Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.

  1. OpenAI model optimization OpenAI · Living reference · Maintained documentationReviewed: evaluation cycle and fine-tuning availability warning. Accessed September 25, 2026.

Questions or corrections? Write to Rubrex. Read our editorial approach.

From reading to a repeatable evaluation

Put a quality bar around your use case.

The AI Reliability Sprint covers one use case, a baseline, failure analysis, and one improvement experiment in 10 business days.

Explore the Sprint