On this page

The short answer

Prompt caching can reduce repeated input processing when requests share a reusable prefix. It does not mean the application is returning an old answer. Measure actual billed input, writes where applicable, output, latency, and accepted outcomes; a high hit rate alone does not establish lower total operating cost.

What to take away

  • Keep reusable instructions stable where practical.
  • Separate prompt caching from application response caching.
  • Check retention and billing rules for the exact model.

Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.

Understand what is being reused

OpenAI describes prompt caching as reuse of intermediate state for a matching prompt prefix. The model still processes new input and generates a response. Its current documentation makes pricing and retention behavior model dependent. Avoid copying a configuration from an older model without checking support. Our recommendation is to begin with a request inventory: identify repeated instructions and tools, then measure how much of the actual traffic shares those prefixes.

Evidence: OpenAI prompt caching [1]

Build a realistic traffic experiment

Compare equivalent traffic with the same model, task distribution, and output expectations. A repeated synthetic prompt can produce an attractive cache result that disappears under varied customer requests. Include both cold starts and repeat interactions, and record the observation window. If the experiment changes the prompt layout, confirm that answer quality still meets the same rubric. A cost improvement achieved by removing useful context is a different intervention and deserves a separate explanation.

Account for the whole request

Keep cached reads, uncached input, cache writes where billed, and generated output in separate accounting fields. Add tool costs and retries when calculating the task total. As an illustrative example, halving an input component that represents one fifth of total cost reduces the total by one tenth, assuming everything else stays constant. It does not halve the bill. Use actual usage records and current provider rates when replacing that simple example with a forecast.

Suggested measurement plan; examples are illustrative, not industry benchmarks.
MetricHow to calculate or checkDecision it supports
Cached-input shareCached input tokens / total reported input tokensHow much input benefits from reuse?
Observed request costSum of billed input, write, output, and tool componentsAre invoices consistent with the forecast?
Time to first token p9595th percentile from request start to first output tokenDoes reuse improve perceived delay?
Cost per accepted taskTotal attempt and review cost / accepted tasksDoes caching improve the business outcome?

Treat cache changes as operational changes

Check whether the selected retention behavior fits the data-handling commitments for the application. Keep tenant authorization independent of caching, and avoid storing sensitive material simply to raise the hit rate. Watch for traffic changes that reduce prefix reuse, such as frequent prompt edits or highly variable tool definitions. After deployment, compare the forecast with observed cost and quality. If savings disappear, investigate usage composition before assuming that the provider changed its rates.

Limits of the evidence

Caching behavior, accounting fields, and retention options differ by provider and model. The calculation is illustrative; this article reports no measured Rubrex latency or savings and does not promise a particular discount.

Common questions

Is prompt caching the same as saving responses?

No. Prompt caching reuses input processing; application response caching can return a previously stored answer and needs its own freshness policy.

Is a higher cache hit rate always better?

No. Total cost, quality, retention requirements, and latency matter. A cached request can still produce an expensive or unusable answer.

Sources & further reading

Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.

  1. OpenAI prompt caching OpenAI · Living reference · Maintained documentationReviewed: prefix reuse, cache accounting, and model-specific retention guidance. Accessed September 25, 2026.

Questions or corrections? Write to Rubrex. Read our editorial approach.

From reading to a repeatable evaluation

Put a quality bar around your use case.

The AI Reliability Sprint covers one use case, a baseline, failure analysis, and one improvement experiment in 10 business days.

Explore the Sprint