On this page
The short answer
Prompt caching can reduce repeated input processing when requests share a reusable prefix. It does not mean the application is returning an old answer. Measure actual billed input, writes where applicable, output, latency, and accepted outcomes; a high hit rate alone does not establish lower total operating cost.
What to take away
- Keep reusable instructions stable where practical.
- Separate prompt caching from application response caching.
- Check retention and billing rules for the exact model.
Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.
Understand what is being reused
OpenAI describes prompt caching as reuse of intermediate state for a matching prompt prefix. The model still processes new input and generates a response. Its current documentation makes pricing and retention behavior model dependent. Avoid copying a configuration from an older model without checking support. Our recommendation is to begin with a request inventory: identify repeated instructions and tools, then measure how much of the actual traffic shares those prefixes.
Evidence: OpenAI prompt caching [1]
Build a realistic traffic experiment
Compare equivalent traffic with the same model, task distribution, and output expectations. A repeated synthetic prompt can produce an attractive cache result that disappears under varied customer requests. Include both cold starts and repeat interactions, and record the observation window. If the experiment changes the prompt layout, confirm that answer quality still meets the same rubric. A cost improvement achieved by removing useful context is a different intervention and deserves a separate explanation.
Account for the whole request
Keep cached reads, uncached input, cache writes where billed, and generated output in separate accounting fields. Add tool costs and retries when calculating the task total. As an illustrative example, halving an input component that represents one fifth of total cost reduces the total by one tenth, assuming everything else stays constant. It does not halve the bill. Use actual usage records and current provider rates when replacing that simple example with a forecast.
| Metric | How to calculate or check | Decision it supports |
|---|---|---|
| Cached-input share | Cached input tokens / total reported input tokens | How much input benefits from reuse? |
| Observed request cost | Sum of billed input, write, output, and tool components | Are invoices consistent with the forecast? |
| Time to first token p95 | 95th percentile from request start to first output token | Does reuse improve perceived delay? |
| Cost per accepted task | Total attempt and review cost / accepted tasks | Does caching improve the business outcome? |
Treat cache changes as operational changes
Check whether the selected retention behavior fits the data-handling commitments for the application. Keep tenant authorization independent of caching, and avoid storing sensitive material simply to raise the hit rate. Watch for traffic changes that reduce prefix reuse, such as frequent prompt edits or highly variable tool definitions. After deployment, compare the forecast with observed cost and quality. If savings disappear, investigate usage composition before assuming that the provider changed its rates.
Limits of the evidence
Caching behavior, accounting fields, and retention options differ by provider and model. The calculation is illustrative; this article reports no measured Rubrex latency or savings and does not promise a particular discount.
Common questions
Is prompt caching the same as saving responses?
No. Prompt caching reuses input processing; application response caching can return a previously stored answer and needs its own freshness policy.
Is a higher cache hit rate always better?
No. Total cost, quality, retention requirements, and latency matter. A cached request can still produce an expensive or unusable answer.
Sources & further reading
Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.
- OpenAI prompt caching OpenAI · Living reference · Maintained documentationReviewed: prefix reuse, cache accounting, and model-specific retention guidance. Accessed September 25, 2026.
Questions or corrections? Write to Rubrex. Read our editorial approach.