What does a prompt-cache hit change about an LLM request's cost and speed?
answer
- three prices, not one
- the store costs more than sending
- reuse costs roughly a tenth
- output tokens are untouched
- first token arrives sooner
basics
~20 sA hit bills the reused prefix at a fraction of the normal input price — often around a tenth — and the model skips re-processing it, so the first token arrives far sooner. Output tokens are billed and generated exactly as usual.
solid answer
~50 sPrompt caching splits input into three price tiers instead of one: ordinary uncached input at the base rate, cache-write (creation) tokens at a **premium** above base — commonly 1.25x to 2x depending on how long the entry is kept — and cache-read tokens at a deep **discount**, often around a tenth of base. A hit therefore buys two separate things. Financially, the reused prefix becomes nearly free. Operationally, the provider does not have to re-process those tokens, which is what cuts time-to-first-token on a long prompt. Nothing else improves: output tokens are billed at full rate and produced at the same speed, and the cached tokens still occupy the context window. The size of the win scales with how big the reused prefix is relative to the fresh input and the generated answer. Exact multipliers are provider-specific, and some providers apply the discount automatically with no write premium at all.
go deeper
Be able to name the three token classes — uncached input, cache write, cache read — and say that reads are much cheaper while writes cost a little more. Add that output tokens get no discount at all.
Explain that the write premium and read discount are multipliers on the base input price, and that the latency win comes from not re-processing the prefix. Be precise that context-window usage is unaffected.
Show you reason about request shape: prefix size versus variable input versus answer length decides whether caching is worth wiring in at all. Mention that cost and latency are separate benefits with separate conditions.
Own the framing that caching is an input-side lever only. When a workload's spend is dominated by output tokens or by model choice, argue for routing, output-length control or a smaller model rather than presenting caching as the fix.
## The problem caching solves Many production LLM calls repeat a large, identical block of text: a long system prompt, a taxonomy or policy document, a set of few-shot examples, a whole file tree. Without caching the provider charges for and processes that block on every single request, even though it has not changed. Prompt caching lets the provider keep the already-processed form of that prefix for a short time and reuse it, and it prices that reuse very differently from fresh input. ## Three prices, not one Once caching is in play, a request's input tokens fall into three billing classes: - **Uncached input** — the variable part of the prompt, billed at the model's standard input rate (call it 1x). - **Cache write / cache creation** — tokens being stored for reuse. These are billed *above* the standard rate, commonly around 1.25x for a short-lived entry and up to about 2x where a provider offers a longer-lived tier. The premium pays for storing the processed state. - **Cache read** — tokens served from an existing entry. These are billed *well below* standard, frequently in the region of 0.1x. The shape matters more than the exact numbers, which differ per provider and change over time: **you pay a modest surcharge once to store, and a large discount every time you reuse.** Some providers do it implicitly — no explicit write premium, an automatic discount on any prefix they happen to recognise — and others bill storage by duration instead. The reasoning below holds either way; only the arithmetic constants move. ## What a hit does to latency The second, often larger, benefit is time. Before an LLM emits its first token it must process the entire prompt. On a very large prefix that dominates the user-visible wait. When the prefix is served from cache, that processing is largely skipped, so **time-to-first-token drops sharply** — on a six-figure-token prompt this is routinely the difference between a several-second stare at a blank screen and a response that starts almost immediately. ## What a hit does not change Four things stay exactly the same, and confusing them is the classic junior mistake: 1. **Output tokens.** They are never discounted by caching, and they are not generated any faster. The per-token speed of the answer is unchanged. 2. **Context-window consumption.** A cached prefix still occupies its full token count in the window. Caching is a cost-and-latency mechanism, not a compression mechanism. 3. **The model's output.** Caching reuses processed input, not answers; sampling still happens normally, so identical prompts do not suddenly return identical text. 4. **The variable part of the prompt.** Anything that differs between requests is fresh input at full price. ## Where the win is largest and smallest Because the benefit is proportional to the *cached prefix* relative to everything else, the profile of a request predicts the payoff: - **Large stable prefix, small variable input, short answer** — best case. Most of the bill and most of the latency were in the prefix, and both collapse. - **Small prefix, long answer** — near-worthless. The bill and the wall-clock are dominated by generation, which caching does not touch. - **Prefix that changes every call** — worse than useless: you pay the write premium repeatedly and never read. ## How to talk about it A good answer separates the two currencies. Cost and latency are *different* benefits with *different* conditions: the cost benefit needs the prefix to be reused often enough to repay the write premium, while the latency benefit lands on the very first hit and is felt by a single user regardless of how much traffic the system has. Systems have adopted caching purely for the first-token latency and treated the cost reduction as a bonus, and the reverse is equally common in batch pipelines where nobody is waiting. Finally, be precise about *where* the discount applies. Saying "caching makes requests 90% cheaper" is wrong; saying "caching bills the reused prefix at roughly a tenth of input price, so a request that is mostly prefix gets much cheaper" is right, and it is the phrasing that survives follow-up questions.
- Does a cache hit free up room in the context window?No. The cached tokens are still part of the prompt and still count in full against the context limit; only their price and their processing time change. If a prompt is too long for the window, caching does not help — you need to shorten, summarize or offload content by reference instead.
- If output tokens are never discounted, when is caching barely worth enabling?When the answer dominates the request. A 500-token prompt producing a 3,000-token answer is mostly output cost and output time, so discounting a small prefix moves neither the bill nor the wall-clock noticeably. Caching pays when the stable prefix is large relative to the variable input and the generated answer.
- Does a cache read make the model's answers more consistent?No. Caching reuses the processed form of the input, not the output. Sampling still applies, so the same cached prefix with the same variable input can still produce different text. Anyone claiming caching gives determinism has confused prompt caching with response or semantic caching, which stores answers.
saying these in an interview costs you the question
- Thinking caching discounts output tokens too
- Believing cached tokens stop counting toward the context window
- Assuming a cache write costs the same as normal input
- Claiming caching makes responses deterministic
- Saying caching speeds up token-by-token generation