skip to content

Prompt Caching

Getting the provider to reuse a prefix you already sent: how cache breakpoints work, how long entries live, how to structure a prompt so it hits, and what a hit is worth in cost and first-token latency. Interviewers ask because it is one of the largest cheap wins available in a production LLM system.

on this pageshow

explore

questions

22

What does a prompt-cache hit change about an LLM request's cost and speed?

level: juniorimportance: must knowfreq 68%

answer

  1. three prices, not one
  2. the store costs more than sending
  3. reuse costs roughly a tenth
  4. output tokens are untouched
  5. first token arrives sooner

basics

~20 s

A hit bills the reused prefix at a fraction of the normal input price — often around a tenth — and the model skips re-processing it, so the first token arrives far sooner. Output tokens are billed and generated exactly as usual.

solid answer

~50 s

Prompt caching splits input into three price tiers instead of one: ordinary uncached input at the base rate, cache-write (creation) tokens at a **premium** above base — commonly 1.25x to 2x depending on how long the entry is kept — and cache-read tokens at a deep **discount**, often around a tenth of base. A hit therefore buys two separate things. Financially, the reused prefix becomes nearly free. Operationally, the provider does not have to re-process those tokens, which is what cuts time-to-first-token on a long prompt. Nothing else improves: output tokens are billed at full rate and produced at the same speed, and the cached tokens still occupy the context window. The size of the win scales with how big the reused prefix is relative to the fresh input and the generated answer. Exact multipliers are provider-specific, and some providers apply the discount automatically with no write premium at all.

go deeper

for a junior

Be able to name the three token classes — uncached input, cache write, cache read — and say that reads are much cheaper while writes cost a little more. Add that output tokens get no discount at all.

for a middle

Explain that the write premium and read discount are multipliers on the base input price, and that the latency win comes from not re-processing the prefix. Be precise that context-window usage is unaffected.

for a senior

Show you reason about request shape: prefix size versus variable input versus answer length decides whether caching is worth wiring in at all. Mention that cost and latency are separate benefits with separate conditions.

for a principal

Own the framing that caching is an input-side lever only. When a workload's spend is dominated by output tokens or by model choice, argue for routing, output-length control or a smaller model rather than presenting caching as the fix.

## The problem caching solves Many production LLM calls repeat a large, identical block of text: a long system prompt, a taxonomy or policy document, a set of few-shot examples, a whole file tree. Without caching the provider charges for and processes that block on every single request, even though it has not changed. Prompt caching lets the provider keep the already-processed form of that prefix for a short time and reuse it, and it prices that reuse very differently from fresh input. ## Three prices, not one Once caching is in play, a request's input tokens fall into three billing classes: - **Uncached input** — the variable part of the prompt, billed at the model's standard input rate (call it 1x). - **Cache write / cache creation** — tokens being stored for reuse. These are billed *above* the standard rate, commonly around 1.25x for a short-lived entry and up to about 2x where a provider offers a longer-lived tier. The premium pays for storing the processed state. - **Cache read** — tokens served from an existing entry. These are billed *well below* standard, frequently in the region of 0.1x. The shape matters more than the exact numbers, which differ per provider and change over time: **you pay a modest surcharge once to store, and a large discount every time you reuse.** Some providers do it implicitly — no explicit write premium, an automatic discount on any prefix they happen to recognise — and others bill storage by duration instead. The reasoning below holds either way; only the arithmetic constants move. ## What a hit does to latency The second, often larger, benefit is time. Before an LLM emits its first token it must process the entire prompt. On a very large prefix that dominates the user-visible wait. When the prefix is served from cache, that processing is largely skipped, so **time-to-first-token drops sharply** — on a six-figure-token prompt this is routinely the difference between a several-second stare at a blank screen and a response that starts almost immediately. ## What a hit does not change Four things stay exactly the same, and confusing them is the classic junior mistake: 1. **Output tokens.** They are never discounted by caching, and they are not generated any faster. The per-token speed of the answer is unchanged. 2. **Context-window consumption.** A cached prefix still occupies its full token count in the window. Caching is a cost-and-latency mechanism, not a compression mechanism. 3. **The model's output.** Caching reuses processed input, not answers; sampling still happens normally, so identical prompts do not suddenly return identical text. 4. **The variable part of the prompt.** Anything that differs between requests is fresh input at full price. ## Where the win is largest and smallest Because the benefit is proportional to the *cached prefix* relative to everything else, the profile of a request predicts the payoff: - **Large stable prefix, small variable input, short answer** — best case. Most of the bill and most of the latency were in the prefix, and both collapse. - **Small prefix, long answer** — near-worthless. The bill and the wall-clock are dominated by generation, which caching does not touch. - **Prefix that changes every call** — worse than useless: you pay the write premium repeatedly and never read. ## How to talk about it A good answer separates the two currencies. Cost and latency are *different* benefits with *different* conditions: the cost benefit needs the prefix to be reused often enough to repay the write premium, while the latency benefit lands on the very first hit and is felt by a single user regardless of how much traffic the system has. Systems have adopted caching purely for the first-token latency and treated the cost reduction as a bonus, and the reverse is equally common in batch pipelines where nobody is waiting. Finally, be precise about *where* the discount applies. Saying "caching makes requests 90% cheaper" is wrong; saying "caching bills the reused prefix at roughly a tenth of input price, so a request that is mostly prefix gets much cheaper" is right, and it is the phrasing that survives follow-up questions.

  • Does a cache hit free up room in the context window?
    No. The cached tokens are still part of the prompt and still count in full against the context limit; only their price and their processing time change. If a prompt is too long for the window, caching does not help — you need to shorten, summarize or offload content by reference instead.
  • If output tokens are never discounted, when is caching barely worth enabling?
    When the answer dominates the request. A 500-token prompt producing a 3,000-token answer is mostly output cost and output time, so discounting a small prefix moves neither the bill nor the wall-clock noticeably. Caching pays when the stable prefix is large relative to the variable input and the generated answer.
  • Does a cache read make the model's answers more consistent?
    No. Caching reuses the processed form of the input, not the output. Sampling still applies, so the same cached prefix with the same variable input can still produce different text. Anyone claiming caching gives determinism has confused prompt caching with response or semantic caching, which stores answers.

saying these in an interview costs you the question

  • Thinking caching discounts output tokens too
  • Believing cached tokens stop counting toward the context window
  • Assuming a cache write costs the same as normal input
  • Claiming caching makes responses deterministic
  • Saying caching speeds up token-by-token generation

context

open as a page

In prompt caching, how many cache reads repay one cache write?

level: middleimportance: must knowfreq 74%

basics

~20 s

Break-even reads equal the write premium above base divided by the discount below base: with a 1.25x write and a 0.1x read, a single reuse already pays; at a 2x write it takes about two. The real constraint is whether that reuse arrives before the entry expires.

open as a page

Why does a timestamp at the top of a system prompt destroy cache hit rate?

level: middleimportance: must knowfreq 74%

basics

~20 s

Prompt caches match an exact prefix from the first token onward. A line like "Current time: 14:32:07" differs on every request, so the match fails at the very first tokens and the entire prompt behind it is reprocessed and re-stored, forever.

open as a page

Why does editing one word early in a long system prompt void the whole cache?

level: middleimportance: must knowfreq 78%

basics

~20 s

Attention is causal: every later token's cached key and value tensors were computed from all the tokens before it. Change a token near the top and the stored state for everything after it is stale, so the server must rebuild the entire prompt.

open as a page

In prompt caching, what is physically stored for a cached prefix?

level: middleimportance: must knowfreq 72%

basics

~20 s

The per-layer attention key and value tensors for every token in the prefix — not the prompt text, not an embedding, and not the model's answer. A hit reloads those tensors and skips recomputing them.

open as a page

In prompt caching, why must stable content come before per-request content?

level: middleimportance: must knowfreq 70%

basics

~20 s

Prompt caching matches a new request against a stored prompt from the very first token onward, and the first difference ends the match. Anything that varies per request must therefore sit after every block you want reused.

open as a page

Does a prompt cache's TTL reset on each cache hit, or expire a fixed time after the write?

level: middleimportance: must knowfreq 62%

basics

~20 s

Prompt caches typically use a sliding window: every read of a cached prefix restarts its expiry at the full TTL. A continuously used prefix stays warm indefinitely, and only an idle gap longer than the window forces a fresh, full-price write.

open as a page

Which deploy-time changes silently invalidate every cached prompt prefix?

level: seniorimportance: must knowfreq 52%

basics

~20 s

A cache entry is tied to one exact token prefix on one exact model version. Bumping the model, editing the system prompt, or adding, removing or reordering a tool definition makes every existing entry unreachable, so the first request after deploy pays a full write.

open as a page

How do you measure prompt-cache hit rate in a production LLM service?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Measure it in tokens, not requests. Providers report per-response counters for tokens read from cache and tokens written to it, so hit rate is cache-read tokens over the reusable prefix. Chart it per route and per deploy.

open as a page

Why does a prompt-cache hit cut time-to-first-token but not total time?

level: middleimportance: should knowfreq 56%

basics

~20 s

A cache hit removes the work of reading the reused prefix, which is all the wait that happens before the first token. Generating the answer afterwards is untouched, so a long response still takes the same time to stream out.

open as a page

How many cache breakpoints do you set, and at which layers of a prompt?

level: middleimportance: should knowfreq 48%

basics

~20 s

Set one breakpoint per stability tier, not per block — commonly after tool definitions, after the static documents and examples, and after the last completed conversation turn. Explicit breakpoints are capped at a handful, so spend them on real change boundaries.

open as a page

Besides TTL expiry, what can drop a cached prompt prefix before its window ends?

level: middleimportance: should knowfreq 38%

basics

~20 s

Prompt caching is best-effort. Entries can be dropped early under capacity pressure, a request can land on serving capacity that never held the entry, and caches are scoped to your own account, so an identical prompt from another customer never helps you.

open as a page

When does enabling prompt caching make a workload more expensive?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Whenever a written prefix is never reused before it expires. Single-shot calls over unique documents, prompts whose stable portion changes every request, and traffic arriving slower than entries survive all pay the write premium and collect no discount — typically about 25% worse than not caching.

open as a page

How does non-deterministic JSON serialization of tool definitions cause cache misses?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A prompt cache matches bytes, not meaning. If tool schemas are serialized from a map with unstable key or element order, two logically identical prompts produce different token sequences, so one logical prefix splinters into many distinct entries that rarely match.

open as a page

Why does prompt caching cut time-to-first-token but not tokens-per-second?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Caching removes prefill — the single parallel pass that builds key/value state for the prompt. Decoding still runs one sequential forward pass per generated token, and those steps are untouched, so the answer starts sooner but streams at the same rate.

open as a page

How does a serving stack find the longest cached KV prefix for a request?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It matches token ids, not text. The incoming token sequence is walked through a prefix structure — a radix tree of shared spans, or a chain of hashes over fixed-size token blocks — which returns the longest stored prefix; matching halts at the first differing token or block.

open as a page

How do you design a multi-turn chat prompt so each turn reuses the cached prefix?

level: seniorimportance: should knowfreq 56%

basics

~20 s

Treat the conversation as append-only: never rewrite, reorder or drop earlier turns, and move the cache breakpoint to the end of each completed turn so the growing history becomes the stable prefix that the next request reuses.

open as a page

How do you keep a 200-worker agent fleet reusing one cached 30k-token prefix?

level: principalimportance: should knowfreq 34%

basics

~20 s

Warm the prefix with one request before fanning out, then shape traffic so related work reaches whatever holds the entry: batch requests that share a prefix, and on self-hosted replicas add prefix-aware routing. Accept that affinity trades load balance for reuse.

open as a page

How much GPU memory does a long shared KV prefix pin, and what does it displace?

level: principalimportance: should knowfreq 30%

basics

~20 s

Roughly 2 x layers x key/value heads x head width x tokens x bytes per value. For a 70B-class model that is about 320 KB per token — near 9 GB for a 30,000-token prefix — and it comes out of the same pool serving in-flight requests, so retention trades hit rate against concurrency.

open as a page

A document-QA cache with a five-minute window expires before almost every turn because users return after 40 minutes — longer-TTL tier, keep-warm traffic, or neither?

level: principalimportance: should knowfreq 34%

basics

~20 s

Start from the measured gap distribution, not the average. A longer window pays when a real share of follow-ups land inside it; keep-warm heartbeats pay only when one warm prefix serves many users; with per-user prefixes and 40-minute gaps, the honest answer is often neither.

open as a page

How do you measure the savings prompt caching actually delivered in production?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Use the per-request token counts providers report, which separate cache-creation from cache-read from ordinary input. Convert each to base-input-token equivalents with its price multiplier, compare against what the same traffic would have cost uncached, and track the result per request class.

open as a page

How do you resolve cache-friendly prompt ordering against putting instructions last?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

Split the instruction from its bulk. Keep the large stable blocks — persona, tools, corpus, examples — in the cached head, and repeat only a short restatement of the task after the variable content. Duplicating a few dozen tokens beats fragmenting the prefix.

open as a page