Why does a prompt-cache hit cut time-to-first-token but not total time?
answer
- two waits, one request
- only the wait before output changes
- tokens still stream at the same pace
- big prompt, small answer wins most
- perceived speed beats measured speed
basics
~20 sA cache hit removes the work of reading the reused prefix, which is all the wait that happens before the first token. Generating the answer afterwards is untouched, so a long response still takes the same time to stream out.
solid answer
~50 sAn LLM request has two distinct waits: the model must first take in the whole prompt, then produce the answer one token at a time. A cache hit eliminates most of the first wait for the reused portion and none of the second. On a 100k-token codebase prompt this can take time-to-first-token from around four seconds to under one, while a 2,000-token answer still streams at exactly the same per-token pace afterwards. So total wall-clock improves only by the amount the prompt-reading phase contributed. When the prompt is huge and the answer short — classification, extraction, a one-line verdict over a big document — that is nearly the whole request and the speedup is dramatic. When the prompt is small and the answer long, caching barely moves the clock even on a perfect hit. Quote the two metrics separately; conflating them is the usual mistake.
go deeper
Remember that a cache hit means the model does not have to read the reused prompt again, so the answer starts sooner. Say clearly that the answer itself is not produced any faster.
Separate time-to-first-token from the streaming phase and explain that caching only touches the first. Use the ratio rule: the wall-clock gain is roughly the share of the request spent waiting before output began.
Show you would instrument TTFT separately from completion time, and pick workloads by shape — large prompt with a short answer for the latency win, high-reuse batch for the cost win. Be explicit that the two benefits have different preconditions.
Own the decision of which user-facing paths justify caching on responsiveness alone, and when a latency complaint is really an output-length or model-choice problem that caching cannot solve.
## Two clocks, not one Every LLM request contains two sequential waits, and a candidate who does not separate them cannot reason about caching's latency benefit at all. 1. **Before the first token.** The model has to take in the entire prompt. This work grows with prompt length, and while it is happening the user sees nothing. It is what **time-to-first-token (TTFT)** measures. 2. **While the answer streams.** Tokens are produced one after another. The pace here — **inter-token latency** — is essentially independent of how long the prompt was, and it multiplied by the answer length gives the rest of the wall-clock. A cache hit acts on the first clock only. The reused prefix has already been taken in, so the provider does not repeat that work; the fresh part of the prompt and the entire generation phase are unchanged. ## The numbers that make it concrete Take a coding assistant that ships a 100k-token snapshot of a repository as a stable prefix. Cold, TTFT can sit around four seconds — a very visible pause. On a hit it drops to well under a second, because the expensive part of that wait has been removed. Now look at total time for two different requests over the same cached prefix: - "Which file defines the retry policy?" → a 30-token answer. Nearly the whole request was the pause before the first token, and total time collapses roughly in proportion. Users describe this as the assistant becoming instant. - "Refactor this module and explain each change." → a 3,000-token answer. The pause was four seconds; the streaming is tens of seconds. Removing the pause is a real but modest fraction of the total, and a stopwatch on end-to-end completion barely notices. Same cache, same hit, wildly different headline improvement — because the ratio of prompt to answer differs. ## Why perceived latency improves more than the stopwatch says Even in the second case the *experience* improves out of proportion to the measurement. A streaming interface that begins producing text immediately feels responsive; the same total duration preceded by four seconds of nothing feels broken. That is why interactive products adopt caching for TTFT alone, quite separately from any cost argument, and why you should report TTFT and total completion time as two metrics rather than averaging them into one "latency" number. ## Where the latency benefit is worthless - **Batch and offline jobs.** Nobody is waiting. A nightly pipeline cares about throughput and cost; the TTFT improvement is irrelevant to it, though the cost saving very much is not. - **Short prompts.** If the prefix is small, the pre-first-token wait was already short, so there is nothing meaningful to remove. - **Output-dominated requests.** Long-form generation is bounded by streaming, which caching does not touch. If a request is slow because the model is writing a lot, the fix is shorter output, a faster model or better streaming UX — not caching. ## The important asymmetry with the cost benefit The cost and latency benefits of caching are governed by different conditions, and confusing them is a common error: - The **cost** benefit needs *repeated* traffic — a write must be repaid by reads before the entry disappears. - The **latency** benefit lands in full on the *very first hit* and is felt by a single user in a single session. One developer alone with an assistant, re-asking questions over the same cached repository snapshot, gets no meaningful fleet-wide cost story but an enormous interactive improvement. This is why some teams enable caching for user-facing paths purely on responsiveness grounds and treat the discount as incidental, while batch teams do the exact opposite. ## How to answer well Name the two clocks, say plainly that caching removes prompt-reading work and not generation work, then give the ratio rule: the wall-clock improvement approximates the share of the request that was spent waiting before the first token. Close with the workload split — huge prompt with a small answer is where caching transforms latency, and small prompt with a long answer is where it does almost nothing. Avoid claiming that caching makes the model "faster"; it makes it *start* faster.
- Your p50 end-to-end latency barely moved after enabling caching, but users say it feels much faster. How do you explain that to a sceptical PM?End-to-end time is dominated by streaming the answer, which caching does not affect. What changed is the dead time before the first visible token, which fell sharply. Report TTFT as its own metric alongside total completion time — a single averaged latency number hides exactly the improvement users are describing.
- Which workloads gain the cost benefit of caching but not the latency benefit?Offline and batch pipelines. A nightly job over thousands of records reuses one prefix constantly, so the cost saving is large, but no human is waiting for the first token, so faster TTFT buys nothing. Judge such systems on cost per record and throughput, not on responsiveness.
- Does a cache hit reduce the per-token speed of the streamed answer?No. Generation proceeds token by token at essentially the same pace regardless of whether the prompt was cached. Caching removes work that happens before generation starts, so it affects when streaming begins, never how quickly it proceeds once it has.
saying these in an interview costs you the question
- Claiming caching makes the model generate tokens faster
- Reporting a single averaged latency number that hides the TTFT gain
- Expecting big wall-clock gains on output-heavy requests
- Pursuing caching for latency in an offline batch pipeline
- Assuming the cost and latency benefits require the same conditions