Between chat turns, should the server keep a session's KV blocks resident or recompute?
answer
- think time is idle GPU memory
- free but keep, do not pin
- evictable beats reserved
- degrade to baseline, not to a stall
- routing decides if the cache is reachable
basics
~20 sPinning blocks through a user's think time blocks capacity for every idle session, so it only pays for short gaps and high-value sessions. The usual answer is neither extreme: free the blocks but leave them in the prefix cache, where they are reused if still present and recomputed if not.
solid answer
~60 sHolding a session's blocks between turns converts think time — often tens of seconds — into occupied GPU memory, and memory is the term that sets concurrency. A thousand idle sessions each pinning a 6k-token cache will out-consume the sessions actually generating, so pinning is a capacity decision disguised as a latency optimization. The middle path is what modern engines already do: release the blocks on completion but keep them registered in the prefix cache, unreferenced and evictable in least-recently-used order. A quick follow-up hits a warm cache and skips prefill; a slow one finds the blocks reclaimed and pays a normal prefill. You get the best case free and degrade to the baseline rather than to a stall, with no capacity held hostage. Beyond that, the levers are offloading blocks to host or external KV storage — trading transfer against prefill compute, which only wins for long prefixes — and shortening the resent history so the recompute is smaller. Note that whether the follow-up even reaches the replica holding the cache is a routing question, and a warm cache with round-robin routing is mostly wasted.
go deeper
Know that a chat turn resends the whole conversation, so the server can sometimes reuse the earlier turn's cached work instead of recomputing it.
Be able to say why holding blocks between turns is expensive — idle memory is memory no other request can use — and that engines instead free the blocks while keeping them reusable until evicted.
Show the arithmetic: think time against generation time sets the fraction of pinned memory that idles, and a strategy whose latency target depends on a cache hit will miss it exactly at peak.
Own the ladder and the couplings — pin, keep-and-evict, offload, shorten history — plus the dependency on cache-affine routing and the fact that the scheduler will always prefer a request generating now over a speculative future turn.
## What the decision actually is A multi-turn chat resends the whole conversation on every turn. Turn five's prompt is turn four's prompt plus the assistant's answer plus the new user message — so a large, exactly-matching prefix is guaranteed. Between turns the user is reading and typing, often for tens of seconds. The question is what the server should be holding during that gap. The options form a ladder, and the interviewer wants the ladder, not a single answer. ## Option 1: pin the blocks Keep the session's blocks allocated and referenced so the next turn resumes instantly. Time-to-first-token on turn two is near-zero prefill. The cost is unforgiving. Those blocks are unavailable to every request that arrives during the gap. Model the ratio: if the average think time is 30 seconds and the average turn takes 4 seconds to generate, then roughly 88% of a pinned session's residency is idle. Serve a thousand concurrent sessions and you are sizing GPU memory for a thousand caches while only a hundred-odd are producing tokens. You have reinvented the worst-case reservation that paging existed to kill — the reservation is now over time instead of over length. Pinning is defensible only in narrow cases: a small, bounded number of sessions; a hard first-token commitment; think times measured in a second or two; or a workload where the traffic is sparse enough that memory is genuinely free. ## Option 2: free and rely on the prefix cache On completion, drop the reference count. The blocks are no longer owned by anyone, but they remain registered under their content hashes and stay usable until the memory is actually needed, at which point they are evicted in least-recently-used order. This is the default behaviour of prefix-caching engines. The economics are strictly better than pinning under load, because the memory is *lendable*. In a quiet minute the follow-up hits warm; in a busy minute the blocks were reclaimed by requests that needed them right then, and the follow-up pays a prefill it would have paid anyway. The failure mode is graceful degradation to the baseline, not a stall — which is exactly the property you want the system to have when it is busy. The design work here is mostly prompt layout: keep the fixed system preamble first so it survives as a shared, hot prefix across all sessions even when per-session tails are evicted. ## Option 3: offload the blocks off-GPU Copy a finished session's blocks to host memory or an external KV store and pull them back on the next turn. vLLM exposes connector plumbing for this through `--kv-transfer-config`. The trade is transfer time against prefill time, and it only wins when the prefix is long enough that recomputing it costs more than moving it — plus you now own a cache tier with its own capacity, eviction and consistency questions. Treat it as an optimization for long-document or long-history workloads with measurable prefill cost, not as a default. ## Option 4: shrink what must be rebuilt The cheapest recompute is a shorter prompt. Summarizing or truncating old turns reduces both the prefill on a miss and the memory on a hit. It changes model behaviour, so it is a product decision, but it is often the largest lever and the one candidates forget because it lives outside the serving layer. ## The coupling people miss All of this assumes the follow-up lands on the replica whose memory holds the state. Cache affinity across replicas is a routing concern, and without it a fleet of eight replicas gives a warm-cache probability of one in eight. Any answer that optimizes residency without acknowledging that dependency is incomplete. The second coupling is with preemption. Cached-but-unreferenced blocks are precisely the memory the scheduler reclaims first when running sequences need space. So the value of option 2 is highest exactly when load is lowest, and near zero at saturation. That is the correct priority — a request generating tokens now should always outrank a speculative future turn — and it means the honest way to state the strategy is: free reuse when we have room, progress for live requests when we do not. ## How to decide with numbers Measure think-time distribution, cache bytes per session, arrival rate, and the prefill cost of a full history. Compute what fraction of memory pinning would idle at your concurrency, and what prefix-cache hit rate you already achieve at peak. If hit rate at peak is already decent, pinning buys little and costs a lot. If prefill dominates end-to-end latency and histories are very long, that is the case where offloading earns its complexity.
- What measurement would make you pin blocks between turns after all?A think-time distribution short relative to generation time, a bounded session count you can size memory for, and a first-token commitment the prefix cache misses at peak. If sessions think for two seconds and generate for six, pinning idles only a quarter of the residency and may be worth it. If they think for a minute, the arithmetic never works at any interesting concurrency.
- Why is offloading KV blocks to host memory not simply a free win?Because you pay transfer both ways over a bus shared with everything else, and you take on a second cache tier with its own capacity and eviction policy. It wins only when the prefix is long enough that recomputing it would cost more GPU time than moving it costs in bandwidth and latency. For short chat histories, prefill is cheap and batched, and the transfer is pure overhead.
- How does this interact with the scheduler under load?Directly and correctly: unreferenced cached blocks are the first memory reclaimed when running sequences need space. So reuse is most available when the server is quiet and least available at peak. Any residency strategy has to assume its benefit disappears under load, which is another argument against designs whose latency target depends on a hit.
saying these in an interview costs you the question
- Treats pinned session cache as free because the GPU is idle anyway
- Ignores think time when sizing per-session memory
- Assumes the follow-up request reaches the same replica
- Proposes an external KV store before measuring prefill cost
- Says never evict a session's cache, to guarantee low first-token latency