How does prefix-aware routing across LLM replicas raise cache hit rate?
answer
- the cache lives in one GPU's memory
- same conversation, different replica
- hash the leading tokens
- consistent hashing survives a rollout
- affinity fights balance
basics
~20 sA prefix cache lives inside one replica's own KV memory, so a shared preamble or a continuing conversation only skips prefill if the request lands on the replica that already holds it. Prefix-aware routing hashes the leading tokens and steers matching requests to the same replica.
solid answer
~60 sPrefix caching is a per-replica asset: when a server reuses cached KV blocks for a repeated preamble, those blocks are in that GPU's memory and nowhere else. Under round-robin, turn two of a conversation lands on a different replica than turn one, finds nothing cached, and re-prefills the entire history — so the cache hit rate falls roughly as 1/N with fleet size, and TTFT on long shared prefixes stays as bad as a cold request. Prefix-aware routing fixes this by making placement a function of content: hash the first N tokens of the prompt, or the session or tenant identifier, and use consistent hashing so that adding or removing a replica reshuffles only about 1/N of the keys. The tension is that affinity and balance pull against each other — one popular prefix will hotspot a single replica. Practical routers therefore treat the prefix match as a strong preference, not a rule: choose the best-matching replica among those below a load threshold, and fall back to least-busy when the affine replica is saturated.
go deeper
Know that a cached prompt prefix lives on one server, so sending a follow-up message to a different server means the whole conversation is processed again.
Explain the 1/N hit-rate collapse under round-robin and the fix: derive a routing key from the prompt prefix or session id so repeat traffic returns to the same replica.
Show the production shape: consistent hashing so membership changes cost 1/N, an escape hatch to least-busy when the affine replica saturates, and hit-rate plus per-replica queue graphs to prove you did not trade latency for a hotspot.
Decide whether content-aware routing is worth its operational surface for your traffic mix, who owns the router, and how deploys, tenant skew and hot prefixes are handled as policy rather than per-incident.
## Why the cache does not follow the request An engine's prefix cache stores KV blocks computed for a token sequence and reuses them when a later request starts with the same tokens. Those blocks physically occupy the GPU memory of one process. There is no shared cache tier across replicas: replica B cannot use replica A's blocks, and nothing replicates them. Cache locality is therefore entirely a routing property. This matters because the workloads with the most to gain are exactly the ones with long repeated prefixes: - **A shared system prompt or tool schema.** Every request carries the same 2,000-token preamble. - **Few-shot prompting.** The examples are identical across calls. - **Multi-turn chat.** Turn *k* is turn *k−1*'s prompt plus the reply plus the new message — the longest prefix match in the corpus. - **Agent loops.** Each step re-sends the whole trace so far. For these, prefill can be the majority of the request's cost, and a hit turns a multi-second TTFT into a fraction of a second. ## The round-robin failure Spread a conversation's turns over N replicas at random and the probability that turn *k* lands where turn *k−1* did is 1/N. At eight replicas, seven out of eight continuations pay a full prefill of the entire history — and that history grows every turn, so the waste compounds. The cache is enabled, the dashboards show it exists, and it does almost nothing. This is one of the few situations where scaling out makes each individual request *slower*. ## How prefix-aware routing works The router computes a key from the front of the request and maps the key to a replica: - **Key choice.** Hash a fixed-length prefix of the tokenized prompt (block-aligned to the engine's KV block size so partial blocks do not spoil matches), or use a session/conversation id that the client already carries, or a tenant id when tenants share a system prompt. Session ids are cheap and accurate for chat; content hashing generalizes to stateless callers. - **Consistent hashing.** Map keys onto a ring (ring-hash or maglev in an Envoy-class proxy) so that adding or removing a replica moves only about 1/N of keys. Naive modulo-by-replica-count remaps nearly everything on any membership change, which throws away every cache in the fleet during a rollout — exactly when you can least afford it. - **Longest-prefix routing.** More sophisticated routers keep an approximate index of which prefixes each replica has recently served and pick the replica with the longest match, which handles branching conversations better than a single fixed-length hash. ## The tension with balance Pure affinity is a load-balancing policy chosen without reference to load, so it will hotspot. Two shapes recur. If one prefix is enormously popular — a single shared system prompt in front of every request — hashing on it sends everything to one replica. Fix that by including a discriminator (session or tenant) in the key, or by hashing a prefix long enough to differ across callers. And if one tenant is simply much busier than the others, its shard is hot by construction; that tenant may need several replicas in its own hash group. The robust design treats affinity as a preference with an escape hatch: among replicas whose queue depth and KV-cache occupancy are below a threshold, pick the best prefix match; if the affine replica is saturated or preempting, take the least-busy replica and accept the cold prefill. That single rule captures most of the cache benefit while keeping the tail latency behaviour of a load-aware balancer. ## Cache lifetime is not a guarantee Cached blocks are evicted under memory pressure and lost entirely on restart. Routing affinity raises the *probability* of a hit; it never promises one. So never make correctness depend on it — the request must carry the full prompt and be servable by any replica — and be aware that during a rolling deploy your hit rate goes to zero for a while, which shows up as a TTFT spike that is not a regression. ## Measuring it Judge the change on the engine's own prefix-cache hit-rate counters and on TTFT p95 for the affected traffic, broken down per replica. Watch the balance metrics at the same time: if hit rate rises while one replica's queue depth diverges from the rest, you have traded a latency win for a hotspot, and the threshold on the escape hatch is the knob to adjust. Expect the win to be large for long shared prefixes and negligible for short, unique prompts — routing complexity is only worth it when the prefixes are genuinely long and genuinely shared.
- Why use consistent hashing rather than hashing modulo the replica count?Because modulo remaps almost every key whenever membership changes. A single replica added during a scale-out, or removed during a rolling deploy, would invalidate cache locality fleet-wide precisely when load is already elevated. A hash ring moves only about 1/N of keys, so the other N−1 shards keep their warm caches through the change.
- All your traffic shares one 2,000-token system prompt. What does prefix hashing do?It sends everything to one replica, because every key is the same. Add a discriminator — session or tenant id — or hash a longer prefix that includes the per-caller portion. Ironically the fully shared preamble is the easy case: any replica warms it after a handful of requests, so the routing effort should go to the per-conversation tail instead.
- What happens to TTFT during a rolling deploy of a prefix-affine fleet?It spikes. New pods start with empty caches and, if the hash ring changes, some keys move to replicas that never held them. Expect a hit-rate trough and a TTFT bump for the duration, plan the rollout for a low-traffic window, and do not page on it as a regression — surge one pod at a time so only a fraction of keys is cold at once.
saying these in an interview costs you the question
- Assumes the prefix cache is shared across replicas
- Uses modulo hashing over the replica count
- Hashes only a universally shared system prompt
- Treats a cache hit as guaranteed rather than likely
- Ignores the hotspot that pure affinity creates