How do you keep a 200-worker agent fleet reusing one cached 30k-token prefix?
answer
- one write, then many reads
- who holds the entry decides everything
- cold fan-out races the first write
- warm once with identical bytes
- affinity costs you load balance
basics
~20 sWarm the prefix with one request before fanning out, then shape traffic so related work reaches whatever holds the entry: batch requests that share a prefix, and on self-hosted replicas add prefix-aware routing. Accept that affinity trades load balance for reuse.
solid answer
~50 sTwo questions decide the design: **who holds the entry**, and **when does it become readable**. With a hosted provider the entry is scoped to your account, so all 200 workers can share it — but it only exists once some request has written it. Fire all 200 cold requests simultaneously and they race: many pay the write and few get a read. The fix is a warm-up — one request carrying the exact prefix, completed before the burst — then fan out. With self-hosted serving the prefix cache lives on a specific replica, so a round-robin balancer scatters the fleet across replicas and each one re-warms independently. There you need prefix-aware (cache-affinity) routing: hash the reusable prefix and steer requests carrying that hash to the replica that already holds it. Affinity is not free. It fights load balancing, creates hot replicas for popular prefixes, and needs a fallback when the target is saturated or draining. Route on prefix first and fall back to least-loaded past a utilisation threshold.
go deeper
Know that a cached prefix has to be written by some request before any other request can reuse it, so the very first call of a burst cannot hit.
Explain the cold-stampede effect and the warm-up fix: one completed request carrying the byte-identical prefix, then fan out. Say why the warm-up must be built by the same code the fleet uses.
Separate account-scoped provider caches from replica-local self-hosted caches, and show that round-robin balancing destroys reuse in the second case. Describe prefix-aware routing with a utilisation-based fallback and what you would measure to confirm it works.
Own the tradeoff: affinity turns a stateless tier stateful, complicating drain, autoscale and failover, and can concentrate load on a hot prefix. Be able to argue when warm-up plus prefix stability is sufficient and affinity is over-engineering, using measured re-warm rates rather than intuition.
## Frame the problem correctly first "200 workers share a 30k-token policy prefix" is not one problem. It is two, and conflating them is the most common way to answer this badly. **Who holds the cached prefix?** With a hosted provider, the entry is generally scoped to your account and model configuration, so any worker's request can match a prefix another worker created — the fleet is already sharing. With self-hosted or replica-based serving, the processed prefix lives in the memory of one serving replica. Nothing is shared implicitly; a request that lands on a different replica finds nothing. **When does the entry become readable?** In both worlds, an entry has to be written before it can be read. A cold fan-out of 200 simultaneous requests, all carrying a prefix that has never been sent, has no entry to match — they race, several pay the write, and the reads only start arriving for whatever fraction is scheduled after a write completes. ## The cold stampede and the warm-up call The practical consequence: the shape of your traffic ramp determines your hit rate on the first minute of a burst. If a scheduled job wakes 200 workers at the top of the hour on a freshly deployed prompt, that first wave is largely uncached, and you will see a spike of cache-creation tokens followed by a normal read pattern. A warm-up call fixes it cheaply. Before releasing the burst, send **one** request carrying the exact reusable prefix and wait for it to complete. It can be a trivial request — the point is the write, not the answer. Then fan out. The rules that make this work in practice: - The warm-up prefix must be **byte-identical** to what production sends. A warm-up built by different code, or with a slightly different tool order, warms an entry nobody will match — a classic and embarrassing failure. - It must **complete** before the fan-out; firing it concurrently defeats the purpose. - It must be **repeated** whenever the prefix changes: after each deploy that touches the prompt, and periodically if a route's traffic is sparse enough that entries go idle between bursts. - Ramping — a small first wave, then full concurrency — achieves the same effect without a dedicated call, and is a good default for autoscaling workers. ## Routing and affinity when the cache is replica-local On self-hosted serving, prefix reuse is a *scheduling* problem. Standard round-robin or least-connections balancing actively destroys it: the same logical prefix gets warmed independently on every replica, multiplying prefill work by the replica count and wasting the memory that holds those entries. Prefix-aware routing (often called cache-aware routing) is the answer: compute a hash of the reusable prefix at the client or router, and prefer the replica that has recently served that hash. This is the same shape as consistent hashing in any sharded cache, with the same costs: - **Hot shards.** A popular prefix concentrates load on one replica. Allow a prefix to spread across a small set of replicas rather than exactly one, sized by that prefix's traffic. - **Balance versus reuse.** Strict affinity can queue requests behind a busy replica while others idle. Route on prefix *until* the target crosses a utilisation threshold, then fall back to least-loaded and accept a cold call. Reuse is an optimisation; latency SLOs are not. - **Churn.** Scale-up, scale-down, restarts and drains all move entries. The router must tolerate a miss gracefully rather than pinning to a replica that no longer holds anything. ## Batching and grouping related work Even where the cache is account-scoped, grouping still pays. Requests sharing a prefix should be run close together in time rather than interleaved with unrelated work, because entries are kept alive by use — sparse, scattered traffic on a prefix lets it go cold between calls, and each revival costs a write. Concretely: sort or partition a queue by prefix identity so a worker pool drains one prefix family before moving to the next, and prefer a dedicated pool per prompt family over a general pool that mixes them. This is also why a fleet with three prompt variants running A/B often shows worse aggregate hit rate than either variant alone would — the traffic is split three ways. ## What to measure and when to stop Instrument cache-read versus cache-creation tokens per route and per replica, and watch creation spikes at burst boundaries — that is the signature of an unwarmed ramp. On self-hosted serving, also watch the distribution of a given prefix hash across replicas; if one hash appears on every replica, affinity is not working. And be willing to conclude it is not worth it. Affinity routing adds a stateful component to what was a stateless tier, complicates draining and autoscaling, and can degrade tail latency. For a small prefix, or a fleet whose traffic is dense enough that entries stay warm naturally, plain warm-up and prefix stability capture most of the benefit with none of the operational cost. Reach for affinity when the prefix is large, the fleet is wide, and the measurement shows replicas warming the same prefix repeatedly.
- A scheduled job wakes all 200 workers at once each hour and you see a cache-creation spike every hour. What do you change?Warm before the burst. Send one request carrying the byte-identical prefix, wait for it to complete, then release the fleet — or ramp, starting with a handful of workers and scaling to full concurrency once the first responses return. The hourly spike is the fleet racing to write the same entry because none existed when the wave started. Either approach converts most of that creation volume into reads.
- What is the risk in a warm-up call that is built by separate code from the production path?It warms an entry production never matches. Any divergence — a different tool order, a different serializer, an extra newline, a slightly different system string — produces a different prefix, so you pay for the warm-up, the fleet still fans out cold, and the metrics look like the warm-up simply did not work. Build the warm-up request through the same prefix-assembly function production uses, and assert the two hashes are equal.
- When would you decline to add prefix-affinity routing despite a large shared prefix?When the cost lands on availability rather than spend. Affinity makes a stateless serving tier effectively stateful: draining, autoscaling and failover all get harder, and a hot prefix can queue behind one busy replica while others idle. If traffic is dense enough to keep entries warm naturally, or the fleet is small, warm-up plus a stable prefix captures most of the benefit. Add affinity only once measurement shows replicas repeatedly re-warming the same prefix.
saying these in an interview costs you the question
- Assuming sticky routing is free of balance and failover costs
- Firing all workers at once and expecting a single cache write
- Believing a hosted provider's cache is scoped per process
- Warming with a prefix that differs slightly from production's
- Pinning a hot prefix to one replica with no overflow path