skip to content

Why can a low-traffic site get slower responses when its route runs at fifty locations instead of one region?

level: middleimportance: should knowfreq 44%

answer

  1. each location keeps its own warm state
  2. nothing is shared between locations
  3. divide traffic by number of locations
  4. hit ratio depends on density, not totals

basics

~20 s

Caches and warm state belong to a location and are never shared. Traffic split across fifty locations gives each a fiftieth of the requests, so most hit a cold cache that one busy region would have served warm.

solid answer

~50 s

Running a route in many places multiplies the number of caches that have to be warmed and divides the traffic that warms them. A response cached at one location does nothing for a request that lands at another, so with a fixed freshness window the hit ratio depends on requests *per location*, not on total traffic. A site taking one request a second worldwide, spread over fifty locations and cached for a minute, has roughly one request per location per minute — near enough to the window that most requests miss. The same site in one region would serve almost all of them warm. The effect inverts at high volume: once every location sees steady traffic, each stays warm and near-user placement wins on both distance and hit ratio. So the value scales with traffic density per location, not with the number of locations.

go deeper

for a junior

The point to retain: caches live at each location separately. Spreading a site with little traffic over many locations means most requests find nothing cached.

for a middle

Be ready to do the arithmetic out loud — total requests divided by locations, compared against the freshness window — and to say when the effect inverts at high volume.

for a senior

Show that you measure the hit ratio per location rather than globally, and that you would narrow the spread or prerender before blaming the cache configuration.

for a principal

Treat spread as a parameter to choose, not a default to accept: match the number of locations to where your audience actually is, and revisit it as traffic shifts.

## Warm state belongs to a location When a route runs in one region, everything it accumulates accumulates in one place: cached responses, cached data reads, warmed-up connections, anything held between requests. When the same route runs at fifty locations, each location has its own copy of all of that and none of it is shared. There is no mechanism by which the response one location computed becomes available to a request that arrives at another, and there is no reason to expect there to be one — the locations are independent execution sites that happen to run the same code. That independence is the whole story behind an unintuitive result: **for a low-traffic site, more locations can mean slower responses.** ## The hit-ratio arithmetic A cached entry survives for some freshness window. Whether the *next* request finds it depends on whether that next request arrives at the same location inside the window. So the number that matters is requests per location per window, not requests per site. | total traffic | locations | requests per location per minute | a one-minute cache is… | |---|---|---|---| | 1 req/s | 1 | 60 | warm almost always | | 1 req/s | 50 | ~1.2 | cold most of the time | | 100 req/s | 50 | ~120 | warm almost always | The left column did not change between rows two and three; the placement did not change either. Only the density per location changed, and it is what decides whether caching is doing any work at all. Two related costs ride along with the same mechanism: - Whatever a location has to do the first time it serves a route, it does **per location** rather than once. Spreading the route wider means paying that first-time cost more times. - Traffic is not evenly spread. Locations near your real audience stay warm; the ones covering geographies you barely serve are permanently cold, and a request that lands there gets the worst of both worlds — a cold path *and* a long hop to the data. ## Who this hurts and who it helps - **Hurt:** a thin, globally scattered audience; long freshness windows that still cannot outlast the gaps between requests; responses that are expensive to produce, so a miss is costly. - **Helped:** heavy, steady traffic in every geography served, where every location stays warm and the shortened user hop is pure profit. - **Indifferent:** request-only handlers. A redirect or a header check has no warm state to lose, which is one more reason those are the routes that move near the user most comfortably. ## What to do about it 1. **Judge the hit ratio per location.** A single global hit-ratio number hides the whole effect; break it down by location before concluding the cache is configured badly. 2. **Narrow the spread.** Where the platform allows choosing a subset of locations, serving from a handful near the real audience keeps each of them warm and still shortens the hop for most users. 3. **Lengthen the window where correctness allows it.** The window has to outlast the gap between requests at a single location, which at low traffic can mean minutes rather than seconds. 4. **Prerender what can be prerendered.** A route built ahead of time has nothing to warm: the response already exists and is served as a file, so the per-location cold-cache problem disappears for it entirely. 5. **Leave the expensive, rarely requested routes in the region.** A route that is costly to produce and seldom asked for is the worst possible thing to scatter. ## The rule this leaves you with Near-user placement pays for itself through two effects — a shorter user hop and cheap repeat responses — and only the first is guaranteed. The second is conditional on there being enough traffic at each location to keep it warm. Meta-frameworks and platforms differ in how much control they give here: some spread a deployment as widely as possible by default, some let you name the locations, some settle it per route. Whatever the control, the question to ask before widening is the same: how many requests will each location actually see inside the freshness window?

  • Does prerendering a route at build time avoid this problem?
    Yes, for that route. A prerendered response already exists before any request arrives, so there is no per-location work to warm up and no freshness window to outlast — the file is simply served from wherever it has been distributed. The problem only applies to responses that some location has to compute on demand and then keep.
  • How would you tell a cold-cache problem from a slow-handler problem?
    Break the latency down by location and by hit or miss. A cold-cache problem shows a low hit ratio per location with fast hits and slow misses, and it improves as traffic grows. A slow-handler problem shows slow responses on misses regardless of traffic, and the hit ratio is healthy wherever there is any traffic at all.

Fifty corner shops each stocking the same product from scratch sell less from the shelf than one busy supermarket that turns its stock over every hour. Thin demand spread thin means most customers find an empty shelf.

saying these in an interview costs you the question

  • Assumes more locations always mean faster responses
  • Thinks one location's cache warms the others
  • Judges the hit ratio from total traffic, not per location
  • Reads a global average as proof every location is warm
  • Recommends widening the spread to fix a cold-cache problem