skip to content

A data-heavy route got slower after it was deployed to many locations near users instead of one region. Why, and what do you change?

level: middleimportance: must knowfreq 52%

answer

  1. the code moved, the data did not
  2. count round trips, not locations
  3. depth of the waterfall multiplies the penalty
  4. pin the route to the data's home

basics

~20 s

The database stayed in one region, so every query the route makes is now a cross-region round trip, and sequential queries multiply it. Pin the route back to the data's region, or collapse it to one round trip.

solid answer

~50 s

Moving the code did not move the data. Inside the region a query cost a millisecond or two; from a location near the user it costs a full inter-continental round trip, and because the handler awaits its queries one after another, the penalty multiplies by the depth of that waterfall. Saving 150 ms on the user hop while adding three dependent 120 ms hops is a net loss for every visitor, including the ones who used to be closest to the region. The repairs, in order: declare the data's region as this route's home so it runs beside the store again; collapse the waterfall so at most one long hop remains; or split the route, keeping the request-only part near the user and leaving the rendering in the region. Cacheable responses are a fourth option, but only where the response is not per-user.

go deeper

for a junior

Remember the rule of thumb: code near the user is close to the visitor and far from the database. A route that reads a lot before it can answer is the wrong candidate.

for a middle

Explain the arithmetic: one hop saved on the user side against one long hop per data round trip, multiplied by how deep the sequence of awaits runs.

for a senior

Demonstrate the diagnosis — break the route's time into hop, CPU and data wait, confirm the waterfall's depth, then pick between pinning the route and collapsing the queries.

for a principal

Own the policy rather than the incident: make placement reversible per route and make a route that gains its first data call trigger a placement review.

## Why the numbers went the wrong way Near-user placement is a trade, not an upgrade. It shortens the hop between the browser and the machine running the route, and it lengthens the hop between that machine and everything the route reads. A route that is mostly reading is on the losing side of that trade, and the loss is easy to miss because the change looks like a pure win in a demo where the developer sits next to the new location and the dataset is tiny. The measurement that settles it is a per-route breakdown: **time to the handler**, **handler CPU**, and **time waiting on data**. Before the move, the third term was small because the handler and the store shared a region. After it, the third term is the response. ## The arithmetic of a waterfall Depth is what hurts, more than count. Consider a route that reads a session, then a user record, then that user's items, each awaited before the next begins. | shape | in the data's region | at a location near the user | |---|---|---| | one round trip | a few milliseconds | one long hop | | three independent trips issued together | a few milliseconds | one long hop | | three dependent trips, one after another | still only a few milliseconds | three long hops, added up | In-region, a waterfall three deep is nearly free, so nobody writes it as a problem. That is exactly why the pattern is everywhere in code that was written for a single region, and why the same code falls apart the moment the handler is moved away from the store. ## Four repairs, strongest first 1. **Give the route a home region.** Most platforms let placement be declared per route, so the deployment as a whole can stay near users while this route says “run me beside the data”. It is the smallest change, it is reversible, and it restores the original numbers exactly. 2. **Collapse the waterfall.** If the route genuinely must run near the user, reduce it to a single round trip — one query that returns everything, or independent queries issued together rather than in sequence. One long hop is often affordable; three are not. 3. **Split the route.** Keep the part that is request-only — the redirect, the variant choice, the header work — near the user, and leave the part that reads data in the region. The two pieces are separate handlers with separate placements. 4. **Serve a cached response.** If the response is the same for many users for some window, most requests never reach the data at all. This only helps when the response is genuinely shared; caching a per-user response for everyone is a correctness bug, not an optimisation. ## What not to do - **Adding more locations.** The problem is the distance from the handler to the store, and more locations do not shorten it; several of them make it longer. - **Moving the reads into the browser.** The same cross-region hops still happen, now on a slower network and after the document has already loaded, and the change costs you server-side rendering as well. - **Retrying or raising the timeout.** That treats a structural latency as flakiness and usually makes the tail worse. - **Blaming the framework.** The framework ran the code you gave it in the place you asked for. The defect is a waterfall that was invisible while it was cheap. ## Regional pinning as the normal outcome It is worth saying plainly that pinning a route to the data's region is not a retreat. A healthy app usually ends up mixed: the pre-routing work and the thin, request-only handlers answer near the user, and anything that renders from the store runs beside the store. The placement question is settled per route because the deciding property — how much the route must talk to its data before it can answer — is a property of the route. One more thing follows from this. The right answer changes when the inputs change. A route that was request-only and gains its first data call has quietly switched sides, and so has a route left near the user after the data moved to a different region. Re-examine placement whenever either of those happens, rather than treating the original decision as permanent.

  • Why do three sequential queries hurt so much more than three issued together?
    Each sequential query pays a full cross-region round trip before the next one is even sent, so the penalty multiplies by the depth of the chain. Issuing independent queries together pays that cost once, in parallel. Depth, not count, is what near-user placement punishes — which is why a waterfall that was free in-region falls apart after the move.
  • When is pinning a single route better than moving the whole app back?
    Almost always, when the platform supports it. Pinning keeps one deployment model: the request-only routes keep their win near the user, and only the route that reads heavily returns to the data's region. Moving everything back throws away the genuine gains to fix a problem that belongs to a handful of handlers.
  • How would you catch this before shipping rather than after?
    Measure the route's high-percentile latency from a client in a far geography, with the data in its real region, and instrument the handler's time-waiting-on-data separately from its CPU. A route whose data wait dominates should never be promoted to near-user placement in the first place.

saying these in an interview costs you the question

  • Blames the framework instead of counting cross-region round trips
  • Thinks one slow query is the cause when three run in sequence
  • Assumes adding more locations will eventually fix the latency
  • Believes running code near users replicates the database automatically
  • Caches a per-user response for everyone to hide the latency