Moving a route's server code from one region to many locations near users: what does that make faster, and what stays the same?
answer
- only one hop gets shorter
- the handler's own work is unchanged
- the data did not move with the code
- request-only work wins, queries lose
- count round trips before the first byte
basics
~20 sRunning a route's server code near the user shortens only the network trip between browser and server. The work inside the handler takes the same time, and the distance from that code to its data usually gets longer.
solid answer
~50 sPlacement changes one term in the latency budget: the round trip between the user and whatever machine runs the route's server code. If the handler needs only the request itself — a rewrite, a redirect, a cookie or header check, a country- or experiment-based decision — that term is most of the cost, so answering from a location near the user can turn a couple of hundred milliseconds into a few. Nothing else improves. The handler's own CPU work takes exactly as long, and the database, cache and upstream services did not move, so every call the route makes is now a long hop from wherever the code ran to wherever the data lives. That is why the same change makes a request-only route dramatically faster and a three-query route slower, and why it is a per-route decision rather than an app-wide switch.
go deeper
Hold on to the one-line version: placement shortens the hop between the user and the server and nothing else. The work inside the handler and the distance to the database are unaffected.
Be able to split a route's response time into user-to-server latency, handler CPU and data round trips, and say out loud which of the three the placement decision touches.
Show that you measure before moving anything: the win exists only when the user-to-server hop dominates that route's high-percentile latency in the geographies you actually serve.
Frame it as a budget question. The milliseconds saved on request-only routes have to pay for operating two execution environments with two sets of constraints and two debugging stories.
## The decision in one sentence A route that renders on the server needs a machine to run its server code on. The common default is a **single region**: one place, chosen once, where the process or the function that handles the route lives. Most deployment platforms also offer a second placement — run the same route's server code in a large number of small locations spread around the world, and let each request be handled by whichever location is nearest to the user. The source code is identical; only *where it executes* changes, and frameworks usually expose that as a setting on the route or on the deployment rather than as a rewrite. ## The three terms in a response's latency Think of a server-rendered response as a sum: - **user → server** — the network time for the request to reach the machine running the route, plus the response travelling back; - **the handler's own work** — parsing the request, running the route's logic, rendering the markup, serialising the payload; - **server → data** — every round trip the handler makes to a database, a cache or an upstream service before it can answer. Placement moves the **first** term, leaves the **second** untouched, and changes the **third** — usually for the worse, because the data did not travel with the code. | term | one region | many locations near users | |---|---|---| | user → server | long for distant users, short for nearby ones | short for almost everyone | | handler CPU | unchanged | unchanged | | server → data | short: the data is in the same region | long: a cross-region hop per round trip | ## Why request-only work wins Some route work needs nothing but the incoming request. Rewriting a URL, redirecting an unauthenticated visitor, reading a cookie to pick a variant, choosing a locale from a geography hint the platform attaches to the request, setting response headers — all of the inputs are already in memory the moment the handler starts. The handler's own work is microseconds, `server → data` is zero, so the response time *is* the user hop. A visitor on another continent may spend well over a hundred milliseconds simply reaching a single region; a location a few network hops away answers in a small fraction of that. This is the case where near-user placement is not a marginal gain but a different order of magnitude. ## Why data-hungry work loses The opposite case is a route that cannot answer until it has read something. Placement moved the reader, not the store. A query that cost a millisecond or two inside the region now pays a full inter-continental round trip, and if the handler makes several queries **in sequence**, each one pays that penalty before the next begins. A route that saved 150 ms on the user hop and added three cross-region round trips of 120 ms is net slower — and slower for *everyone*, including the users who were previously closest to the region. ## What placement does not change - The handler's CPU time. The same code doing the same work takes the same time; splitting traffic across locations raises capacity, not per-request speed. - The size of the response. Bytes on the wire are unaffected by where they were produced. - Where the data lives. No copy of the database appears alongside the code. - Anything about caching. A cached response is fast because it was cached, not because of where the handler sits; conversely, whatever warm state a location holds is that location's alone. - Correctness. Placement is a performance and operability decision, not a behavioural one — except where the constrained runtime some platforms use for near-user execution refuses a dependency the route needs. ## The test to apply per route Count the round trips the handler must complete before it can produce its first byte. 1. **Zero round trips** — a strong candidate for running near the user; the whole response time is the hop you are about to shorten. 2. **One round trip** — measure. The saving on the user hop may or may not cover one long hop to the data. 3. **Several, especially dependent ones** — keep it beside the data, or restructure it before you move it. Meta-frameworks differ in how granular this choice is: some let every route declare its own placement, others settle it once for a whole deployment or for one category of handler. When the granularity is per route, the answer should be per route too — an app almost always has both kinds of route in it.
- Which kinds of route work are genuinely request-only?Work whose every input is already in the request or in the code: rewriting a URL, issuing a redirect, reading a cookie or header to choose a variant, setting response headers, picking a locale from a geography hint the platform attaches. It reads no store and calls no service, so it completes in microseconds wherever it runs.
- Does serving the HTML from near the user help if the page then fetches its data from the browser?Only partly. The document arrives sooner and the browser starts parsing earlier, but the data request still travels to wherever the API lives. You have shortened the first hop and left the second. If the page cannot show anything useful until that data lands, the time to something meaningful barely moves.
- Why is this a per-route decision rather than an app-wide one?Because the deciding factor is a property of the route, not of the app: how much it must talk to its data before it can answer. A typical app contains both request-only handlers, which gain a lot, and read-heavy pages, which lose. Settling it once for everything guarantees one of the two groups is placed wrongly.
Opening a branch office on every high street helps the customers who just need a form stamped. It does nothing for the ones whose file still sits in the single head-office archive — those now wait for a courier both ways.
saying these in an interview costs you the question
- Thinks running code near the user speeds up database queries
- Assumes every route gets faster once the app runs near users
- Believes placement removes the network round trip entirely
- Confuses running server code near users with caching static files
- Treats placement as an app-wide switch rather than a per-route call