skip to content

Why does a route rendered on the server for each request send nothing to the browser until its data and render finish?

level: middleimportance: must knowfreq 70%

answer

  1. the body is the render output
  2. a data dependency, not a policy
  3. sequential reads add up
  4. queueing shows up as a cliff
  5. latency moved, not removed

basics

~20 s

The response body is the rendered markup, and markup cannot be produced before the values it prints are known. So the request waits for the route's server data work plus render time, and the first byte carries both of those costs.

solid answer

~50 s

The body of the response *is* the render output, so it does not exist until the render has run, and the render cannot run until the data it prints has resolved. A visitor's wait therefore stacks up: network time to the server, any cold start of the runtime, the route's data reads, render CPU, serialisation of the data for the client, and the trip back. Two of those are usually dominant — the slowest data read, because reads that must happen in sequence add up rather than overlap, and the render itself on a busy machine. The consequence is that upstream latency becomes user-visible as a delay before anything at all appears, not as a spinner on a page already on screen. Frameworks that can flush the response in pieces break this rule deliberately; that is a separate mode.

go deeper

for a junior

Remember the causal order: data first, then markup, then the response. Nothing can be sent before the string exists, so a slow data source delays the entire page rather than one part of it.

for a middle

Break the wait into its segments — network, runtime start, data reads, render, serialisation — and say which usually dominates and why sequential reads are worse than concurrent ones.

for a senior

Show you have measured it: separate queueing from CPU from upstream latency, bound every dependency with a timeout, and know that saturation appears as a cliff rather than a gradual slowdown.

for a principal

Argue about where the render should live relative to the data, and what latency budget a held response is allowed to consume before the route's mode or its data path has to change.

## Why the response is held An HTTP response is a status line, headers, then a body. When a route is rendered per request, the body **is** the string the renderer produces. A string cannot be written before it has been built, and it cannot be built while the values it interpolates are still unresolved. So the server holds the connection open, does the work, and writes everything at once. That ordering is not a framework quirk, it is a data dependency: 1. The router resolves the URL to a route and its server-side data work. 2. That data work is awaited — database queries, upstream service calls, session lookups. 3. The component tree renders to markup using the resolved values. 4. The data is serialised into the document for the client to take over with. 5. Status, headers and body are written; the first byte leaves the server. Step 5 cannot be pulled in front of steps 2 and 3 without changing what the response says. Note the one real exception: a framework that flushes the response in pieces can send an early part of the document before later parts are ready. That is a deliberately different mode with its own failure behaviour, and it does not change the rule for the simple case described here. ## What the wait is actually made of | Segment | Typical cause | Who controls it | |---|---|---| | Client to server network | Distance, connection setup | Hosting location, protocol | | Runtime start | A cold start where the render runs in a short-lived process | Hosting shape | | Route data work | Upstream queries, especially reads that must happen in sequence | Route code and data sources | | Render | Tree size, expensive component work, machine load | Application code, capacity | | Serialisation | The size of the data embedded for the client | What the route returns | | Server to client network | Document size, bandwidth | Payload and compression | Two of these usually dominate. The first is the **slowest data read**: if a route makes reads that depend on one another, their latencies add, so one extra sequential hop can cost more than the whole render. The second is **render time under load**: rendering is CPU work, and when a machine is saturated the render sits in a queue before it even starts, which shows up as a latency cliff rather than a gentle slope. ## What the user sees during the wait Nothing from the new page. On a full navigation the browser keeps showing the previous document (or an empty tab) with the loading indicator spinning. There is no skeleton, no header, no layout — the shell itself is inside the body that has not been written. This is why a slow upstream on a per-request route feels different from a slow upstream on a page that is already on screen: it delays the *whole* page rather than one region of it. It also means the visitor cannot tell a slow render from a slow network. From the browser's side both are simply a long gap before the first byte. ## What is inside your control - **Make the reads overlap instead of stacking** wherever they do not depend on each other. The route's wait is the longest read, not the sum, only when they run concurrently. - **Bound every upstream call** with a timeout, so one unhealthy dependency cannot hold the response open indefinitely. What to do when that timeout fires is a policy decision, not a default. - **Keep render work proportionate.** Expensive formatting, large lists rendered whole, and deep trees all cost CPU per view, and they cost it again for every visitor. - **Watch the serialised payload.** The data embedded for the client is written into the same body; a large blob delays the last byte and inflates the document. - **Measure the segments separately.** A single end-to-end number cannot tell you whether the time went to a query, to CPU or to queueing, and each has a different fix. ## Where the shape of the deployment changes the answer On a long-running server process the runtime is already warm and the first request pays nothing extra; the risk is saturation, where concurrent renders compete for the same CPU. On a platform that starts a short-lived process per request, an idle route can pay a start-up cost before any of your code runs, and that cost lands on exactly the unlucky visitor who arrives first. On an edge runtime the network segment shrinks while the distance to the data source may grow, which can move the bottleneck rather than remove it. The dependency chain is identical in all three; only the constant terms differ. ## The honest summary Rendering per request does not create latency, it *relocates* it: work the browser would have done after the document lands is done before it is sent. That is a good trade when the server is close to the data and the output must be fresh or personal, and a poor one when the render adds a long serial wait to information that did not need to be current for this visitor.

  • Why can two routes with identical render code have very different first-byte times?
    Because the render is rarely the dominant term. One route may await several dependent upstream reads while the other reads one fast source, and one may run on a warm process while the other pays a start-up cost. Same render, different data chain and different runtime state.
  • If a route awaits three independent upstream calls, what is its data wait?
    The longest single call if they are issued concurrently, or the sum of all three if each is awaited before the next is started. The difference is entirely in how the route issues them, and it is one of the cheapest wins available on a per-request route.
  • Does putting the render closer to the user always reduce the wait?
    No. Moving the render to the network edge shortens the visitor-to-server hop but can lengthen the server-to-data hop, and the data hop is often the larger term. If the render awaits a database in one region, running it far from that region usually makes the held response longer, not shorter.

saying these in an interview costs you the question

  • Assumes every framework flushes markup while the render still runs
  • Assumes render CPU is always the dominant cost
  • Awaits independent upstream reads one after another
  • Leaves upstream calls with no timeout at all
  • Believes moving the render to the edge always helps
  • Treats one end-to-end number as enough to diagnose it