skip to content

A server-rendered page holds its whole HTML response until a 600 ms database query finishes, then sends it in one piece. What changes for the browser if the server instead streams the HTML in chunks, flushing the <head> before that query completes — and what does it not change?

level: middleimportance: must knowfreq 65%

answer

  1. server think time is dead air
  2. head first, body when ready
  3. subresources download during the query
  4. TTFB falls, total work does not
  5. status code is committed at first flush

basics

~20 s

Streaming sends the <head> immediately, so the browser parses it and starts downloading CSS, fonts and scripts while the server is still querying. The query time, the byte count and the work the browser must do are unchanged.

solid answer

~50 s

With a buffered response the connection sits idle for the full 600 ms: the browser has no bytes, so it cannot discover a single stylesheet or script. Streaming turns that dead air into useful work. The server writes `<!doctype html>`, the `<head>` and any above-the-fold shell first and flushes; the browser's parser and preload scanner see those bytes right away and open connections for the subresources in parallel with the query. When the data is ready the server writes the rest of the body on the same response. Time to first byte falls, first paint usually falls with it, and the critical requests are already in flight instead of starting after the query. What does not change: the query still takes 600 ms, the total bytes are the same, and if the slow data *is* the largest visible element, the largest contentful paint may barely move.

code

javascript · 11 lines
javascript
import http from 'node:http';

const slowQuery = () => new Promise((resolve) => setTimeout(() => resolve('Hello'), 600));

http.createServer(async (req, res) => {
  res.writeHead(200, { 'Content-Type': 'text/html; charset=utf-8' });
  res.write('<!doctype html><html><head><link rel="stylesheet" href="/app.css"></head><body>');
  const data = await slowQuery();
  res.write(`<main>${data}</main>`);
  res.end('</body></html>');
}).listen(3000);

go deeper

for a junior

Be able to say that HTML can arrive in pieces and the browser starts parsing the first piece straight away, so styles and scripts begin downloading before the server has finished the page.

for a middle

Explain the overlap concretely: head flushed first, preload scanner discovers subresources during server think time, body written when data resolves. Name what improves (TTFB, first paint) and what does not (query time, total bytes).

for a senior

Show that you would verify streaming end to end in production, comparing time to first byte against total document transfer time, and that you know buffering proxies or late-flushing compression can silently erase the benefit.

for a principal

Own the tradeoff: committing headers at the first flush constrains error handling, redirects and cookies, and it interacts with full-page caching. Decide where streaming is worth the added complexity rather than mandating it everywhere.

## The shape of the problem A dynamic page is produced in two phases: the server thinks (queries a database, calls services, renders a template), then bytes travel to the browser, then the browser works (parses HTML, fetches CSS/JS/images, paints). In a buffered response those phases are strictly sequential. During the server's 600 ms of thinking, the browser holds an open connection and nothing else. It cannot know the page needs `/app.css`, because it has not seen a single byte of markup. Streaming overlaps the phases. HTTP responses do not have to be delivered as one blob — the server can write part of the body, flush it, keep the response open, and write more later. On HTTP/1.1 this is chunked transfer encoding; on HTTP/2 and HTTP/3 the body simply arrives as more frames. The important part for a front-end engineer is not the wire format but the consequence: **the browser can begin work on a document the server has not finished producing.** ## What the browser does with a partial document HTML parsing is incremental by design. The parser consumes bytes as they arrive and builds DOM as it goes; it does not wait for `</html>`. Separately, browsers run a lightweight scanner ahead of the parser that looks for URLs in the bytes already received and starts fetching them. So the moment the `<head>` lands: - stylesheets and preloaded fonts begin downloading; - deferred and module scripts begin downloading; - DNS, TCP and TLS for third-party origins can be set up. All of that happens concurrently with the server's query. By the time the body chunk arrives, the CSS may already be parsed and the page can paint almost immediately. ## Ordering: what to flush first The practical rule is to flush everything that does not depend on slow data as early as possible: 1. Doctype and `<head>` — charset, stylesheet links, hints, anything needed to fetch the critical path. 2. The static shell — header, navigation, layout containers, and a placeholder for the slow region. 3. The slow region, written when its data resolves. A server that renders its template only after `await`ing every data source cannot do this, which is why streaming usually forces a code change: the template has to be produced in pieces, and data fetches have to start before rendering rather than inside it. ```javascript res.writeHead(200, { 'Content-Type': 'text/html; charset=utf-8' }); res.write('<!doctype html><html><head><link rel="stylesheet" href="/app.css"></head><body>'); const rows = await slowQuery(); // browser is fetching CSS during this await res.write(renderRows(rows)); res.end('</body></html>'); ``` ## What improves, and what does not Time to first byte improves, because first byte no longer waits for the query. First paint usually improves, because the render-blocking CSS was fetched during the server's think time. But streaming is a *scheduling* change, not a *work reduction*: the query still costs 600 ms, the bytes are identical, and the JavaScript still has to execute. If the slow data is the biggest visible element on the screen, the largest contentful paint may hardly move — the user simply sees the frame and a placeholder sooner. That is real value (the page stops looking broken), but claiming streaming "made the page 600 ms faster" is a claim the numbers will not support. ## The costs to be honest about Once the first chunk is flushed, the status line and headers are committed. The server can no longer switch to a 500, issue a redirect, or set a cookie — an error mid-render has to be handled inside the already-streaming body. Intermediaries can also undo the benefit: a reverse proxy or CDN that buffers the response will collect the whole body before forwarding it, and compression that only flushes at the end has the same effect. Streaming therefore needs verifying end to end in production, not just against a local dev server. ## How you would verify it Compare time to first byte with total transfer time for the document request. Buffered: they are nearly equal and both large. Streaming: first byte is small while total stays large. In the browser's network panel the document row shows a long "content download" phase rather than a long "waiting" phase, and the subresource requests start near the beginning of the document request rather than after it.

  • If total load time barely moves, how do you justify the change to a sceptical product owner?
    By separating the two numbers. Time to first byte and first paint drop measurably, and the user sees a laid-out page with a placeholder instead of a white screen for most of the wait. The honest framing is that streaming removes idle time from the front of the page load; it does not make the slow query faster, so it is worth pairing with fixing that query.
  • What has to change in application code before streaming actually helps?
    Data fetching has to start before rendering, and the template has to be emitted in pieces. If the handler awaits every data source and then renders one string, flushing early is impossible because no markup exists yet. The usual refactor is to kick off the slow fetches, immediately write the head and shell, then write each region as its promise resolves.
  • Once the first chunk is flushed, what can the server no longer do?
    Anything that lives in the status line or headers: it cannot change a 200 into a 500, issue a redirect, or set a cookie. An error discovered halfway through rendering has to be handled inside the body — by writing an inline error region, or by buffering just enough of the response that most failure paths are detected before the first flush.

A buffered response is a waiter who takes your order, disappears until every dish is plated, then carries them all out at once. Streaming is the same kitchen sending bread and water immediately — the meal is not cooked faster, but you stop staring at an empty table.

saying these in an interview costs you the question

  • Says streaming makes the database query itself faster
  • Claims total page load time always drops with streaming
  • Thinks the parser waits for the closing html tag
  • Confuses streaming HTML with client-side rendering
  • Assumes streaming works even behind a buffering proxy
  • Believes you can still send a redirect after the first chunk

context