skip to content

A page served over HTTP/1.1 loads slowly. The network waterfall shows dozens of requests with long 'Stalled' or 'Queueing' time before any bytes move, while each server response time is small. How do you interpret that, and what would you do?

level: seniorimportance: should knowfreq 38%

answer

  1. Queueing large + TTFB small = connection starvation
  2. Concurrency plateaus at ~6 per origin
  3. Check the protocol column first
  4. Root fix = HTTP/2 or HTTP/3
  5. Fewer, better-ordered requests

basics

~20 s

Large queueing with small service time means requests are waiting for a free connection, not for the server. The client has more requests than the roughly six per-origin HTTP/1.1 connections allow. Fix it by moving to HTTP/2 or reducing and reordering requests.

solid answer

~50 s

That signature — big 'Stalled'/'Queueing', small 'Waiting (TTFB)' — says the bottleneck is client-side connection availability, not the backend. Over HTTP/1.1 each connection carries one exchange at a time and the browser holds about six per origin, so request 7 onward simply waits for a slot. I would confirm before acting: check the protocol column is `http/1.1`, count concurrent in-flight requests (it will plateau near six per host), verify server timings are genuinely small, and check whether one slow response is holding a connection hostage. Scaling backend capacity will do nothing here, which is the expensive mistake. Fixes in order of value: enable HTTP/2 or HTTP/3 so one connection multiplexes everything — the root fix. Then reduce request count (bundle, sprite, inline critical CSS) and defer non-critical assets. Prioritize render-blocking resources and use preload hints. Move third-party assets off the critical path. Under HTTP/1.1 only, limited domain sharding can help — but it becomes harmful after upgrading.

go deeper

for a junior

Recognize that waiting-for-a-connection time and waiting-for-the-server time are different phases, and that this one is the former.

for a middle

Confirm the protocol and the concurrency plateau, then reduce request count and enable HTTP/2.

for a senior

Rule out a hostage connection and third-party origins, sequence the fixes by leverage, and state explicitly that backend scaling will not help.

for a principal

Treat it as delivery architecture: protocol termination, CDN topology, third-party budget and a critical-path policy, plus retiring the HTTP/1.1-era workarounds the upgrade makes harmful.

## Reading the waterfall A browser waterfall splits each request into phases, roughly: queueing/stalled, DNS, connecting, TLS, request sent, waiting for the first byte (TTFB), and content download. Each phase points somewhere different: - **Long queueing/stalled, short TTFB** — the request had nowhere to go. The browser has more requests than free connections. This is connection-level serialization. - **Long TTFB** — the server is slow: the request reached it and it took time to answer. - **Long content download** — bandwidth or response size. - **Long DNS/connect/TLS** — setup cost, often the fingerprint of too many distinct hostnames. The described case is unambiguous: the browser holds the requests before sending them, so no amount of backend optimization moves the number. Scaling application servers or tuning queries is the classic wasted response. ## Confirming the diagnosis 1. **Check the protocol.** Enable the protocol column; `http/1.1` on the hot assets confirms the premise. Frequently only a subset (a legacy asset host, an internal service, a misconfigured origin) is still on 1.1. 2. **Count concurrency.** Plot in-flight requests over time; a hard plateau around six per origin is the fingerprint. If the plateau sits at six times the number of hosts, the budget is already being multiplied by sharding. 3. **Check for a hostage connection.** One genuinely slow response occupying a connection reduces effective parallelism from six to five for its whole duration. A slow analytics beacon or an unresponsive third party can do real damage this way. 4. **Separate origins.** Third-party hosts have their own budgets; a page slow because of a third party shows queueing concentrated on that host. 5. **Verify server timings independently.** Compare with server-side latency metrics so you are not trusting one measurement. ## Fixes, in order of leverage **Upgrade the protocol.** HTTP/2 multiplexes many streams on one connection, so the queue disappears at the application layer and one connection per origin suffices. HTTP/3 over QUIC additionally removes transport-level head-of-line blocking, which matters on lossy mobile networks. On modern stacks this is a load-balancer or CDN setting, and it is the root fix rather than a mitigation. **Reduce request count.** Every request removed frees a slot. Bundle scripts and styles, inline small critical CSS, use sprites or icon fonts for many tiny images, and lazy-load below-the-fold media. These practices lose value under HTTP/2 — excessive bundling then hurts caching — so treat them as HTTP/1.1-era mitigations. **Reorder rather than just shrink.** With a fixed number of slots, what occupies them matters more than the total. Render-blocking CSS and critical fonts should be requested first; scripts should be `defer`/`async`; `<link rel=preload>` promotes genuinely critical assets. A single misplaced synchronous script can hold a slot and block rendering. **Contain third parties.** Tags and beacons compete for slots on their own origins and can hold connections open. Load them late, asynchronously, or from a worker. **Shard — only if you must stay on HTTP/1.1.** Two extra hostnames multiply the budget, at the cost of DNS and handshakes and cache fragmentation, and it must be undone when you upgrade. **Move bytes closer.** A CDN shortens the round trip, and since queueing cost is a multiple of round trips, that shrinks the total even without changing concurrency. ## What not to conclude Do not read this as a server capacity problem, do not add application instances, and do not raise backend connection pools — the bottleneck is on the client side of the wire. Equally, do not assume upgrading to HTTP/2 fixes everything: if one backend endpoint really is slow, multiplexing lets the other requests proceed but the slow one stays slow, and transport-level blocking from packet loss still affects a multiplexed HTTP/2 connection. ## The one-line takeaway Queueing time is a client-side connection-availability signal; TTFB is the server signal. Read which phase is large before choosing what to optimize.

  • How do you distinguish this from a genuinely slow backend in the same waterfall?
    Look at which phase is long. Connection starvation shows a long stalled/queueing segment with a short waiting-for-first-byte segment, and requests start in batches as slots free up. A slow backend shows the opposite: the request is sent immediately and the time is spent waiting for the first byte, which server-side latency metrics will corroborate.
  • After upgrading to HTTP/2, which parts of the problem remain?
    Any genuinely slow endpoint is still slow — multiplexing lets other requests overtake it but does not speed it up. Transport-level head-of-line blocking also remains, because a lost TCP segment stalls every stream on that connection until retransmission; only HTTP/3 over QUIC addresses that. Bandwidth limits and oversized payloads are unaffected as well.

saying these in an interview costs you the question

  • Adding backend capacity in response to a client-side queueing signal
  • Reading queueing time as server latency
  • Not checking which protocol version the slow requests actually used
  • Introducing domain sharding without first checking whether HTTP/2 is available
  • Assuming HTTP/2 removes all head-of-line blocking, including the TCP-level kind

context