skip to content

On a single HTTP/1.1 TCP connection, what happens when a client wants to issue three requests at once, and what is the resulting delay when a slow response holds up the ones behind it called?

level: juniorimportance: must knowfreq 55%

answer

  1. One exchange at a time per connection
  2. No request IDs — order is the only matching
  3. Slow first response delays the queue
  4. App-layer HOL vs TCP-level HOL
  5. ~6 connections per origin as the workaround

basics

~10 s

HTTP/1.1 handles one request-response exchange at a time per connection, so the second and third wait for the first to finish. A slow first response stalls the queue behind it; that is head-of-line blocking.

solid answer

~50 s

An HTTP/1.1 connection is strictly serial: the client writes a request, reads the complete response, and only then may send the next one. Responses have no identifiers, so the only way to know which response belongs to which request is ordering — which forces one exchange at a time. With three requests on one connection, requests two and three sit in a queue. If the first takes two seconds — slow database query, large body — the others are delayed by two seconds no matter how fast they would have been. That is **head-of-line blocking**: the item at the front holds up everything behind it. The workaround built into every browser is concurrency through multiple connections, conventionally about six per origin, so six exchanges can be in flight simultaneously and a slow one blocks only its own connection. HTTP/2 solves it properly at the application layer by multiplexing independent streams with identifiers over one connection.

code

http · 15 lines
http
GET /a HTTP/1.1
Host: example.com

HTTP/1.1 200 OK
Content-Length: 12

hello world

GET /b HTTP/1.1
Host: example.com

HTTP/1.1 200 OK
Content-Length: 3

bye

go deeper

for a junior

State plainly that one HTTP/1.1 connection carries one request-response at a time and that a slow response delays whatever is queued behind it.

for a middle

Explain why ordering is the only correlation mechanism, and name the workarounds: several connections per origin, and multiplexing in HTTP/2.

for a senior

Separate application-layer from transport-layer blocking, quantify the round-trip cost for asset-heavy pages, and recognize the waterfall signature.

for a principal

Connect it to delivery strategy: protocol version, connection reuse, bundling policy, and where the remaining transport-level blocking still bites.

## The serial model HTTP/1.1 defines a conversation on a connection as a strict alternation: request, response, request, response. There is no request identifier anywhere in the message. A response is matched to a request purely by position — the first response received answers the first request sent. That design is simple and human-readable, and it makes concurrency on one connection impossible without extra rules. So a client that wants three resources at once from one connection must send request 1, wait for the whole of response 1, send request 2, wait, send request 3. Total time is the sum of the three exchanges, and each includes at least one network round trip. ## Head-of-line blocking When the first exchange is slow, everything queued behind it inherits that delay even though those requests could have been served instantly. The name comes from queueing theory: the item at the head of the line blocks the line. Concretely, one endpoint that takes two seconds delays a 20-millisecond request behind it to 2.02 seconds. This is **application-layer** head-of-line blocking — caused by HTTP/1.1's message framing, not by the network. It is worth separating from **transport-layer** head-of-line blocking, where a single lost TCP segment forces the kernel to withhold all subsequent, already-arrived bytes until the retransmission fills the gap. The first is fixed by multiplexing at the HTTP layer; the second persists on any protocol that runs over TCP, which is a large part of why QUIC exists. ## Why it mattered so much A typical web page needs dozens to hundreds of subresources — scripts, stylesheets, images, fonts. Serialized over one connection, each pays a round trip, and on a 100 ms link 100 assets is at least ten seconds of pure latency before any bandwidth argument enters. This single fact shaped a decade of web performance practice: bundling scripts, spriting images, inlining small assets as data URIs, and sharding across hostnames all exist to fight per-connection serialization. ## The workarounds 1. **Persistent connections.** Reusing one connection avoids repeating TCP and TLS handshakes, but does nothing about serialization — exchanges are still one at a time. 2. **Multiple connections per origin.** Browsers open several connections to the same host, conventionally about six, and dispatch requests across them. That is real parallelism: six exchanges in flight, and a slow response blocks only its own connection while the other five continue. This is the workaround that actually carried the web through the HTTP/1.1 era. 3. **Pipelining.** HTTP/1.1 allows sending several requests without waiting, but responses must still come back in order, so the blocking is not removed — only the request-side round trips are. Browsers abandoned it. 4. **Domain sharding.** Serve assets from several hostnames to multiply the per-origin connection cap. ## How it was actually fixed HTTP/2 gives every exchange a stream identifier and interleaves frames from many streams over a single connection. Responses may arrive in any order and in pieces, so a slow response no longer blocks anything at the HTTP layer, and one connection per origin becomes sufficient. Transport-level blocking from TCP loss remains, which HTTP/3 addresses by running over QUIC where streams are independently delivered. ## What to observe in practice In a browser's network panel over HTTP/1.1, requests beyond the connection limit show substantial 'Stalled' or 'Queueing' time before any bytes move, while their server timings are small. That pattern — large queueing, small service time — is the visible fingerprint of connection-level serialization, and it is the thing to recognize rather than to memorize the exact terminology. ## Summary One connection, one exchange at a time, ordering as the only correlation mechanism. Everything else in this area — the six-connection convention, sharding, bundling, and eventually multiplexing — is a response to that one constraint.

  • Do persistent connections solve head-of-line blocking?
    No. Keeping a connection open removes the repeated TCP and TLS handshake cost, which is a real saving, but the exchanges on it remain strictly serial. A slow response still blocks the next request on that same connection; only parallel connections or multiplexing address the blocking itself.
  • How is head-of-line blocking caused by TCP packet loss different from the HTTP/1.1 kind?
    HTTP/1.1 blocking is caused by the protocol's message framing: one exchange at a time, matched by order. TCP blocking happens below HTTP: a lost segment makes the kernel withhold later bytes that already arrived until the gap is filled, which stalls every multiplexed stream sharing that connection. HTTP/2 fixes the first but not the second; HTTP/3 over QUIC addresses the second.

A single-lane supermarket checkout: the shopper with the full trolley in front holds up everyone with a single item, no matter how fast those purchases would have been.

saying these in an interview costs you the question

  • Believing HTTP/1.1 can interleave two responses on one connection
  • Thinking keep-alive makes requests concurrent
  • Confusing HTTP-level blocking with TCP retransmission blocking
  • Claiming the six-connection limit is imposed by the protocol rather than by browser convention
  • Assuming bandwidth is the bottleneck when the real cost is serialized round trips

context