skip to content

questions

21

On a single HTTP/1.1 TCP connection, what happens when a client wants to issue three requests at once, and what is the resulting delay when a slow response holds up the ones behind it called?

level: juniorimportance: must knowfreq 55%

answer

  1. One exchange at a time per connection
  2. No request IDs — order is the only matching
  3. Slow first response delays the queue
  4. App-layer HOL vs TCP-level HOL
  5. ~6 connections per origin as the workaround

basics

~10 s

HTTP/1.1 handles one request-response exchange at a time per connection, so the second and third wait for the first to finish. A slow first response stalls the queue behind it; that is head-of-line blocking.

solid answer

~50 s

An HTTP/1.1 connection is strictly serial: the client writes a request, reads the complete response, and only then may send the next one. Responses have no identifiers, so the only way to know which response belongs to which request is ordering — which forces one exchange at a time. With three requests on one connection, requests two and three sit in a queue. If the first takes two seconds — slow database query, large body — the others are delayed by two seconds no matter how fast they would have been. That is **head-of-line blocking**: the item at the front holds up everything behind it. The workaround built into every browser is concurrency through multiple connections, conventionally about six per origin, so six exchanges can be in flight simultaneously and a slow one blocks only its own connection. HTTP/2 solves it properly at the application layer by multiplexing independent streams with identifiers over one connection.

code

http · 15 lines
http
GET /a HTTP/1.1
Host: example.com

HTTP/1.1 200 OK
Content-Length: 12

hello world

GET /b HTTP/1.1
Host: example.com

HTTP/1.1 200 OK
Content-Length: 3

bye

go deeper

for a junior

State plainly that one HTTP/1.1 connection carries one request-response at a time and that a slow response delays whatever is queued behind it.

for a middle

Explain why ordering is the only correlation mechanism, and name the workarounds: several connections per origin, and multiplexing in HTTP/2.

for a senior

Separate application-layer from transport-layer blocking, quantify the round-trip cost for asset-heavy pages, and recognize the waterfall signature.

for a principal

Connect it to delivery strategy: protocol version, connection reuse, bundling policy, and where the remaining transport-level blocking still bites.

## The serial model HTTP/1.1 defines a conversation on a connection as a strict alternation: request, response, request, response. There is no request identifier anywhere in the message. A response is matched to a request purely by position — the first response received answers the first request sent. That design is simple and human-readable, and it makes concurrency on one connection impossible without extra rules. So a client that wants three resources at once from one connection must send request 1, wait for the whole of response 1, send request 2, wait, send request 3. Total time is the sum of the three exchanges, and each includes at least one network round trip. ## Head-of-line blocking When the first exchange is slow, everything queued behind it inherits that delay even though those requests could have been served instantly. The name comes from queueing theory: the item at the head of the line blocks the line. Concretely, one endpoint that takes two seconds delays a 20-millisecond request behind it to 2.02 seconds. This is **application-layer** head-of-line blocking — caused by HTTP/1.1's message framing, not by the network. It is worth separating from **transport-layer** head-of-line blocking, where a single lost TCP segment forces the kernel to withhold all subsequent, already-arrived bytes until the retransmission fills the gap. The first is fixed by multiplexing at the HTTP layer; the second persists on any protocol that runs over TCP, which is a large part of why QUIC exists. ## Why it mattered so much A typical web page needs dozens to hundreds of subresources — scripts, stylesheets, images, fonts. Serialized over one connection, each pays a round trip, and on a 100 ms link 100 assets is at least ten seconds of pure latency before any bandwidth argument enters. This single fact shaped a decade of web performance practice: bundling scripts, spriting images, inlining small assets as data URIs, and sharding across hostnames all exist to fight per-connection serialization. ## The workarounds 1. **Persistent connections.** Reusing one connection avoids repeating TCP and TLS handshakes, but does nothing about serialization — exchanges are still one at a time. 2. **Multiple connections per origin.** Browsers open several connections to the same host, conventionally about six, and dispatch requests across them. That is real parallelism: six exchanges in flight, and a slow response blocks only its own connection while the other five continue. This is the workaround that actually carried the web through the HTTP/1.1 era. 3. **Pipelining.** HTTP/1.1 allows sending several requests without waiting, but responses must still come back in order, so the blocking is not removed — only the request-side round trips are. Browsers abandoned it. 4. **Domain sharding.** Serve assets from several hostnames to multiply the per-origin connection cap. ## How it was actually fixed HTTP/2 gives every exchange a stream identifier and interleaves frames from many streams over a single connection. Responses may arrive in any order and in pieces, so a slow response no longer blocks anything at the HTTP layer, and one connection per origin becomes sufficient. Transport-level blocking from TCP loss remains, which HTTP/3 addresses by running over QUIC where streams are independently delivered. ## What to observe in practice In a browser's network panel over HTTP/1.1, requests beyond the connection limit show substantial 'Stalled' or 'Queueing' time before any bytes move, while their server timings are small. That pattern — large queueing, small service time — is the visible fingerprint of connection-level serialization, and it is the thing to recognize rather than to memorize the exact terminology. ## Summary One connection, one exchange at a time, ordering as the only correlation mechanism. Everything else in this area — the six-connection convention, sharding, bundling, and eventually multiplexing — is a response to that one constraint.

  • Do persistent connections solve head-of-line blocking?
    No. Keeping a connection open removes the repeated TCP and TLS handshake cost, which is a real saving, but the exchanges on it remain strictly serial. A slow response still blocks the next request on that same connection; only parallel connections or multiplexing address the blocking itself.
  • How is head-of-line blocking caused by TCP packet loss different from the HTTP/1.1 kind?
    HTTP/1.1 blocking is caused by the protocol's message framing: one exchange at a time, matched by order. TCP blocking happens below HTTP: a lost segment makes the kernel withhold later bytes that already arrived until the gap is filled, which stalls every multiplexed stream sharing that connection. HTTP/2 fixes the first but not the second; HTTP/3 over QUIC addresses the second.

A single-lane supermarket checkout: the shopper with the full trolley in front holds up everyone with a single item, no matter how fast those purchases would have been.

saying these in an interview costs you the question

  • Believing HTTP/1.1 can interleave two responses on one connection
  • Thinking keep-alive makes requests concurrent
  • Confusing HTTP-level blocking with TCP retransmission blocking
  • Claiming the six-connection limit is imposed by the protocol rather than by browser convention
  • Assuming bandwidth is the bottleneck when the real cost is serialized round trips

context

open as a page

In HTTP/1.1, what does it mean that connections are persistent by default, and what effect does sending the HTTP header `Connection: close` have on a request or response?

level: juniorimportance: must knowfreq 62%

basics

~20 s

HTTP/1.1 leaves the TCP connection open after a response so the next request reuses it instead of paying for a new handshake. Connection: close announces this is the last message on that connection; the sender closes it once the message ends.

open as a page

What does a client communicate by sending the HTTP request header Expect: 100-continue, and what is the server expected to do when it receives it?

level: middleimportance: must knowfreq 33%

basics

~20 s

The client sends headers only and waits before sending the body. The server either replies 100 Continue, telling it to send the body, or answers with a final status such as 401 or 413, letting the client skip the upload entirely.

open as a page

What does a client actually pay, in round trips and CPU, to open a brand-new HTTPS connection, and why does that cost justify engineering effort to reuse connections?

level: middleimportance: must knowfreq 55%

basics

~20 s

A new HTTPS connection costs a DNS lookup, a TCP three-way handshake (1 RTT), and a TLS handshake (1 RTT with TLS 1.3, 2 with TLS 1.2), plus asymmetric-crypto CPU on both ends and a cold TCP congestion window. Reuse skips all of it.

open as a page

When an HTTP client follows a redirect, which parts of the original request does it send again, and why do clients strip the `Authorization` header when the redirect points to a different origin?

level: middleimportance: must knowfreq 55%

basics

~20 s

Most request headers are re-sent to the new URL, and the body is re-sent only when the redirect preserves the method. Host is recomputed. Credentials are the exception: clients drop Authorization (and proxy credentials) on a cross-origin redirect so a redirect cannot leak your token to another host.

open as a page

An HTTP client library exposes several separate timeout settings. What is the difference between a connect timeout, a read (socket) timeout, and an idle keep-alive timeout, and which failure does each one catch?

level: middleimportance: must knowfreq 62%

basics

~20 s

Connect timeout bounds establishing the TCP (and often TLS) connection. Read timeout bounds waiting for the next bytes of a response - it is a per-read gap, not the total duration. Idle keep-alive timeout bounds how long an unused pooled connection stays open before being closed.

open as a page

How does an HTTP client connection pool work, and what happens to a request when the pool has already hit its maximum number of connections to that host?

level: seniorimportance: must knowfreq 52%

basics

~20 s

A pool keeps open connections keyed by scheme+host+port and leases one per in-flight request, returning it when the response is fully consumed. If every connection is leased, the request waits in a queue until one frees up or an acquire timeout fires - it does not silently open an extra one.

open as a page

An HTTP call to a payment service times out after the client has finished sending the request body. Is it safe to retry, and what determines the answer?

level: seniorimportance: must knowfreq 50%

basics

~20 s

A timeout means unknown, not failed - the server may have processed the request and lost the response. Retry only if the failure proves the request never ran (connect failure, refused connection, reset with no bytes received) or the operation is genuinely idempotent server-side.

open as a page

A service intermittently fails with 'connection reset by peer' or an empty response on roughly one request in a thousand, always just after a quiet period, while the target service is healthy. What race is happening and how do you address it?

level: seniorimportance: must knowfreq 45%

basics

~20 s

The server's keep-alive idle timeout closed a pooled connection at the same moment the client leased it and wrote a request. The client never saw the FIN, so it sends into a dying socket. Keep the client's idle timeout below the server's, validate connections before use, and retry safe requests once.

open as a page

A redirect response arrives with the HTTP header `Location: /v2/users?page=2` instead of a full URL. How does the client work out the exact URL to request next, and what rules govern that resolution?

level: juniorimportance: should knowfreq 42%

basics

~20 s

The Location value may be a relative reference. The client resolves it against the URL of the request that produced the redirect, inheriting scheme, host and port, then applies normal URI resolution rules - so /v2/users?page=2 becomes https://api.example.com/v2/users?page=2.

open as a page

What does HTTP status 408 Request Timeout mean, who sends it and when, and how does it differ from 504 Gateway Timeout?

level: juniorimportance: should knowfreq 30%

basics

~20 s

408 is the origin server saying it gave up waiting for the client to send a complete request in time. 504 is a gateway or proxy saying it gave up waiting for an upstream server's response. 408 blames the inbound direction, 504 the outbound one.

open as a page

A client sends a request with the HTTP header Expect: 100-continue. What are the possible server behaviours, including HTTP status 417 Expectation Failed, and what must the client do if no response arrives at all?

level: middleimportance: should knowfreq 28%

basics

~20 s

The server may send 100 Continue, a final status such as 413 rejecting it early, or 417 Expectation Failed. It may also stay silent and just wait for the body, so the client must start a short timer and send the body anyway when it expires.

open as a page

HTTP/1.1 defines request pipelining. What exactly is it, what does it improve, and why did every major browser end up disabling it?

level: middleimportance: should knowfreq 45%

basics

~20 s

Pipelining sends several requests back-to-back without waiting for each response. It saves request round trips, but responses must still return in order, so head-of-line blocking remains, and buggy proxies mismatched responses. Browsers disabled it by default.

open as a page

Browsers cap parallel HTTP/1.1 connections at roughly six per origin. Why is there a cap at all, and what was domain sharding meant to do about it?

level: middleimportance: should knowfreq 44%

basics

~20 s

More connections mean more parallel exchanges but also more server sockets and more competing TCP flows, so browsers settled on about six per origin as a compromise. Domain sharding served assets from extra hostnames to multiply that budget.

open as a page

Uploads made with curl consistently stall for about a second before the request body is sent, and the stall disappears when the request goes direct instead of through a proxy. What is curl doing, and how would you confirm and remove the delay?

level: seniorimportance: should knowfreq 30%

basics

~20 s

curl adds Expect: 100-continue for bodies over about 1 KB and waits up to a second for 100 Continue. The proxy never sends it, so curl waits out the timer then uploads. Confirm with verbose tracing; remove it by sending an empty Expect header.

open as a page

A page served over HTTP/1.1 loads slowly. The network waterfall shows dozens of requests with long 'Stalled' or 'Queueing' time before any bytes move, while each server response time is small. How do you interpret that, and what would you do?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Large queueing with small service time means requests are waiting for a free connection, not for the server. The client has more requests than the roughly six per-origin HTTP/1.1 connections allow. Fix it by moving to HTTP/2 or reducing and reordering requests.

open as a page

Why do many HTTP servers and load balancers cap how many requests they will serve on a single connection - for example advertising `Keep-Alive: timeout=5, max=100` - and what does a client observe when that cap is reached?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Caps reclaim per-connection memory and leaks, and force long-lived clients back through the load balancer so newly added backends get traffic. On the last permitted request the server answers normally but adds Connection: close; the client should transparently open a new connection.

open as a page

A page never loads and the logs show a long chain of 30x responses. How do HTTP clients detect and stop redirect loops, and what are the usual server-side causes?

level: seniorimportance: should knowfreq 38%

basics

~20 s

HTTP defines no loop limit, so clients impose a maximum hop count - typically 20 in browsers, 30 in curl, and configurable in libraries - and abort with a too-many-redirects error. Typical causes: https-redirect rules behind a TLS-terminating proxy, trailing-slash rewrites fighting each other, and login redirects where the session cookie never sticks.

open as a page

Your backend service calls third-party URLs with an HTTP client that automatically follows redirects. What controls would you place on that redirect-following behaviour before shipping it?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Cap the number of hops, re-run your URL validation on every hop (a redirect bypasses the check you did on the first URL), refuse scheme downgrades and private or link-local addresses, verify credentials are dropped cross-origin, and apply a total time budget across the chain - or disable auto-follow and handle redirects yourself.

open as a page

How would you choose timeout values for a service that sits in the middle of a call chain - an inbound caller, your service, and three downstream HTTP dependencies?

level: principalimportance: should knowfreq 40%

basics

~20 s

Start from the caller's deadline, not from downstream latency. Subtract your own processing, then give each downstream a timeout that fits the remaining budget including retries. Every inner timeout must be shorter than the outer one, and connect timeouts stay small.

open as a page

For a service accepting multi-gigabyte uploads, would you rely on the HTTP Expect: 100-continue mechanism to reject bad requests before the body is transferred? What would you weigh, and what alternatives exist?

level: principalimportance: nice to knowfreq 20%

basics

~20 s

Use it as an optimization, never as the guarantee. It only helps when every hop honours it, and it saves nothing for checks that need the body. Pair it with server-side limits, and prefer pre-authorized or resumable upload flows for the real design.

open as a page