skip to content

HTTP/2 can carry many requests over a single TCP connection at the same time. Explain the mechanism that makes that possible, and what HTTP/1.1 had to do instead.

level: juniorimportance: must knowfreq 66%

answer

  1. 9-byte frame header, 31-bit stream ID
  2. odd = client, even = server, 0 = connection
  3. HEADERS + DATA interleave on one socket
  4. h1: ~6 connections + sharding
  5. removes app-layer HOL, not TCP-level

basics

~20 s

HTTP/2 chops every message into binary frames, each tagged with a stream ID. Frames from many streams interleave on one connection, so requests and responses overlap. HTTP/1.1 finished one message per connection at a time, so browsers opened about six connections per host.

solid answer

~50 s

HTTP/1.1 is text-based and strictly serial per connection: a response must finish before the next one starts on that connection. Browsers worked around it by opening ~6 parallel connections per host, and sites did domain sharding to get more. HTTP/2 adds a **binary framing layer**. Every message is split into frames with a 9-byte header carrying length, type, flags, and a 31-bit **stream ID**. A stream is an independent bidirectional exchange: one request/response pair. Because each frame says which stream it belongs to, the endpoints can interleave HEADERS and DATA frames from dozens of streams over one connection and reassemble them at the other end. So concurrency moves from "many connections" to "many streams on one connection": one TLS handshake, one congestion-control state, shared header-compression context, and no per-connection slow start for each new request.

code

http · 9 lines
http
HEADERS   stream=1  END_HEADERS   GET /index.html
HEADERS   stream=3  END_HEADERS   GET /app.css
HEADERS   stream=5  END_HEADERS   GET /app.js
HEADERS   stream=3  END_HEADERS   200 OK text/css
DATA      stream=3  <4096 bytes>
HEADERS   stream=1  END_HEADERS   200 OK text/html
DATA      stream=1  <8192 bytes>
DATA      stream=3  END_STREAM <900 bytes>
DATA      stream=1  END_STREAM <2048 bytes>

go deeper

for a junior

Be able to say: binary frames, stream IDs, many requests at once on one connection, instead of six connections in HTTP/1.1.

for a middle

Add the frame vocabulary (HEADERS/DATA/SETTINGS) and stream ID rules, and name what actually got faster: no repeated handshakes and slow start.

for a senior

Frame it as a change in the unit of cost — the connection is now long-lived and shared — and note the residual TCP head-of-line blocking plus the h1 optimizations that become anti-patterns.

for a principal

Discuss the system consequences: connection concentration on load balancers, per-connection memory limits, and the fact that stream concurrency exposes backend capacity limits that six connections used to hide.

## The problem HTTP/1.1 had In HTTP/1.1 a connection carries one request/response exchange at a time. There is no way to say which bytes belong to which message — the response body is simply the bytes after the headers until Content-Length is reached or the chunked encoding ends. Pipelining (sending request 2 before response 1 arrives) was specified but responses still had to come back in request order, so one slow response stalled everything behind it, and buggy proxies made pipelining unsafe. Browsers gave up on it and opened roughly six TCP connections per host instead, and developers spread assets over extra hostnames ("domain sharding") to multiply that budget. Each extra connection costs a handshake, its own TCP slow start, and memory on the server. ## The binary framing layer HTTP/2 keeps HTTP's semantics — the same methods, status codes, and header fields — but replaces the wire serialization. Everything is a **frame**: a 9-byte header followed by a payload. The header holds a 24-bit length, an 8-bit type, 8-bit flags, and a 31-bit stream identifier. The frame types you should be able to name: - **HEADERS** — carries a compressed header block; opens a stream. - **DATA** — carries a message body. - **SETTINGS** — connection-level configuration exchanged at startup and any time after; always acknowledged. - **WINDOW_UPDATE** — grants flow-control credit. - **RST_STREAM** — abruptly terminates one stream. - **GOAWAY** — tells the peer to stop opening streams; used for shutdown. - **PING** — liveness/RTT probe. - **PRIORITY** / **PUSH_PROMISE** — from the original design; both since deprecated in practice. ## Streams A **stream** is an independent, bidirectional sequence of frames within the connection, identified by its stream ID. Client-initiated streams use odd IDs, server-initiated use even, and stream 0 is reserved for connection-level frames (SETTINGS, connection WINDOW_UPDATE, GOAWAY, PING). IDs increase monotonically and are never reused; the 31-bit space is the reason a connection eventually has to be replaced. A typical exchange: the client sends HEADERS on stream 1 (request line and headers, compressed with HPACK), maybe DATA frames on stream 1 for the body, and END_STREAM to close its half. Meanwhile it can open streams 3, 5, 7 for other requests. The server replies with HEADERS then DATA on each stream, in whatever order it can produce them. The two endpoints interleave frames freely; each side reassembles per stream ID. ## What multiplexing buys 1. **True concurrency without extra sockets.** Dozens or hundreds of requests share one connection, and a slow response no longer blocks the others at the HTTP layer — this is the removal of *application-layer* head-of-line blocking. 2. **One handshake, one congestion window.** The connection stays warm, so later requests do not pay TCP slow start again. This is often the biggest real-world win. 3. **Shared compression state.** Repeated header fields cost a byte or two after the first request on the connection. 4. **Less server memory and fewer file descriptors** per client than six-plus connections. ## What it does not fix Multiplexing is an HTTP-layer property. Underneath, TCP still delivers one ordered byte stream, so a lost segment stalls delivery of *every* stream's bytes that follow it in the stream until retransmission. That transport-level head-of-line blocking is exactly what a different transport was later designed to remove. Multiplexing also does not make the server faster: if the backend can only produce four responses at a time, opening 100 streams just queues work. ## Practical consequences Because the connection is now the expensive, long-lived unit, HTTP/2 clients use **one connection per origin** and old HTTP/1.1 optimizations invert: domain sharding costs handshakes for no benefit, and aggressive concatenation of assets hurts cache granularity. Debugging changes too — you cannot read the wire with a text dump; you read a frame log (Chrome's net-export, `nghttp -v`, Wireshark's HTTP/2 dissector) and think in terms of stream IDs.

  • If HTTP/2 multiplexes, why do browsers still cap concurrency?
    The server advertises SETTINGS_MAX_CONCURRENT_STREAMS (commonly 100–128), and the client will not exceed it; extra requests queue in the client. Beyond that, per-stream and per-connection flow-control windows limit how much data can be in flight, and the shared TCP congestion window limits actual throughput. Unlimited streams would not make the backend produce responses any faster.
  • Does HTTP/2 require TLS?
    The specification allows cleartext HTTP/2 (h2c), but every major browser only speaks HTTP/2 over TLS and negotiates it with ALPN during the handshake. In practice you see h2c only between infrastructure components, such as a load balancer and an internal service.

HTTP/1.1 is one lane where each truck must fully pass before the next enters. HTTP/2 breaks the cargo into labelled pallets from many trucks and sends them mixed down the same lane; the receiver sorts them by label.

saying these in an interview costs you the question

  • Claiming HTTP/2 is just HTTP/1.1 pipelining done properly — pipelining forced in-order responses, multiplexing does not
  • Saying HTTP/2 eliminates head-of-line blocking entirely, ignoring the TCP layer underneath
  • Thinking HTTP/2 changed methods, status codes, or header semantics — only the serialization changed
  • Claiming domain sharding still helps on HTTP/2, when it splits the single connection and adds handshakes

context