skip to content

questions

22

HTTP/2 can carry many requests over a single TCP connection at the same time. Explain the mechanism that makes that possible, and what HTTP/1.1 had to do instead.

level: juniorimportance: must knowfreq 66%

answer

  1. 9-byte frame header, 31-bit stream ID
  2. odd = client, even = server, 0 = connection
  3. HEADERS + DATA interleave on one socket
  4. h1: ~6 connections + sharding
  5. removes app-layer HOL, not TCP-level

basics

~20 s

HTTP/2 chops every message into binary frames, each tagged with a stream ID. Frames from many streams interleave on one connection, so requests and responses overlap. HTTP/1.1 finished one message per connection at a time, so browsers opened about six connections per host.

solid answer

~50 s

HTTP/1.1 is text-based and strictly serial per connection: a response must finish before the next one starts on that connection. Browsers worked around it by opening ~6 parallel connections per host, and sites did domain sharding to get more. HTTP/2 adds a **binary framing layer**. Every message is split into frames with a 9-byte header carrying length, type, flags, and a 31-bit **stream ID**. A stream is an independent bidirectional exchange: one request/response pair. Because each frame says which stream it belongs to, the endpoints can interleave HEADERS and DATA frames from dozens of streams over one connection and reassemble them at the other end. So concurrency moves from "many connections" to "many streams on one connection": one TLS handshake, one congestion-control state, shared header-compression context, and no per-connection slow start for each new request.

code

http · 9 lines
http
HEADERS   stream=1  END_HEADERS   GET /index.html
HEADERS   stream=3  END_HEADERS   GET /app.css
HEADERS   stream=5  END_HEADERS   GET /app.js
HEADERS   stream=3  END_HEADERS   200 OK text/css
DATA      stream=3  <4096 bytes>
HEADERS   stream=1  END_HEADERS   200 OK text/html
DATA      stream=1  <8192 bytes>
DATA      stream=3  END_STREAM <900 bytes>
DATA      stream=1  END_STREAM <2048 bytes>

go deeper

for a junior

Be able to say: binary frames, stream IDs, many requests at once on one connection, instead of six connections in HTTP/1.1.

for a middle

Add the frame vocabulary (HEADERS/DATA/SETTINGS) and stream ID rules, and name what actually got faster: no repeated handshakes and slow start.

for a senior

Frame it as a change in the unit of cost — the connection is now long-lived and shared — and note the residual TCP head-of-line blocking plus the h1 optimizations that become anti-patterns.

for a principal

Discuss the system consequences: connection concentration on load balancers, per-connection memory limits, and the fact that stream concurrency exposes backend capacity limits that six connections used to hide.

## The problem HTTP/1.1 had In HTTP/1.1 a connection carries one request/response exchange at a time. There is no way to say which bytes belong to which message — the response body is simply the bytes after the headers until Content-Length is reached or the chunked encoding ends. Pipelining (sending request 2 before response 1 arrives) was specified but responses still had to come back in request order, so one slow response stalled everything behind it, and buggy proxies made pipelining unsafe. Browsers gave up on it and opened roughly six TCP connections per host instead, and developers spread assets over extra hostnames ("domain sharding") to multiply that budget. Each extra connection costs a handshake, its own TCP slow start, and memory on the server. ## The binary framing layer HTTP/2 keeps HTTP's semantics — the same methods, status codes, and header fields — but replaces the wire serialization. Everything is a **frame**: a 9-byte header followed by a payload. The header holds a 24-bit length, an 8-bit type, 8-bit flags, and a 31-bit stream identifier. The frame types you should be able to name: - **HEADERS** — carries a compressed header block; opens a stream. - **DATA** — carries a message body. - **SETTINGS** — connection-level configuration exchanged at startup and any time after; always acknowledged. - **WINDOW_UPDATE** — grants flow-control credit. - **RST_STREAM** — abruptly terminates one stream. - **GOAWAY** — tells the peer to stop opening streams; used for shutdown. - **PING** — liveness/RTT probe. - **PRIORITY** / **PUSH_PROMISE** — from the original design; both since deprecated in practice. ## Streams A **stream** is an independent, bidirectional sequence of frames within the connection, identified by its stream ID. Client-initiated streams use odd IDs, server-initiated use even, and stream 0 is reserved for connection-level frames (SETTINGS, connection WINDOW_UPDATE, GOAWAY, PING). IDs increase monotonically and are never reused; the 31-bit space is the reason a connection eventually has to be replaced. A typical exchange: the client sends HEADERS on stream 1 (request line and headers, compressed with HPACK), maybe DATA frames on stream 1 for the body, and END_STREAM to close its half. Meanwhile it can open streams 3, 5, 7 for other requests. The server replies with HEADERS then DATA on each stream, in whatever order it can produce them. The two endpoints interleave frames freely; each side reassembles per stream ID. ## What multiplexing buys 1. **True concurrency without extra sockets.** Dozens or hundreds of requests share one connection, and a slow response no longer blocks the others at the HTTP layer — this is the removal of *application-layer* head-of-line blocking. 2. **One handshake, one congestion window.** The connection stays warm, so later requests do not pay TCP slow start again. This is often the biggest real-world win. 3. **Shared compression state.** Repeated header fields cost a byte or two after the first request on the connection. 4. **Less server memory and fewer file descriptors** per client than six-plus connections. ## What it does not fix Multiplexing is an HTTP-layer property. Underneath, TCP still delivers one ordered byte stream, so a lost segment stalls delivery of *every* stream's bytes that follow it in the stream until retransmission. That transport-level head-of-line blocking is exactly what a different transport was later designed to remove. Multiplexing also does not make the server faster: if the backend can only produce four responses at a time, opening 100 streams just queues work. ## Practical consequences Because the connection is now the expensive, long-lived unit, HTTP/2 clients use **one connection per origin** and old HTTP/1.1 optimizations invert: domain sharding costs handshakes for no benefit, and aggressive concatenation of assets hurts cache granularity. Debugging changes too — you cannot read the wire with a text dump; you read a frame log (Chrome's net-export, `nghttp -v`, Wireshark's HTTP/2 dissector) and think in terms of stream IDs.

  • If HTTP/2 multiplexes, why do browsers still cap concurrency?
    The server advertises SETTINGS_MAX_CONCURRENT_STREAMS (commonly 100–128), and the client will not exceed it; extra requests queue in the client. Beyond that, per-stream and per-connection flow-control windows limit how much data can be in flight, and the shared TCP congestion window limits actual throughput. Unlimited streams would not make the backend produce responses any faster.
  • Does HTTP/2 require TLS?
    The specification allows cleartext HTTP/2 (h2c), but every major browser only speaks HTTP/2 over TLS and negotiates it with ALPN during the handshake. In practice you see h2c only between infrastructure components, such as a load balancer and an internal service.

HTTP/1.1 is one lane where each truck must fully pass before the next enters. HTTP/2 breaks the cargo into labelled pallets from many trucks and sends them mixed down the same lane; the receiver sorts them by label.

saying these in an interview costs you the question

  • Claiming HTTP/2 is just HTTP/1.1 pipelining done properly — pipelining forced in-order responses, multiplexing does not
  • Saying HTTP/2 eliminates head-of-line blocking entirely, ignoring the TCP layer underneath
  • Thinking HTTP/2 changed methods, status codes, or header semantics — only the serialization changed
  • Claiming domain sharding still helps on HTTP/2, when it splits the single connection and adds handshakes

context

open as a page

HTTP/3 runs over QUIC instead of TCP. What is QUIC, why was it built on UDP, and where does TLS fit into it?

level: juniorimportance: must knowfreq 55%

basics

~20 s

QUIC is a reliable, multiplexed transport implemented on top of UDP, with TLS 1.3 built into the protocol rather than layered above it. UDP was chosen because it passes through existing networks and lets the transport live in user space, so it can evolve without OS or middlebox changes.

open as a page

Walk through how HPACK, the header compression format used by HTTP/2, encodes a header field. What do the static table, the dynamic table and Huffman coding each contribute?

level: middleimportance: must knowfreq 45%

basics

~20 s

HPACK encodes each field as an index or a literal. A 61-entry static table covers common fields; a per-connection dynamic table (FIFO, byte-bounded) holds fields already seen, so repeats become one index byte. Literal strings are Huffman-coded. Both peers must keep tables in sync.

open as a page

On an HTTP/2 connection carrying ten concurrent requests, one TCP segment is lost. What happens to the other nine responses, and how does QUIC's per-stream loss recovery change the outcome for HTTP/3?

level: middleimportance: must knowfreq 48%

basics

~20 s

TCP delivers one ordered byte stream, so the kernel holds back every byte after the lost segment — all nine other responses stall for a retransmission round trip even though their data arrived. QUIC tracks loss per packet and delivers each stream independently, so only streams with data in the lost packet wait.

open as a page

A browser opens https://example.com. Explain the mechanism by which it ends up speaking HTTP/2 rather than HTTP/1.1, and how it would ever discover that HTTP/3 is available on that origin.

level: middleimportance: must knowfreq 45%

basics

~20 s

HTTP/2 is chosen by ALPN inside the TLS handshake: the client offers h2 and http/1.1, the server picks one, at no extra round trip. HTTP/3 cannot be negotiated that way because it needs a QUIC connection first, so the server advertises it with an Alt-Svc response header or a DNS HTTPS record.

open as a page

A team enabled HTTP/2 on their servers and page-load times barely moved. What kinds of workloads does HTTP/2 actually speed up, and which ones does it leave essentially unchanged?

level: middleimportance: must knowfreq 42%

basics

~20 s

HTTP/2 helps when many small requests go to one origin over a high-latency link: one connection, no six-connection cap, no client-side queueing, compressed headers. It barely helps a few large transfers, low-request-count APIs, or fast local networks, where the bottleneck is bandwidth or the server.

open as a page

HTTP/1.1 sends request headers as plain ASCII text on every single request. What changed in HTTP/2, and why was compressing header fields considered worth the extra complexity?

level: juniorimportance: should knowfreq 40%

basics

~20 s

HTTP/2 encodes headers in binary and compresses them with HPACK. Real requests repeat almost identical headers (cookies, user-agent, accept) worth hundreds of bytes each. On pages making many small requests that overhead dominated; HPACK turns repeats into tiny index references.

open as a page

On an HTTP/2 connection, what limits how many requests a client can have in flight at once, and how does a peer change that limit?

level: middleimportance: should knowfreq 45%

basics

~20 s

The server advertises SETTINGS_MAX_CONCURRENT_STREAMS in a SETTINGS frame; the client must not have more open streams than that. It can be changed at any time by sending a new SETTINGS frame. Excess streams are refused with RST_STREAM (REFUSED_STREAM) and are safe to retry.

open as a page

How does an HTTP client tell a server which responses it wants delivered first, and how did that mechanism change from HTTP/2's original dependency-tree design to RFC 9218 Extensible Priorities with its u= and i parameters?

level: middleimportance: should knowfreq 32%

basics

~20 s

HTTP/2 originally used a tree of stream dependencies and weights sent in PRIORITY frames. It was complex and barely implemented, so it was deprecated. RFC 9218 replaces it with a simple Priority header field: u=0..7 urgency (lower is more urgent, default 3) and i for incremental delivery.

open as a page

HTTP/2 server push let a server send a response the client never asked for, announced with a PUSH_PROMISE frame. Explain how it worked and why browsers such as Chrome ended up removing support for it.

level: middleimportance: should knowfreq 36%

basics

~20 s

The server sent PUSH_PROMISE reserving an even stream ID, then delivered a response the client had not requested. Servers could not know what the client already had cached, so most pushes were wasted bytes competing with the HTML for bandwidth. Chrome removed it in 2022.

open as a page

QUIC identifies a connection by connection IDs rather than by the source IP address and port. What does that enable when a phone moves from Wi-Fi to cellular, and what routing and privacy issues follow?

level: middleimportance: should knowfreq 30%

basics

~20 s

Because the connection is keyed by a connection ID, a client that changes IP address keeps the same QUIC connection: no new handshake, no lost streams. The server validates the new path first. Load balancers must route by connection ID, and IDs are rotated so observers cannot link a user across networks.

open as a page

Early HTTP/2 work (SPDY) compressed header fields with gzip, and that approach was abandoned. What attack made generic compression of HTTP headers unsafe, and how does HPACK handle sensitive header values?

level: seniorimportance: should knowfreq 30%

basics

~20 s

CRIME. When a secret (a session cookie) and attacker-chosen text are compressed together, the compressed size leaks whether they match, letting the attacker guess the secret one character at a time through TLS. HPACK removes substring matching - it indexes whole field values only - and adds a never-indexed representation for secrets.

open as a page

HTTP/3 does not reuse HPACK; it defines QPACK instead. What property of QUIC made HPACK unusable, and what does QPACK do differently to avoid the problem?

level: seniorimportance: should knowfreq 28%

basics

~20 s

HPACK's dynamic table is mutated in strict order, which TCP guarantees but QUIC does not across streams: a header block referencing an entry from a delayed stream would stall everything. QPACK moves table updates onto a dedicated ordered encoder stream and lets the encoder bound how many streams may block.

open as a page

HTTP/2 defines its own flow-control windows on top of TCP. Explain how the per-stream and per-connection windows work, and what symptom appears when the window is left at the default 65,535 bytes on a high-bandwidth, high-latency link.

level: seniorimportance: should knowfreq 34%

basics

~20 s

Each stream and the whole connection have a credit window in bytes; DATA frames consume it and WINDOW_UPDATE frames replenish it. With the 65,535-byte default on a long fat link, a sender stalls waiting for credit each round trip, capping throughput at roughly window divided by RTT regardless of available bandwidth.

open as a page

In HTTP/2, what do stream identifiers signify, and how do the RST_STREAM and GOAWAY frames let a server shut a connection down without discarding requests that are already being processed?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Stream IDs are monotonically increasing: odd from the client, even from the server, 0 for connection frames. RST_STREAM kills one stream. GOAWAY names the highest stream the server will process, so anything above it is untouched and safely retryable. Graceful shutdown sends GOAWAY twice.

open as a page

You want to use HTTP status 103 Early Hints in production to speed up page loads. Describe how it travels through the request path, what must be true for it to help, and what can break.

level: seniorimportance: should knowfreq 26%

basics

~20 s

The origin sends a 103 informational response with Link: rel=preload or rel=preconnect fields, then later the real response. It only helps when origin think-time is large enough to cover a fetch, every hop forwards 1xx responses, and the hints match what the page actually needs. Wrong hints waste bandwidth.

open as a page

QUIC with TLS 1.3 lets a resuming client send application data in its very first flight, known as 0-RTT. What does that buy, and what must a server do about replay?

level: seniorimportance: should knowfreq 32%

basics

~20 s

0-RTT saves one round trip on resumed connections by sending data under keys derived from a previous session. That data has no replay protection and is not forward-secret, so servers must accept only safe, idempotent requests in it, or reject with HTTP status 425 Too Early and make the client retry after the handshake.

open as a page

On what kind of network can moving from HTTP/1.1 to HTTP/2 make performance worse rather than better, and what is the underlying reason?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Lossy networks. HTTP/2 puts every request on one TCP connection, and TCP delivers bytes strictly in order, so one lost segment stalls all multiplexed streams. HTTP/1.1's six connections isolate the damage to one sixth of the traffic. HTTP/3 over QUIC fixes this with independent streams.

open as a page

HTTP/2 clients normally keep one connection per origin and may coalesce several hostnames onto the same connection. Explain when coalescing is allowed, and what a single long-lived connection per client means for a service sitting behind load balancers.

level: principalimportance: should knowfreq 26%

basics

~20 s

A client may reuse a connection for another hostname when that host resolves to the same IP address and the presented certificate is authoritative for it. One durable connection per client means load balancing happens once, at connect time, so servers must age connections out with GOAWAY to rebalance, and a server can reject misrouted requests with HTTP status 421.

open as a page

HTTP/2 over cleartext, known as h2c, is defined in the specification but almost never seen on the public web. Why is that, and where does cleartext HTTP/2 still get used?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

No browser implements h2c, so public traffic never uses it. Without TLS there is no ALPN, so h2c needs prior knowledge or an Upgrade handshake, and middleboxes mangle unfamiliar cleartext traffic. It survives inside trusted networks: gRPC between services and proxy-to-origin hops.

open as a page

With HTTP/2 server push removed from browsers, how would you decide between HTTP 103 Early Hints, Link: rel=preload on the final response, inlining critical CSS, and doing nothing, for delivering render-blocking resources on a high-traffic site?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

Decide by where the time actually goes. Long origin think-time favours 103 Early Hints; a fast origin makes preload on the final response the cheap default; a tiny, stable critical set justifies inlining despite losing cacheability; otherwise do nothing and fix caching or the number of critical resources instead.

open as a page

Your product already runs on HTTP/2 behind a CDN. How would you decide whether enabling HTTP/3 is worth it, and how would you know afterwards whether it actually helped?

level: principalimportance: nice to knowfreq 20%

basics

~20 s

Decide from your traffic mix: HTTP/3 pays off for lossy, high-latency mobile users making many requests, and little for datacenter or clean-broadband traffic. Behind a CDN it is usually a cheap, reversible toggle. Measure field percentiles by network class, plus h3 share and QUIC fallback rate.

open as a page