skip to content

HTTP/2 clients normally keep one connection per origin and may coalesce several hostnames onto the same connection. Explain when coalescing is allowed, and what a single long-lived connection per client means for a service sitting behind load balancers.

level: principalimportance: should knowfreq 26%

answer

  1. one connection per origin (scheme+host+port)
  2. coalesce: same IP + certificate authoritative
  3. 421 Misdirected Request = retry on new connection
  4. ORIGIN frame widens/narrows the coalescing set
  5. sticky L4 routing → age connections out with GOAWAY

basics

~20 s

A client may reuse a connection for another hostname when that host resolves to the same IP address and the presented certificate is authoritative for it. One durable connection per client means load balancing happens once, at connect time, so servers must age connections out with GOAWAY to rebalance, and a server can reject misrouted requests with HTTP status 421.

solid answer

~60 s

HTTP/2 asks clients to use **one connection per origin**, and RFC 9113 §9.1.1 permits **coalescing**: reusing an existing connection for a different hostname if (a) the new host resolves to an IP address the connection is already established to, and (b) the server's certificate is authoritative for that host, typically via a SAN wildcard. Chrome and Firefox do this aggressively; a server that cannot serve the coalesced host answers **421 Misdirected Request**, and the client retries on a fresh connection. The ORIGIN frame (RFC 8336) lets a server declare exactly which origins it will serve, expanding or narrowing coalescing without a DNS round trip. The systems consequence: routing decisions become sticky. An L4 load balancer picks a backend once, then that client's entire session lives there — new instances after a scale-up receive nothing, and a hot client concentrates load on one node. Countermeasures: bounded connection age with a graceful GOAWAY drain, L7 proxies that rebalance per request, and per-connection stream/rate limits so one connection cannot monopolize a node.

code

http · 9 lines
http
C->S  HEADERS stream=9
      :method: GET
      :scheme: https
      :authority: other-tenant.example.com
      :path: /assets/app.js

S->C  HEADERS stream=9
      :status: 421
S->C  DATA    stream=9  END_STREAM

go deeper

for a junior

Know that HTTP/2 uses a single connection per origin and that opening extra hostnames no longer helps.

for a middle

State both coalescing conditions and what 421 Misdirected Request means for the client.

for a senior

Connect stickiness to real operational symptoms — idle new instances, skewed nodes, spiky deploys — and apply connection ageing or L7 balancing.

for a principal

Argue the isolation-versus-efficiency tradeoff explicitly: which traffic deserves its own origin, what connection-age policy and jitter you choose, where you terminate, and what retry budget makes drains invisible.

## The one-connection rule HTTP/2's specification says clients SHOULD NOT open more than one connection to a given origin (scheme + host + port). All the benefits — a warm congestion window, a shared HPACK table, coherent prioritization, low server memory — depend on traffic being concentrated. This reverses HTTP/1.1 practice, where six connections plus domain sharding was the tuning advice. ## Connection coalescing Browsers go further: they reuse one connection for *different* hostnames. RFC 9113 §9.1.1 allows this when both conditions hold: 1. **Same destination.** The new origin's DNS resolution yields an IP address the client already has a connection to. (Firefox is stricter than Chrome here about requiring a fresh resolution rather than trusting an existing certificate alone.) 2. **Authoritative certificate.** The certificate already presented on that connection covers the new host — usually a SAN entry or a wildcard such as `*.example.com` — and passes normal validation. So `img.example.com` and `api.example.com`, both behind the same CDN IP and the same wildcard certificate, share one connection and one TLS handshake. This is why domain sharding is now harmful: it either splits into several connections (extra handshakes, cold congestion windows) or coalesces anyway, giving you nothing for the added DNS lookups. Two escape hatches exist: - **421 Misdirected Request**: a server that receives a request whose `:authority` it is not configured to serve returns 421, and the client must retry on a new connection to that origin. This is the correct answer for multi-tenant edges where one certificate covers hosts served by different backends. - **The ORIGIN frame** (RFC 8336): a server sends the set of origins it is authoritative for on this connection, letting it *expand* coalescing (accept hosts the client would not have tried) or *contract* it (avoid a 421 round trip entirely). ## What breaks at scale **Load balancing becomes a connect-time decision.** An L4 (TCP) balancer hashes or round-robins once. With HTTP/1.1 and six short connections per client, imbalance self-corrects; with one HTTP/2 connection that lives for hours, whatever backend you landed on serves every request. Consequences: - **Scale-up does nothing.** New pods sit idle because no client reconnects. Autoscaling appears broken. - **Skew.** A few heavy clients (mobile app fleets, service-to-service gRPC callers) can pile hundreds of concurrent streams onto one node while its peers idle. - **Deploys hurt more.** Killing an instance drops many multiplexed sessions at once, so the retry storm is spikier. **Mitigations, roughly in order of preference:** 1. **Bounded connection age.** Have servers send a graceful two-phase GOAWAY after a maximum age (plus jitter) so clients re-resolve DNS and re-balance. gRPC exposes this directly as `MAX_CONNECTION_AGE` + grace; Envoy and most edge proxies have an equivalent drain. 2. **Layer-7 balancing.** Terminate HTTP/2 at a proxy that balances *per request* to upstream pools, so client connection stickiness never reaches the application tier. 3. **Client-side load balancing.** For internal RPC, let the client hold connections to several backends (subsetting) and spread streams across them; this is standard gRPC practice. 4. **Per-connection budgets.** Cap concurrent streams, request rate, and reset rate per connection so one connection cannot exhaust a node — the same controls that mitigate Rapid Reset also bound noisy-neighbour effects. **Blast radius**: coalescing means a single connection may carry several hostnames' traffic. If that connection stalls (flow-control deadlock, a slow TCP path, a lost packet triggering transport-level head-of-line blocking), everything on it stalls together. Splitting genuinely independent, latency-critical traffic onto its own origin — with a certificate or ORIGIN frame that prevents coalescing — is a legitimate isolation decision, and one of the few remaining reasons to use a separate hostname. ## Judgement calls to raise in an interview - Isolation versus efficiency: coalescing saves handshakes but couples failure domains. - Connection age versus handshake cost: shorter ages rebalance better but pay TLS more often; jitter avoids synchronized reconnect storms. - Where to terminate: L4 keeps state cheap but inherits stickiness; L7 rebalances per request but costs CPU and adds a hop. - Who owns retries: a drain is only invisible if clients retry the streams above the GOAWAY cutoff and have a retry budget that stops the storm from amplifying.

  • How do you stop a browser from coalescing two of your hostnames onto one connection?
    Break one of the two conditions: serve the hostname from a different IP address, or from a certificate that is not authoritative for the other host — no shared wildcard or SAN entry. A server can also send the ORIGIN frame listing only the origins it will serve, and return 421 Misdirected Request for anything else so the client opens a separate connection.
  • Your autoscaler adds pods but traffic stays on the old ones. What is going on and how do you fix it?
    Long-lived HTTP/2 connections were balanced once at connect time, so existing clients never touch the new pods. Fix it by bounding connection age with a jittered graceful GOAWAY drain so clients reconnect and re-resolve, or by balancing at layer 7 per request, or by having internal clients hold a subset of backend connections and spread streams across them.

It is a single shared pipeline into one building rather than six separate deliveries: efficient, but everything for that customer now depends on one door — and once the door is chosen, no amount of opening new buildings redirects them.

saying these in an interview costs you the question

  • Recommending domain sharding for HTTP/2, which either splits the connection or coalesces anyway
  • Believing coalescing requires only a matching certificate, ignoring the same-IP condition
  • Treating 421 Misdirected Request as a client bug rather than the protocol's designed correction
  • Assuming an L4 load balancer keeps HTTP/2 traffic balanced without connection-age limits

context