skip to content

On an HTTP/2 connection, what limits how many requests a client can have in flight at once, and how does a peer change that limit?

level: middleimportance: should knowfreq 45%

answer

  1. SETTINGS_MAX_CONCURRENT_STREAMS (0x3)
  2. no default limit; RFC says don't go below 100
  3. open + half-closed count; closed don't
  4. REFUSED_STREAM = not processed = safe retry
  5. lowering it doesn't kill open streams

basics

~20 s

The server advertises SETTINGS_MAX_CONCURRENT_STREAMS in a SETTINGS frame; the client must not have more open streams than that. It can be changed at any time by sending a new SETTINGS frame. Excess streams are refused with RST_STREAM (REFUSED_STREAM) and are safe to retry.

solid answer

~50 s

Concurrency is governed by **SETTINGS_MAX_CONCURRENT_STREAMS**, sent in the SETTINGS frame each side transmits at connection start (and may resend later; the peer must ACK). It counts streams in the open or half-closed states. The protocol default is unlimited, but the RFC says do not advertise below 100, and typical servers pick 100–128. If a client exceeds the limit, the server answers RST_STREAM with REFUSED_STREAM, which guarantees the request was not processed, so the client may safely retry it on the same or a new connection. When a server *lowers* the setting, existing streams above the new limit are not killed; the client simply cannot open more until it drops below. Separately, flow-control windows limit bytes in flight, not stream count, and clients queue requests locally once the limit is hit. So "unlimited concurrency" is never true — it is a negotiated budget backed by server memory.

code

http · 4 lines
http
S->C  SETTINGS  MAX_CONCURRENT_STREAMS=100, INITIAL_WINDOW_SIZE=65535
C->S  SETTINGS  ACK
C->S  HEADERS   stream=201  GET /img/42.png
S->C  RST_STREAM stream=201  error=REFUSED_STREAM

go deeper

for a junior

Know the name of the setting and that the server decides the cap, with extra requests queuing in the client.

for a middle

Explain the SETTINGS/ACK exchange, which stream states count, and that REFUSED_STREAM is retry-safe.

for a senior

Separate stream count from flow control and backend concurrency when diagnosing, and check the proxy's upstream limit as well as the edge one.

for a principal

Treat the limit as a capacity contract per connection and pair it with rate-based defences such as reset budgets, since concurrency caps alone did not stop Rapid Reset.

## The setting itself HTTP/2 endpoints exchange a **SETTINGS** frame immediately after the connection preface, and either side may send another at any time. SETTINGS is a list of 16-bit identifiers with 32-bit values; the receiver must acknowledge with an empty SETTINGS frame carrying the ACK flag. The relevant parameter here is `SETTINGS_MAX_CONCURRENT_STREAMS` (0x3). Its value tells the *peer* how many streams that peer may have open at once toward the sender of the setting. The specification defines no initial limit, which formally means "unlimited until told otherwise", and it explicitly advises against advertising a value below 100 because low values throttle parallelism badly. Real servers commonly settle on 100 (nginx `http2_max_concurrent_streams` default 128 in recent versions, Go's `http2.Server.MaxConcurrentStreams` default 250 for servers, etc.). Because there is no initial value, a client that starts blasting streams before the server's SETTINGS arrives can overshoot; that is legal and handled by refusal. ## Which streams count Streams in the **open**, **half-closed (local)**, and **half-closed (remote)** states count against the limit. Streams in **idle**, **reserved**, or **closed** do not. So a request whose body has been fully sent but whose response has not arrived still occupies a slot — the limit really measures outstanding work. ## What happens on overshoot The server responds to the offending HEADERS with **RST_STREAM** carrying either `REFUSED_STREAM` or `PROTOCOL_ERROR`. `REFUSED_STREAM` is the useful one: it is a promise that the request was **not** processed, which makes it retry-safe even for non-idempotent methods. A stricter server may treat repeated overshoot as a connection error and send GOAWAY. A client that is at the limit does not error — it queues the request in its own connection pool. This is why you can see queuing time in browser dev tools even on HTTP/2: 300 image requests against a 100-stream limit means two-thirds are waiting in the client. ## Interaction with the rest of the connection Stream count and byte budget are different knobs: - `SETTINGS_MAX_CONCURRENT_STREAMS` limits how many exchanges are open. - `SETTINGS_INITIAL_WINDOW_SIZE` and WINDOW_UPDATE limit how many *bytes* may be in flight per stream and per connection. - The TCP congestion window limits what actually reaches the network. All 100 streams can be open while only a few are transferring data, because flow control and the backend's own concurrency decide who makes progress. ## Lowering the limit A server under pressure can send a new SETTINGS frame with a smaller value. Streams already open above the new value are **not** terminated — the RFC requires that the change only prevents *new* streams. If a server truly needs to shed load it must additionally reset streams or send GOAWAY. Clients should treat the value as advisory-until-ACKed: the setting applies from the moment it is received, not from the ACK. ## Security angle: Rapid Reset The stream-count limit is not by itself a work limit. In the HTTP/2 Rapid Reset attack (CVE-2023-44487, 2023) a client opens a stream, immediately sends RST_STREAM, and repeats. Each reset stream leaves the counted set instantly, so the attacker never exceeds MAX_CONCURRENT_STREAMS while causing the server to start unbounded backend work. Mitigations added since: count recently-reset streams against a budget, cap resets per connection per second, or close the connection with GOAWAY (ENHANCE_YOUR_CALM) when the reset rate is abusive. Mention this if an interviewer asks whether the setting protects the server — it constrains concurrency, not request rate. ## Practical tuning Raising the limit is rarely the fix for a slow page; the bottleneck is normally backend concurrency, flow-control windows, or bandwidth. Lower it if per-stream memory (header tables, buffers) is your constraint, or if a single abusive client can occupy your worker pool. If you sit behind a proxy, check the *proxy's* limit too — the client-facing and upstream connections each negotiate their own value, and a proxy with a low upstream limit silently serializes traffic that looked parallel at the edge.

  • A client hits the limit — what should it do?
    Queue locally and wait for a stream slot to free; that is what browsers do. Opening a second connection to the same origin is allowed but defeats the point of one connection per origin and is normally only done when the server advertises a very low limit. Requests refused with REFUSED_STREAM can be retried immediately because the server promises it did not process them.
  • Does a high MAX_CONCURRENT_STREAMS protect the server from overload?
    No. It bounds simultaneously open streams, not the rate of work. The Rapid Reset attack showed that a client can open and immediately reset streams so the open count stays tiny while backend work grows without bound. Real protection needs reset-rate accounting, per-connection request budgets, and load shedding.

saying these in an interview costs you the question

  • Saying HTTP/2 gives unlimited parallel requests with no negotiated cap
  • Confusing the stream-count limit with flow-control windows, which bound bytes not streams
  • Believing a lowered SETTINGS value terminates streams already open
  • Treating any RST_STREAM as unsafe to retry — REFUSED_STREAM explicitly means the request was not processed

context