A cue desk holds 40,000 WebSocket connections with `permessage-deflate` on and memory climbs per socket — which negotiated parameters change that, and what do you give up?
answer
- the default keeps state
- per direction, per socket
- two parameters reset, two shrink
- ratio traded for memory
- agreed once, never renegotiated
basics
~20 sBy default each endpoint keeps its deflate context between messages, so every socket holds compressor state in each direction. server_no_context_takeover and client_no_context_takeover reset it per message, and the max_window_bits parameters shrink it; both cost compression ratio.
solid answer
~40 sDeflate compresses by referring back to bytes it has already seen, so by default an endpoint keeps its **compression context** across messages on a connection — that is context takeover, and it is where the ratio on repetitive messages comes from. The price is per-socket state in each direction, which one connection never notices and forty thousand do. Four parameters, agreed once in `Sec-WebSocket-Extensions`, move the dial: `server_no_context_takeover` and `client_no_context_takeover` make that side reset its compressor after every message it sends, and `server_max_window_bits` and `client_max_window_bits` cap the sliding window each side may use. Both reduce memory and both cost ratio, because a compressor that cannot look back at the previous message cannot exploit the fact that your messages are nearly identical. Nothing can be renegotiated later, so this is a handshake decision.
code
pseudocode · 8 linesfor each open socket:
for each direction that compresses:
if context takeover applies:
keep window and compressor tables until close
else:
reset compressor state after each message
resident = sockets * retained_directions * state_per_directiongo deeper
The takeaway is that compression is not free: an endpoint remembers earlier messages in order to compress the next one, and remembering costs memory on every connection it holds open.
Explain context takeover as a deflate window kept across messages, name the parameters that reset it or cap it, and state that they are agreed in the handshake and cannot change afterwards.
Do the multiplication out loud — per direction, per socket, times the fleet — and separate this fixed per-socket cost from queue growth behind a slow reader, because the two look identical on a memory graph.
This is a bandwidth-against-memory bet taken once, for every client that will ever connect, and unwound only at the speed your fleet reconnects. Decide it with the connection count you expect, not the one you have.
## What context takeover actually is Deflate works by referring back to byte sequences it has already emitted, using a sliding window of recent history. Across a WebSocket connection, an endpoint may keep that window and the compressor state alive **from one message to the next**. That is context takeover, and it is the default: nothing in the negotiation says otherwise unless a parameter does. It is also where most of the ratio comes from on a realtime feed. Successive messages on such a connection are usually near-duplicates of each other — the same structure, the same field names, a handful of changed values. A compressor that remembers the last message encodes the next one as a short set of references. A compressor that has been reset sees each message cold and must describe the whole structure again. ## The four parameters | Parameter | Effect when agreed | |---|---| | `server_no_context_takeover` | the server resets its compressor after each message it sends | | `client_no_context_takeover` | the client resets its compressor after each message it sends | | `server_max_window_bits` | caps the window the server's compressor uses, at most 15, i.e. 32 KiB | | `client_max_window_bits` | caps the window the client's compressor uses, on the same scale | Four things about how they are exchanged: - They are carried as parameters on a `permessage-deflate` offer in `Sec-WebSocket-Extensions`, and the server returns the set that will actually apply on the `101 Switching Protocols` response. - A client may offer `client_max_window_bits` with no value at all, which says "I support being capped; you choose the number". - A server must not put `client_max_window_bits` in its response unless the client's offer mentioned it. - None of them can be changed later. There is no renegotiation on an open WebSocket connection. ## The arithmetic that makes this a senior question One connection's compression state is uninteresting. The cost is the multiplication: 1. State is held **per direction**. A socket that compresses in both directions holds two sets. 2. Each retained context holds its sliding window — up to 32 KiB at the maximum setting — plus the compressor's internal tables, which are several times the window. The **compressing** side is the expensive one; the decompressing side is roughly the window plus a little. 3. Multiply by every open socket. If retained state runs to the order of a hundred kilobytes per direction, forty thousand connections with both directions retained is on the order of several gigabytes, before a single message has been buffered for a slow reader. That is why the parameters exist at all: they were added precisely because a process holding many sockets cannot afford the default. ## What you give up, and how to choose - **Ratio, on exactly the payloads where compression was winning.** Repetitive structured messages are the case that benefits most from context takeover and therefore loses most when it is turned off. - **Predictability in the other direction, gained.** Per-message reset makes each message's compressed size independent of history, which is easier to reason about and removes a cross-message information leak. - **CPU is roughly unchanged.** Resetting a compressor is cheap; you are trading bytes for memory, not for processor time. The practical reading is by connection shape: - **Many sockets, small repetitive messages, memory-bound process:** ask for no context takeover on the side you operate, and cap the window. Accept the ratio loss; it is bounded and your memory is not. - **Few sockets, large or already-varied messages:** keep the context. The state is affordable and the ratio is worth having. - **Anything holding a serving tier open at scale:** decide this deliberately in the handshake rather than inheriting the default, because it cannot be changed once the socket is open and you will only discover the default under load. One more thing to check before blaming the extension: memory that climbs per socket also climbs when messages queue for readers that cannot keep up. Compression contexts are a fixed per-socket cost that appears as soon as the socket opens; a queue grows with the backlog. The two look alike in a coarse memory graph and have different fixes, so measure which one is moving before renegotiating anything.
- Why can't you just turn context takeover off once memory starts climbing?Because extension parameters are agreed in the opening handshake and the protocol defines no renegotiation. Existing connections keep what they negotiated until they close; a change takes effect only on sockets opened afterwards, so the fix arrives at the pace clients reconnect.
- Which side does `server_no_context_takeover` constrain, and who can ask for it?It constrains the server's own compressor: the server resets its context after each message it sends. It can appear in the client's offer, asking the server to behave that way, and in the server's response, where it states what the server will do.
- Does capping the window with `client_max_window_bits` save memory on both sides?It bounds the client's compression window, and the server sizes its decompression state to match, so both ends benefit — but the larger saving is on whichever side is compressing, because a compressor's tables are several times its window while a decompressor holds little beyond the window itself.
Keeping the previous cue sheet on the desk so the next one can be written as a handful of amendments. Cheap on one desk; expensive when the building holds forty thousand of them.
saying these in an interview costs you the question
- Thinks compression state is shared per process rather than per connection
- Says no-context-takeover saves CPU rather than memory
- Believes extension parameters can be changed on an open connection
- Assumes turning off context takeover leaves the compression ratio unchanged
- Confuses fixed per-socket compressor state with a growing send queue