skip to content

A WebSocket endpoint buffers each message's fragments until it is complete. What does a peer sending an enormous message cost it, and how can it refuse?

level: seniorimportance: should knowfreq 40%

answer

  1. the sender chooses how much you allocate
  2. no total length is ever announced
  3. count as you append, not after
  4. stop before the final fragment arrives
  5. close status 1009, message too big

basics

~20 s

A fragmented WebSocket message declares no total length, so an unbounded reassembly buffer is a memory-exhaustion vector: one peer, one message. Count bytes as fragments arrive and close with status 1009 once the running total passes your limit.

solid answer

~50 s

Fragmentation exists so a sender can start a message before it knows how long it will be, which means the receiver never learns the total in advance - only each fragment's own size. An endpoint that buffers until the final fragment is therefore agreeing to allocate whatever the peer decides to send, and one peer can exhaust its memory with a single message it simply never ends. The defence is a declared maximum message size enforced against a running byte count, closing with status `1009` (message too big) the moment the total crosses it rather than discovering the size after allocating it. The alternative is not to buffer at all and consume fragments incrementally, which caps peak memory at one fragment but forces the application to cope with work it has already done on a message that turns out to be truncated or invalid.

code

pseudocode · 19 lines
pseudocode
MAX MESSAGE BYTES = 1048576

buffer = empty
total = 0

on data fragment(f):
    total = total + length of f.payload

    if total > MAX MESSAGE BYTES:
        discard buffer
        close connection with status 1009
        return            # do not wait for a final fragment

    append f.payload to buffer

    if f is the final fragment:
        deliver message(buffer)
        buffer = empty
        total = 0

go deeper

for a junior

Know that a message can arrive in fragments and that nothing tells the receiver how many there will be. Buffering until the end means trusting the sender about how much memory you will use.

for a middle

Explain the enforcement: keep a running byte total across fragments, abandon the message the moment it crosses the configured maximum, and close with status 1009 rather than waiting for a final fragment that may never arrive.

for a senior

Show that you have operated this. One connection can take a whole node down because the memory is the process's, not the connection's, and the traffic is indistinguishable from a large legitimate upload until the moment it is not.

for a principal

The real decision is which payloads are allowed to be large at all. A command channel with a hard, small ceiling and a separate path for bulk data is easier to reason about than one channel whose limit must satisfy both.

## A fragmented message has no declared total length When a WebSocket message is sent in one piece, its size is known before the payload arrives. When it is **fragmented**, it is not: each fragment carries its own length and nothing carries the sum. That is not an oversight - fragmentation exists so a sender can begin a message it has not finished producing, streaming a file or a serialisation as it goes. The consequence for the receiver is blunt. Buffering fragments until the message is complete means committing to allocate an amount of memory that the **peer** chooses, that is never announced, and that has no protocol-defined ceiling. A hostile or merely careless peer can open one message, send fragments indefinitely, and never send the final one. ## The buffer is the vulnerability This is one of the cheapest denial-of-service shapes there is against a long-lived connection tier, because it costs the attacker almost nothing: - It needs **one** connection rather than a flood of them, so connection-rate defences never see it. - It needs no credential beyond whatever opened the socket. - It looks like normal traffic - a large upload in progress - right up to the moment the process runs out of memory. - The damage is not scoped to that connection: the memory consumed is the whole process's, so every other connection the node holds dies with it. And it happens accidentally too. A client that serialises an unexpectedly large structure, or relays a payload it did not generate, sends exactly the same traffic with none of the intent. ## Refusing: a running count and a close The defence is a **declared maximum message size** the endpoint enforces itself, because the protocol will not enforce one for you. The check has to be incremental: 1. Decide the limit from what the application actually needs to accept per message, not from what happens to fit in memory today. 2. Keep a running byte total as each fragment is appended. 3. The moment the total crosses the limit, stop buffering and free what you have - do not wait for the final fragment, which may never come. 4. Close the connection with status **1009**, the status defined for a message too big for the endpoint to process, so the peer learns why rather than guessing at an abrupt disconnect. Step 3 is the one people get wrong: a limit checked only after reassembly completes is barely a limit at all, because the allocation it was meant to prevent has already happened. ## Or do not buffer: consuming incrementally The other shape hands each fragment to a consumer as it arrives - a digest, a streaming parser, a write to storage - so peak memory is one fragment regardless of message size. It is strictly better on memory and strictly worse on simplicity. | | Buffer the whole message | Consume fragments incrementally | |---|---|---| | Peak memory | the largest message you allow | one fragment | | Limit enforcement | running count, close at the threshold | still bound the total, elsewhere | | The application sees | one complete, validated message | a partial message that may never complete | | Truncation | never reaches the application | must be handled explicitly | | Late invalidity | rejected before anyone acts on it | discovered after work has begun | | Complexity | low | real, and it lands in the application | The last three rows are the price: - **Truncation.** The peer can vanish mid-message. Something has to decide what a half-processed message means, and be able to undo or discard the work already done. - **Late invalidity.** A text message's encoding is a property of the whole message, so a fragment that arrives last can condemn everything you already consumed. - **Partial side effects.** If the first fragment moved a crane, the last fragment failing to arrive is an operational problem rather than a parsing one. ## Choosing For control and command traffic - small, structured, acted on as a unit - buffer the whole message and set a limit legitimate traffic would never reach. For genuinely large payloads, consume incrementally and make the partial-message case explicit in the application rather than pretending it cannot happen. What is never acceptable is the third option people actually ship: buffer everything, check nothing, and trust the peer to be reasonable.

  • Why is a limit checked after the message is reassembled worse than one checked as it arrives?
    Because the allocation the limit exists to prevent has already happened by the time it runs. It protects the code downstream of reassembly, not the memory of the process. The check has to run against a running total as fragments are appended, and abandon the message mid-flight.
  • What should the endpoint do with the fragments it has already buffered when the limit is crossed?
    Free them immediately and stop reading that message. Holding them while the close completes keeps the memory pinned for exactly as long as the peer wants to delay, which reproduces the problem the limit was meant to solve.
  • Does incremental consumption remove the need for a size limit?
    It removes the memory cliff, not the need for a bound. A message that never ends still consumes whatever the consumer accumulates downstream - storage, a digest's worth of processing, a held transaction - so a total-bytes ceiling and a maximum duration are still worth enforcing, just at a different layer.

saying these in an interview costs you the question

  • Assumes a fragmented message announces its total size
  • Checks the size limit only after reassembly completes
  • Believes the protocol caps message size by default
  • Thinks the blast radius is limited to that one connection
  • Treats an unended message as impossible without malice