skip to content

A recipient attached to the store's broadcast surface reads slower than messages are sent — where does the undelivered backlog sit, and what ends it?

level: seniorimportance: should knowfreq 46%

answer

  1. the sender is not the one paying
  2. undelivered bytes queue on the server
  3. same memory as the entries
  4. over the bound, connection closed
  5. resolved in the server's favour

basics

~20 s

On the server, in the delivery queue the store holds for that recipient's connection, drawing on the same memory as the data. Stores that bound those bytes close the slow recipient's connection: the tier is protected, the recipient sacrificed.

solid answer

~50 s

A broadcast is fire-and-forget for the sender but not for the server: the store has to write each copy out to each attached connection, and a recipient that reads slower than the sender sends leaves undelivered bytes queued on the store itself. That memory competes with the data in the same tier. Stores that offer a broadcast surface generally bound how much they will hold per recipient, and when a recipient stays over the bound the server closes its connection rather than keep growing — the argument is resolved in the server's favour. What varies is how the bound is expressed: a byte size, how long a recipient has been over a softer level, or nothing explicit at all, in which case the only real bound is the tier's own memory ceiling. The design lesson is that a bigger bound buys seconds and converts one slow recipient into pressure every other caller shares.

go deeper

for a junior

Remember that a broadcast still costs the server work: it has to write a copy to every attached recipient, and a recipient that is slow to read leaves that copy waiting on the server.

for a middle

Explain that the waiting bytes are held per connection on the store and come out of the same memory as the data, and that a store which bounds them closes the connection when a recipient stays over the bound.

for a senior

Show the production reflex: diagnose the flapping recipient, refuse the bigger-bound fix, cut payload size and send rate, and name the signals — connection memory and disconnect counts — you would alert on.

for a principal

Treat it as a shared-tier isolation question. Decide whether one team's slow recipient may endanger everybody's memory, and whether a consumer that truly needs every message belongs on a broker instead.

## Fire-and-forget for the sender, not for the server Sending a broadcast looks free to the sender: one call, no acknowledgement, no entry written. The work is on the other side. The store has to hand a copy of that message to every recipient currently attached to the name, and handing it over means writing bytes to a socket that the recipient must read. If a recipient reads more slowly than the sender sends — a browser on a bad link, a worker paused by garbage collection, a process stopped at a breakpoint — the bytes the recipient has not taken yet have to live somewhere until it does. They live **on the store**, in the delivery queue it keeps for that connection. That is the whole point of the question: a slow *recipient* consumes memory on the *server*, and that memory comes out of the same pool as the entries the tier is holding for everyone else. ## How the argument ends No server can grow that queue forever, so stores that offer a broadcast surface put a bound on it and enforce the bound by **closing the connection**. The slow recipient is disconnected; the tier keeps running. Stated as a principle: *the server resolves the conflict in its own favour*, because sacrificing one recipient is cheap and running the tier out of memory is not. What that costs the disconnected recipient is exactly the loss described by transient fan-out generally — everything sent while it is away is gone for it, and nothing tells it how much it missed. A client that reconnects automatically and is still slow will be disconnected again, which is the familiar production signature: a recipient that flaps rather than one that fails. ## What genuinely varies here - **Whether a broadcast surface exists at all.** Several stores in this class have none, and there the question does not arise. - **How the bound is expressed.** Some stores bound the bytes held for a recipient; some use a softer level plus a duration, so a brief burst is tolerated and sustained lag is not; some expose no explicit bound, leaving the tier's memory ceiling as the real one. - **Whether bounds differ by kind of connection.** Some stores treat a recipient of broadcasts differently from an ordinary request/response caller, because the two have very different traffic shapes. - **How the copies are held.** Implementations differ in whether the message is encoded once and shared between connections or copied per connection. Either way the *undelivered* bytes grow with the number of lagging recipients. None of that changes the shape of the answer, and a candidate who names the shape while flagging that stores differ on the bound is answering at the right level. ## Why a bigger bound is the wrong first move The instinct on seeing recipients disconnected is to raise the bound. Consider what that buys: | Change | What improves | What it costs | |---|---|---| | Raise the per-recipient bound | A brief stall no longer disconnects anyone | Sustained lag now holds far more memory on the shared tier | | Send larger payloads | Fewer messages | Each lagging recipient holds more, multiplied by how many lag | | More recipients on one name | Broader fan-out | The same lag now multiplies across every one of them | A larger bound does not make the recipient faster. It converts a bounded, local failure — one disconnected recipient — into an unbounded, shared one: memory pressure on a tier that other callers depend on, and on stores set to remove entries under pressure it starts costing *other people's data*. ## What to do instead 1. **Shrink the payload.** Broadcast a small notice that something changed and let the recipient read the detail from the source of record. Fan-out cost scales with payload size times attached recipients, so this is the biggest lever you have. 2. **Lower the send rate.** Coalesce updates into one message per interval instead of one per change. Recipients that only render a screen rarely need every intermediate value. 3. **Move the slow consumer off the broadcast.** If a particular recipient genuinely needs every message and cannot keep up, it does not need a faster broadcast — it needs a surface that keeps messages for it, which is a durable log behind a broker. 4. **Watch for it.** Disconnections caused by a lagging recipient, and the memory held for connections rather than for data, are the two signals that tell you this is happening before it becomes an incident. The senior point being probed is ownership of the failure: a slow reader is not only the slow reader's problem, because the memory it wastes belongs to the shared tier.

  • Why does raising the per-recipient bound usually make things worse rather than better?
    It does not speed the recipient up; it only lets it hold more of the tier's memory for longer. A bounded local failure — one disconnected recipient — becomes shared memory pressure, and on a store set to remove entries when it reaches its ceiling that pressure starts costing other callers their data.
  • What is the cheapest way to cut the memory a fan-out holds for lagging recipients?
    Shrink the payload. The cost scales with message size times the number of lagging recipients, so broadcasting a short notice that something changed, and letting the recipient read the detail from the source of record, cuts it by far more than any tuning of the bound. Coalescing several changes into one message per interval is the next lever.
  • The recipient reconnects immediately after being disconnected. What has it lost?
    Everything sent between the disconnect and the reattach, with no way to discover how much. If it is still slow it will be disconnected again, which shows up as a recipient that flaps rather than one that fails outright — the signature worth alerting on.

saying these in an interview costs you the question

  • Saying the backlog waits on the recipient's side, so the server is unaffected
  • Treating a lagging recipient as that team's problem alone
  • Raising the per-recipient bound as the first fix for repeated disconnections
  • Assuming the disconnected recipient can collect what it missed afterwards
  • Believing every store in this class bounds the backlog the same way