skip to content

Which requirements tell you a fan-out workload has outgrown the volatile tier's broadcast and belongs on a durable log with a broker?

level: seniorimportance: must knowfreq 56%

answer

  1. test against what the surface promises
  2. kept, acknowledged, positioned, replayed
  3. one trigger is enough
  4. retention built on entries is a worse broker
  5. no trigger means broadcast is cheaper

basics

~20 s

Four: something must be kept for a recipient absent at send time; someone must know a recipient handled it; a recipient needs a resumable position; or history must be replayable. Any one means a broker.

solid answer

~50 s

The broadcast surface makes exactly one promise: a copy to every recipient attached at the instant of the send, kept for nobody. So the handover test is a list of things it does not do. If anything must be held for a recipient that is absent right then, if the sender or an operator must know the message was handled, if a recipient needs a position it can resume from after a restart, or if anyone must reprocess history — onboard a new recipient, rerun after a bug, satisfy an audit — the workload has left this tier. The failure mode to name is the middle path: keeping recent messages under ordinary entries so recipients can catch up, which rebuilds retention, positions and acknowledgement by hand on storage that drops them under memory pressure or a restart. If none of the four applies, the broadcast is cheaper, lower-latency and one fewer failure domain.

go deeper

for a junior

Learn the one promise the broadcast makes — a copy to whoever is attached right now, kept for nobody — because every part of this decision is measured against it.

for a middle

Be able to list the four things that force a broker: keeping for the absent, acknowledgement, a resumable position, and replay. Say which one a given scenario hits.

for a senior

Argue both directions. Refuse the hand-rolled retention on ordinary entries and explain how it fails under memory pressure, but also defend staying on the broadcast when no trigger applies.

for a principal

Own the consequence: adding a broker adds a failure domain and operational surface, so make the handover on a stated requirement and write down what the product promises when a message is missed.

## The promise you are testing against The broadcast surface on a shared volatile tier promises one thing: at the instant of a send, a copy goes to every recipient attached to that name, and then the message is gone. There is no entry, no retention, no per-recipient state and no history. Everything below is a test of whether your requirements stay inside that promise. ## The four triggers Any one of these on its own means the workload belongs on a durable log behind a broker. 1. **Something must be kept for a recipient that is absent at send time.** A recipient restarting, deploying, or briefly off the network must still get the message. The broadcast holds nothing for the absent, by construction. 2. **Someone must know a recipient actually handled it.** Acknowledgement — the sender, an operator or a retry loop needing evidence that the message was processed, not merely transmitted. The broadcast reports at most how many connections were handed a copy, and never what any of them did with it. 3. **A recipient needs a position it can resume from.** Anything that says "continue from where I stopped" after a restart requires the surface to track per-recipient progress, which this one does not. 4. **History must be replayable.** Onboarding a new recipient that needs the past, rerunning after a bug, or reconstructing a sequence of events for an audit. A broadcast has no past. A fifth, weaker trigger is worth naming separately: **the message must survive the tier's own restart or failover.** The broadcast surface is delivery in flight, so a failover ends every attachment; recipients reconnect and continue from that moment, with the gap lost. If that gap has real cost, you are back at trigger one. ## The test in table form | Requirement in the ticket | Stays on the broadcast | Moves to a broker | |---|---|---| | "Refresh any screen that is open" | Yes — a missed refresh self-heals on the next one | | | "Every order must reach the fulfilment service" | | Yes — absence at send time cannot lose it | | "Tell me which consumers processed it" | | Yes — that is acknowledgement | | "A new consumer must see the last day's events" | | Yes — that is replay | | "Nudge workers that there is something to do" | Yes — the durable job list is the truth | | Notice the pattern in the left column: the message is worthless a moment after it is sent, and the truth lives somewhere durable. That is the shape that belongs on this tier. ## The wrong middle path The move to argue against is building retention out of ordinary entries: keeping the last few minutes of messages under keys, having recipients read them on attach, and stamping them so a recipient can work out where it left off. You have then implemented retention, reader positions and, before long, acknowledgement, on storage that: - can remove those entries under memory pressure, at exactly the moment the system is busiest; - keeps nothing across a restart unless you also chose a persistence posture for it; - has no notion of a consumer, so every recipient's bookkeeping is yours to get right; - costs memory proportional to retention span times message rate, competing with the data in the same tier. That is a broker with worse durability, written by you, maintained by you. Where the requirement is genuine, take the broker. ## What justifies staying It is not a one-way test, and reaching for a broker by reflex has its own costs — another system to operate, another failure domain in the request path, and latency that a broadcast simply does not have. Stay on the broadcast when all of the following hold: - a missed message is invisible and harmless, because the recipient re-reads the truth from a durable source anyway; - nobody needs to know who received it; - there is no history worth having; - the fan-out is live and cheap — a screen refresh, a cache-warming nudge, an operational signal. ## How to answer it out loud Name the four triggers, say which one the interviewer's scenario hits, and state the consequence in the product's terms: *"if a recipient that was deploying at the time must still get this, the tier cannot do it, and re-creating retention here gives you a broker that loses data under memory pressure."* That answer shows the boundary is a decision you make deliberately rather than a wall you hit in production.

  • A team keeps the last five minutes of messages under ordinary entries so recipients can catch up. What is wrong with that?
    It is retention, positions and eventually acknowledgement rebuilt by hand on storage that can remove those entries under memory pressure and keeps nothing across a restart unless a persistence posture was chosen. The memory cost is the retention span times the message rate, competing with the tier's data — a broker with worse durability that you now maintain.
  • None of the four triggers applies. What do you gain by staying on the broadcast?
    Lower latency, no extra system to run, and one fewer failure domain in the path. A broadcast costs a write to attached connections and nothing else, and when the truth already lives in a durable source the miss is self-healing. Reaching for a broker by reflex buys guarantees you are not using.
  • Does a failover of the tier count as a trigger on its own?
    Only through the first trigger. A failover ends every attachment, so recipients reconnect and resume from that moment with the gap lost. If losing that gap has real cost, something had to be kept for an absent recipient — which is trigger one — and the workload belongs on a broker.

saying these in an interview costs you the question

  • Judging the handover by message volume rather than by what must be kept
  • Rebuilding retention and positions out of ordinary entries on the tier
  • Claiming a broker is always the safer default regardless of requirements
  • Assuming a reconnecting recipient can catch up if messages are sent often enough
  • Treating acknowledgement as something the broadcast surface can be configured to add