What is a message broker, and why would two services send messages through one instead of calling each other's APIs directly?
answer
- middleman decouples producer/consumer
- temporal decoupling
- extra hop = extra infra
- at-least-once = idempotent consumers
- queue buildup failure mode
basics
~20 sA message broker is a middleman server that receives messages from senders and delivers them to receivers, so the two sides never talk directly. This lets one side keep working even if the other is slow, down, or busy.
solid answer
~40 sA message broker is an intermediary process that accepts messages from producers and routes them to one or more consumers via named channels (queues or topics), decoupling the two sides in time, space, and pace. Producers don't need to know who consumes a message, how many consumers exist, or whether they're currently available; the broker buffers messages until a consumer is ready. This buys temporal decoupling (a consumer can be down without blocking the producer), load leveling (bursts get smoothed into a queue), and easier fan-out. The cost is an extra hop, an extra piece of infrastructure to run and monitor, and weaker consistency guarantees than a direct synchronous call.
go deeper
Can explain that a broker sits between sender and receiver and buffers messages so they don't need to be online at the same time; doesn't need to know delivery guarantees yet.
Should articulate temporal decoupling as the core value, know that most brokers are at-least-once by default, and name one real broker product.
Should discuss operational trade-offs (broker as shared dependency/SPOF), poison messages and dead-letter handling, and design idempotent consumers.
Should reason about when a broker is the wrong tool, how broker choice affects system-wide consistency and observability strategy, and how to avoid the broker becoming an organizational bottleneck.
## Where the broker sits A **message broker** sits as a separate network service between producers and consumers. Neither side holds a direct network connection to the other; both only ever talk to the broker. - **The producer** opens a connection to the broker and publishes a message to a named destination — a queue or a topic. - **The broker's job** is to accept that message, persist or buffer it (in memory, on disk, or both, depending on configuration), and then deliver it to whichever consumer(s) are entitled to receive it according to the destination's semantics. - **The consumer**, running as a wholly separate process possibly on different hardware, opens its own connection to the broker, subscribes to or polls the destination, and receives messages independently of when the producer sent them. Delivery can be **push-based** or **pull-based**, and most brokers support **acknowledgment** — the consumer tells the broker 'I successfully processed this,' and only then does the broker consider it delivered. ## The problem it solves The core problem a broker solves is **coupling** — specifically **time coupling**, and to a lesser extent **space** and **load** coupling. In a synchronous request/response call, service A must know service B's address, B must be running and responsive right now, and A blocks until B answers. If B is deploying, overloaded, or crashed, A's call fails or hangs, and that failure can cascade back to whoever called A. A broker breaks this: A publishes and moves on, and the message sits in the broker until B is ready to consume it. This is invaluable for workflows where the consumer is slow or temporarily unavailable, and for spreading a load spike over time rather than forcing the receiver to handle it instantaneously. ## What the decoupling costs This decoupling isn't free. 1. **First, an extra hop.** It introduces an extra network hop and a piece of infrastructure that has to be operated and monitored — if the broker goes down, everything using it stalls or loses messages, so it becomes a shared dependency and potential single point of failure unless deployed in a highly available cluster. 2. **Second, eventual consistency.** You trade strong consistency for eventual consistency: the producer typically gets no confirmation that the consumer processed the message successfully, only that the broker accepted it. 3. **Third, harder debugging.** Debugging gets harder — an asynchronous processing bug shows up minutes or hours later, disconnected from the code that produced it, so you need correlation IDs and good observability to follow a message's journey. 4. **Fourth, duplicates.** Most brokers offer at-least-once delivery by default, meaning consumers must be written to tolerate duplicate messages (idempotency). ## Failure modes - **Queue buildup.** In production, the classic broker-related failure is queue buildup: if consumers fall behind or crash, queue depth grows unbounded, memory/disk fills, and eventually the broker starts rejecting messages or crashes, taking down every producer/consumer pair depending on it. - **The 'poison message'.** Another common failure is a malformed message a consumer can never successfully process, redelivered on every crash/retry, looping forever unless a dead-letter mechanism removes it after N attempts. - **Silent message loss.** A third failure mode is silent loss when a broker is configured to hold messages only in memory and the process restarts — anything unacknowledged and unpersisted disappears. - **Duplicate delivery.** Finally, network partitions between broker and consumers can cause duplicate delivery: an acknowledgment gets lost in transit and the broker redelivers a message already handled, which is why idempotent consumers matter. ## A concrete example A concrete example: an e-commerce checkout service places an order and needs to charge a card, email a receipt, and update inventory. If checkout called all three synchronously, a slow email provider could delay or fail the entire checkout. Instead, checkout publishes an `OrderPlaced` message to a broker; payment, email, and inventory services each consume that message independently, at their own pace, and retry on their own if they fail, without checkout ever knowing or caring how long they take. Checkout's response to the customer returns as soon as the message is accepted by the broker, not after all three side effects complete.
- What happens to a producer's request if the broker itself is unreachable when it tries to publish?It depends on the client library, but most block or throw a connection error, since the producer must still successfully reach the broker to hand off the message. Well-designed producers wrap this in retries with backoff, and some queue messages locally as a last resort, but if the broker cluster is fully down, publishing genuinely fails and the caller needs a fallback.
- How does using a broker change your error-handling story compared to a direct HTTP call?With a direct call you get an immediate synchronous failure you can catch in the same request. With a broker, failures happen asynchronously in the consumer, often much later, so you need dead-letter queues, retry policies, and alerting on consumer-side failures rather than a caught exception in the caller's code path.
- Why is 'at-least-once' delivery the common default rather than 'exactly-once'?Exactly-once delivery across a network requires coordinating the send, the durable write, and the acknowledgment in a way that survives crashes without duplicating or losing anything, which is expensive to guarantee end-to-end. Brokers instead default to at-least-once and push deduplication onto the consumer via idempotency keys, which is simpler and cheaper to implement correctly.
A message broker is like a post office: you drop a letter in a mailbox (publish) and walk away; the post office holds it and delivers it whenever the recipient checks their mailbox, so you never need the recipient to be home when you write the letter.
saying these in an interview costs you the question
- says a broker guarantees exactly-once delivery by default
- doesn't mention decoupling as the core benefit
- thinks a broker removes the need for error handling entirely
- confuses a broker with a load balancer
- assumes producer and consumer must be online at the same time