skip to content

Messaging Foundations

The mechanics every messaging system shares: producer and consumer roles, broker guarantees, ordering, duplicate delivery and schema change. Interviewers start here before any pattern question.

part ofEvent-driven architecture & messagingoverview, primer and where to startread it →
on this pageshow

questions

page 1 of 2

What is a message broker, and why would two services send messages through one instead of calling each other's APIs directly?

level: juniorimportance: must knowfreq 85%

answer

  1. middleman decouples producer/consumer
  2. temporal decoupling
  3. extra hop = extra infra
  4. at-least-once = idempotent consumers
  5. queue buildup failure mode

basics

~20 s

A message broker is a middleman server that receives messages from senders and delivers them to receivers, so the two sides never talk directly. This lets one side keep working even if the other is slow, down, or busy.

solid answer

~40 s

A message broker is an intermediary process that accepts messages from producers and routes them to one or more consumers via named channels (queues or topics), decoupling the two sides in time, space, and pace. Producers don't need to know who consumes a message, how many consumers exist, or whether they're currently available; the broker buffers messages until a consumer is ready. This buys temporal decoupling (a consumer can be down without blocking the producer), load leveling (bursts get smoothed into a queue), and easier fan-out. The cost is an extra hop, an extra piece of infrastructure to run and monitor, and weaker consistency guarantees than a direct synchronous call.

go deeper

for a junior

Can explain that a broker sits between sender and receiver and buffers messages so they don't need to be online at the same time; doesn't need to know delivery guarantees yet.

for a middle

Should articulate temporal decoupling as the core value, know that most brokers are at-least-once by default, and name one real broker product.

for a senior

Should discuss operational trade-offs (broker as shared dependency/SPOF), poison messages and dead-letter handling, and design idempotent consumers.

for a principal

Should reason about when a broker is the wrong tool, how broker choice affects system-wide consistency and observability strategy, and how to avoid the broker becoming an organizational bottleneck.

## Where the broker sits A **message broker** sits as a separate network service between producers and consumers. Neither side holds a direct network connection to the other; both only ever talk to the broker. - **The producer** opens a connection to the broker and publishes a message to a named destination — a queue or a topic. - **The broker's job** is to accept that message, persist or buffer it (in memory, on disk, or both, depending on configuration), and then deliver it to whichever consumer(s) are entitled to receive it according to the destination's semantics. - **The consumer**, running as a wholly separate process possibly on different hardware, opens its own connection to the broker, subscribes to or polls the destination, and receives messages independently of when the producer sent them. Delivery can be **push-based** or **pull-based**, and most brokers support **acknowledgment** — the consumer tells the broker 'I successfully processed this,' and only then does the broker consider it delivered. ## The problem it solves The core problem a broker solves is **coupling** — specifically **time coupling**, and to a lesser extent **space** and **load** coupling. In a synchronous request/response call, service A must know service B's address, B must be running and responsive right now, and A blocks until B answers. If B is deploying, overloaded, or crashed, A's call fails or hangs, and that failure can cascade back to whoever called A. A broker breaks this: A publishes and moves on, and the message sits in the broker until B is ready to consume it. This is invaluable for workflows where the consumer is slow or temporarily unavailable, and for spreading a load spike over time rather than forcing the receiver to handle it instantaneously. ## What the decoupling costs This decoupling isn't free. 1. **First, an extra hop.** It introduces an extra network hop and a piece of infrastructure that has to be operated and monitored — if the broker goes down, everything using it stalls or loses messages, so it becomes a shared dependency and potential single point of failure unless deployed in a highly available cluster. 2. **Second, eventual consistency.** You trade strong consistency for eventual consistency: the producer typically gets no confirmation that the consumer processed the message successfully, only that the broker accepted it. 3. **Third, harder debugging.** Debugging gets harder — an asynchronous processing bug shows up minutes or hours later, disconnected from the code that produced it, so you need correlation IDs and good observability to follow a message's journey. 4. **Fourth, duplicates.** Most brokers offer at-least-once delivery by default, meaning consumers must be written to tolerate duplicate messages (idempotency). ## Failure modes - **Queue buildup.** In production, the classic broker-related failure is queue buildup: if consumers fall behind or crash, queue depth grows unbounded, memory/disk fills, and eventually the broker starts rejecting messages or crashes, taking down every producer/consumer pair depending on it. - **The 'poison message'.** Another common failure is a malformed message a consumer can never successfully process, redelivered on every crash/retry, looping forever unless a dead-letter mechanism removes it after N attempts. - **Silent message loss.** A third failure mode is silent loss when a broker is configured to hold messages only in memory and the process restarts — anything unacknowledged and unpersisted disappears. - **Duplicate delivery.** Finally, network partitions between broker and consumers can cause duplicate delivery: an acknowledgment gets lost in transit and the broker redelivers a message already handled, which is why idempotent consumers matter. ## A concrete example A concrete example: an e-commerce checkout service places an order and needs to charge a card, email a receipt, and update inventory. If checkout called all three synchronously, a slow email provider could delay or fail the entire checkout. Instead, checkout publishes an `OrderPlaced` message to a broker; payment, email, and inventory services each consume that message independently, at their own pace, and retry on their own if they fail, without checkout ever knowing or caring how long they take. Checkout's response to the customer returns as soon as the message is accepted by the broker, not after all three side effects complete.

  • What happens to a producer's request if the broker itself is unreachable when it tries to publish?
    It depends on the client library, but most block or throw a connection error, since the producer must still successfully reach the broker to hand off the message. Well-designed producers wrap this in retries with backoff, and some queue messages locally as a last resort, but if the broker cluster is fully down, publishing genuinely fails and the caller needs a fallback.
  • How does using a broker change your error-handling story compared to a direct HTTP call?
    With a direct call you get an immediate synchronous failure you can catch in the same request. With a broker, failures happen asynchronously in the consumer, often much later, so you need dead-letter queues, retry policies, and alerting on consumer-side failures rather than a caught exception in the caller's code path.
  • Why is 'at-least-once' delivery the common default rather than 'exactly-once'?
    Exactly-once delivery across a network requires coordinating the send, the durable write, and the acknowledgment in a way that survives crashes without duplicating or losing anything, which is expensive to guarantee end-to-end. Brokers instead default to at-least-once and push deduplication onto the consumer via idempotency keys, which is simpler and cheaper to implement correctly.

A message broker is like a post office: you drop a letter in a mailbox (publish) and walk away; the post office holds it and delivers it whenever the recipient checks their mailbox, so you never need the recipient to be home when you write the letter.

saying these in an interview costs you the question

  • says a broker guarantees exactly-once delivery by default
  • doesn't mention decoupling as the core benefit
  • thinks a broker removes the need for error handling entirely
  • confuses a broker with a load balancer
  • assumes producer and consumer must be online at the same time

context

open as a page

In a message queue system, what is a dead-letter queue (DLQ), and why does a consumer route a message there instead of retrying it forever?

level: juniorimportance: must knowfreq 65%

basics

~20 s

A DLQ is a separate queue where messages that fail processing repeatedly get moved to, instead of blocking the main queue or being retried forever, so the app keeps working while someone looks at the bad message later.

open as a page

In event-driven messaging systems, what do at-most-once, at-least-once, and exactly-once delivery mean, and which one do most production systems actually rely on by default?

level: juniorimportance: must knowfreq 85%

basics

~20 s

At-most-once can lose a message but never repeats it. At-least-once never loses a message but can repeat it. Exactly-once means every message is processed once, no loss and no duplicates. Most real systems use at-least-once plus deduplication.

open as a page

In an event streaming platform like Kafka, what is a consumer offset, and how does it allow a consumer to replay events it has already processed?

level: juniorimportance: must knowfreq 65%

basics

~10 s

An offset is a bookmark marking how far a consumer has read in a partition. Since the log isn't deleted after reading, moving the bookmark backward lets the consumer re-read and reprocess old events.

open as a page

A message queue delivers the same order-created message to a consumer twice because the consumer crashed after processing it but before acknowledging it. What must the consumer do so processing the message twice doesn't cause a duplicate charge or duplicate shipment?

level: juniorimportance: must knowfreq 75%

basics

~20 s

Idempotent means doing something twice has the same effect as doing it once. The consumer remembers which messages it already handled (by ID) and skips reprocessing the effect, like charging money again, even if the message arrives more than once.

open as a page

When a consumer successfully processes a message from a traditional message queue like RabbitMQ or SQS, what typically happens to that message, and how does that differ from a system like Kafka?

level: juniorimportance: must knowfreq 70%

basics

~20 s

In a queue, once a message is picked up and confirmed done, it's deleted - gone for good. In an event log like Kafka, the message stays stored so other readers can still see it later, or the same reader can go back and read it again.

open as a page

In a partitioned message log such as an Apache Kafka topic, messages within a single partition are delivered to consumers in the exact order they were produced. Why is there no such ordering guarantee across two different partitions of the same topic?

level: juniorimportance: must knowfreq 65%

basics

~10 s

Each partition is its own ordered line of messages, like a single queue. Different partitions are separate queues running independently, so messages from different partitions can arrive in any order relative to each other.

open as a page

When a producer sends a message to a broker instead of calling a consumer directly, what does that decoupling actually buy you, and what do you give up?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A producer drops a message off at a broker (like a queue) without knowing who reads it; a consumer picks messages up without knowing who sent them. Each side can change, restart, or scale on its own without breaking the other.

open as a page

What is the fundamental difference between publish-subscribe (pub-sub) messaging and point-to-point queuing, in terms of how many consumers receive a given message?

level: juniorimportance: must knowfreq 85%

basics

~20 s

In pub-sub, every subscriber gets its own copy of each message (fan-out, one-to-many). In a queue, each message goes to exactly one consumer picked from the pool (point-to-point, one-to-one), even if many workers are listening.

open as a page

What is a schema registry in an event-driven messaging system, and why do producers and consumers use one instead of just embedding the full schema in every message?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A schema registry is a shared service that stores the agreed-upon 'shape' of each message type. Producers register a schema and get a short ID; consumers use that ID to look up the schema and read the message correctly, instead of sending the whole schema every time.

open as a page

In a message broker, what's the practical difference between publishing to a queue versus publishing to a topic, in terms of how many consumers receive each message?

level: middleimportance: must knowfreq 80%

basics

~20 s

A queue delivers each message to exactly one consumer, like a shared to-do list where each task is picked up once. A topic broadcasts each message to every subscriber, like a radio station everyone tuned in hears.

open as a page

Before a message ever reaches a dead-letter queue, how should a retry policy typically combine a backoff schedule with a max-delivery threshold, and why not just retry immediately and forever?

level: middleimportance: must knowfreq 75%

basics

~20 s

You wait a bit longer between each retry (backoff) instead of retrying instantly, and you cap the total number of tries (threshold) so a broken message eventually gets pulled aside into the DLQ instead of hammering the system forever.

open as a page

When a Kafka consumer processes a batch of 10 records and commits only the offset of the last record after the whole batch succeeds, what happens on a crash after record 6 finishes, and why is committing after every single record not simply 'safer'?

level: middleimportance: must knowfreq 75%

basics

~20 s

Committing the offset once per batch is faster but replays the whole batch if a crash happens partway through, since Kafka only tracks one checkpoint per partition, not per message. Committing after every single message shrinks that replay window but is much slower.

open as a page

What does log compaction do to a Kafka-style topic, and how does it differ from simple time/size-based retention?

level: middleimportance: must knowfreq 60%

basics

~20 s

Compaction keeps only the newest record for each key and throws away older ones, instead of deleting everything older than some age. It's like keeping only the latest version of each row in a table, not a full history.

open as a page

In stream processing, what's the difference between tumbling, sliding, and session windows, and when would you reach for each?

level: middleimportance: must knowfreq 75%

basics

~20 s

A tumbling window is a fixed time slice that never overlaps the next one, like every 5 minutes. A sliding window overlaps and updates continuously, like a 'last 5 minutes' view refreshed often. A session window has no fixed size - it groups events until there's a gap of inactivity.

open as a page

A payment consumer stores every message ID it has successfully processed in a 'processed_messages' database table, and checks that table before applying a payment. What race condition can still cause a duplicate payment if two copies of the same message arrive at nearly the same time, and how would you close it?

level: middleimportance: must knowfreq 80%

basics

~20 s

A table that lists 'already handled' message IDs can still let two duplicates through if both are checked at the same instant before either writes its ID down, like two people checking an empty sign-up sheet at once and both writing their name first. Fix it by making the 'write' step atomic, for example a database unique constraint.

open as a page

Concretely, how does a Kafka-style consumer track 'what have I already processed' using offsets, and how is that different from how an SQS or RabbitMQ consumer acknowledges a message?

level: middleimportance: must knowfreq 75%

basics

~20 s

A log consumer remembers a number (an offset) marking its place in an ordered list of messages, and moves that number forward as it reads. A queue consumer instead tells the broker 'I'm done with this one specific message,' and the broker deletes just that message.

open as a page

When producing events to a partitioned topic, what determines which partition a given message is written to, and how should you pick a partition key so that all events belonging to the same business entity (say, a shopping cart or an IoT device) stay in order relative to each other?

level: middleimportance: must knowfreq 75%

basics

~20 s

The system picks a partition based on a formula applied to a 'key' you attach to each message. Use the same identifying value (like the entity's ID) as the key every time, so all its events always land in the same partition and stay in order.

open as a page

In a messaging system where multiple consumer instances share the label 'consumer group' when reading a topic, how does the broker split the work between them so each message is processed by only one instance in that group?

level: middleimportance: must knowfreq 85%

basics

~20 s

The topic's data is split into chunks (partitions), and the broker hands each chunk to exactly one instance in the group. So the group as a whole reads everything once, but each message is only handled by one member.

open as a page

A consumer pulls a message off a queue and crashes before sending an acknowledgement back to the broker. What happens to that message, and what delivery guarantee does this behavior typically produce?

level: middleimportance: must knowfreq 75%

basics

~20 s

The broker assumes the message wasn't handled and gives it to another consumer after a timeout, so the message isn't lost. But this means the same message might get processed twice - that's called at-least-once delivery.

open as a page

In a point-to-point queue with multiple consumer instances attached, how does the broker typically decide which consumer receives which message, and what ordering guarantees (if any) survive this distribution?

level: middleimportance: must knowfreq 78%

basics

~20 s

The broker hands each message to whichever available consumer is free (round-robin or pull-based), so work spreads across consumers. Once messages are spread across multiple consumers, you generally lose the guarantee that they're processed in the exact order they were sent.

open as a page

A topic's schema subject in a message schema registry is set to BACKWARD compatibility mode. What does that mode actually guarantee, and how does it differ from FORWARD and FULL compatibility modes?

level: middleimportance: must knowfreq 80%

basics

~20 s

BACKWARD compatibility means new schema versions can still be read using the code built for the old schema. FORWARD is the opposite: old code can read data written with the new schema. FULL means both directions work at once.

open as a page

In a Confluent-style Avro schema registry integration, walk through how the schema ID gets embedded in a message written to a Kafka topic, and what a consumer does with it on read.

level: middleimportance: must knowfreq 65%

basics

~20 s

The producer's serializer sticks a magic byte and a 4-byte number (the schema's ID) at the very start of the message, followed by the actual encoded data. The consumer reads those first 5 bytes to know which schema to fetch, then uses that schema to decode the rest.

open as a page

When a broker partitions a topic across multiple nodes/logs for parallelism, what ordering guarantee do you keep and what do you lose, and how does the partition key decide this?

level: seniorimportance: must knowfreq 75%

basics

~20 s

Splitting a topic into partitions lets multiple machines share the load, but messages are only guaranteed to arrive in order within a single partition, not across the whole topic. Which partition a message lands in is usually decided by a key you choose, like a customer ID.

open as a page

Once messages start landing in a dead-letter queue, what operational practices — alerting and a 'parking lot' replay strategy — turn the DLQ from a silent graveyard into something actionable, and what should a replay process actually check before resending a message?

level: seniorimportance: must knowfreq 70%

basics

~20 s

You need alarms that page someone when the DLQ starts filling up, and a safe process to review each stuck message, fix whatever caused it (or confirm the world hasn't changed), and only then resend it back to be processed — never just blindly replaying everything.

open as a page

Why do practitioners often say Kafka's 'exactly-once semantics' are really 'effectively-once' once you look at the whole pipeline end-to-end, from source system through to the final side effect?

level: seniorimportance: must knowfreq 70%

basics

~20 s

Because 'exactly-once' guarantees usually only cover the messaging system itself. The moment a message triggers an outside effect — an email, an API call, a non-transactional database write — that effect can still happen more than once, so the real guarantee is 'looks like exactly once if downstream systems are built to tolerate replay.'

open as a page

A producer publishes 'order-placed' events and retries publishing on timeout, so the broker may assign a different message ID to what is logically the same business event. Why is deduplicating on the broker-assigned message ID insufficient here, and what should the dedup key be instead?

level: seniorimportance: must knowfreq 70%

basics

~20 s

If the sender retries and the queue gives the retry a brand-new message ID, checking 'have I seen this exact message ID' won't catch it as a duplicate. Instead, dedupe on something the business itself considers the same thing every time, like the order's own ID, so retries collapse together no matter what ID the broker assigns.

open as a page

A consumer service has a bug that silently corrupts 3 days' worth of derived data before anyone notices. If the upstream system is a Kafka-style event log with 14-day retention, how would you recover, and why would the same recovery be much harder (or impossible) if the upstream were an SQS queue instead?

level: seniorimportance: must knowfreq 65%

basics

~20 s

With a log, you can rewind your reader to before the bug started and reprocess those 3 days from the original messages, since they're all still stored. With a queue, those messages are already deleted once the first, buggy pass consumed them, so there's nothing left to replay - you'd need another way to get that data back.

open as a page

Why does increasing a partitioned topic's partition count from, say, 6 to 24 improve consumer throughput, and what ordering guarantee do you give up (or put at risk) by doing so?

level: seniorimportance: must knowfreq 60%

basics

~20 s

More partitions means more independent lines that can be read and written in parallel, so more machines can work at once and throughput scales up. But you can never guarantee that the whole topic's events arrive in one single overall order - only within each line - and resizing can even split up an entity's history across old and new lines.

open as a page

When consumers process messages slower than producers publish them, consumer lag builds up. What actually breaks as that lag grows, and what back-pressure mechanisms can a system apply to cope?

level: seniorimportance: must knowfreq 80%

basics

~20 s

If readers can't keep up with writers, unprocessed work piles up. Eventually old messages get deleted before anyone reads them, or storage fills up, or downstream data goes stale. Back-pressure means slowing the producer down, scaling up consumers, or shedding load on purpose before things break.

open as a page

showing 1–30 of 59