skip to content

Messaging Foundations

The mechanics every messaging system shares: producer and consumer roles, broker guarantees, ordering, duplicate delivery and schema change. Interviewers start here before any pattern question.

part ofEvent-driven architecture & messagingoverview, primer and where to startread it →
on this pageshow

questions

page 2 of 2

A message repeatedly fails processing - say, due to a malformed payload a consumer can never successfully parse. Without any special handling, what happens to that message in an at-least-once queue system, and how does dead-letter queue (DLQ) routing fix it?

level: seniorimportance: must knowfreq 80%

basics

~20 s

Without a fix, the broker keeps redelivering the bad message forever, and it can jam up processing for everyone behind it. A DLQ routes a message elsewhere automatically after it fails too many times, so the queue keeps moving and someone can look at the bad message later.

open as a page

How does a broker's message retention policy differ between 'delete after acknowledgment' queues and 'keep for a fixed time/size regardless of consumption' log retention, and what does each let you do that the other doesn't?

level: middleimportance: should knowfreq 60%

basics

~20 s

Some brokers throw a message away as soon as it's been successfully handled once. Others keep every message around for a set amount of time no matter who's read it, so you can go back and re-read old messages later, at the cost of using more storage.

open as a page

How does dead-letter queue routing actually get implemented differently in a broker like Amazon SQS or RabbitMQ versus a log-based system like Apache Kafka, where there is no built-in 'move failed message elsewhere' primitive?

level: middleimportance: should knowfreq 55%

basics

~20 s

Queue systems like SQS or RabbitMQ can automatically move a failed message to a separate DLQ for you. Kafka has no such built-in feature — the consumer's own code has to catch the failure and manually publish the message to a separate 'DLQ topic' itself.

open as a page

A consumer applies each 'account balance updated' event using SQL like UPDATE accounts SET balance = balance + :delta WHERE id = :id. Why is this operation NOT naturally idempotent under redelivery, and what change would make it safe to replay?

level: middleimportance: should knowfreq 65%

basics

~20 s

Adding a number to a running total isn't safe to repeat; do it twice and the total is wrong twice. Idempotent updates instead set a value to its final state, like 'balance is now X', or use a unique key so replaying the same update lands on the exact same result every time.

open as a page

If three different services (billing, shipping, and analytics) all need to react to the same 'OrderPlaced' notification, how does supporting that differ structurally between a destructive-read queue like RabbitMQ and an append-only log like Kafka?

level: middleimportance: should knowfreq 60%

basics

~20 s

With a queue, you need three separate queues (or a fan-out router) so each service gets its own copy, because one message can only be taken once. With a log, all three can just read the same shared stream independently, each keeping its own place.

open as a page

In a Kafka-style consumer group, what is a 'rebalance,' and what effect can it have on message ordering and processing guarantees while it's happening?

level: middleimportance: should knowfreq 55%

basics

~20 s

A rebalance is when the group reshuffles which consumer reads which partition (e.g., because someone joined or left). During it, processing pauses briefly, and a partition can move to a different consumer that has no memory of what was being processed, so you can see duplicate or stalled processing, though the log order itself isn't changed.

open as a page

In the competing consumers pattern, several consumer instances pull work from the same shared queue to scale out processing. What delivery-guarantee problems does this introduce that a single consumer wouldn't have, and how are they typically handled?

level: middleimportance: should knowfreq 70%

basics

~20 s

Multiple workers grabbing jobs from one shared line means a job's owner might crash mid-task, so the system needs a way to notice and give that job to someone else — otherwise it's silently lost or, if handled sloppily, done twice.

open as a page

In a broker with an explicit routing layer between publishers and queues (an exchange, in AMQP terms), how does the broker decide which queue(s) a published message ends up in, and what happens if the routing rules match nothing?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The publisher doesn't send directly to a queue - it sends to a routing component that reads a label on the message and copies it into whichever queues match that label. If nothing matches, the message is usually just dropped unless the broker is told to save unmatched messages somewhere.

open as a page

When would routing a failed message to a dead-letter queue be the wrong choice, and what should a team do instead in those cases?

level: seniorimportance: should knowfreq 50%

basics

~20 s

If a message failure means something urgent needs to happen right away (like a fraud alert) or reprocessing it later could cause harm (like applying a stale price), quietly parking it in a DLQ for someone to check later isn't good enough — you need a faster, more active response instead.

open as a page

How does Kafka's transactional producer combine the idempotent producer with a transaction coordinator to give exactly-once semantics for a read-process-write pipeline, and what exactly does that atomicity cover?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Kafka can bundle writing output records and committing the input offset into one all-or-nothing transaction, using a producer ID and sequence numbers to avoid duplicate writes on retry, so a crash mid-write either commits everything or nothing — but only within Kafka.

open as a page

How do stateful stream operators (like running aggregations or joins) maintain state across a potentially unbounded stream, and how do they recover that state after a crash or task reassignment?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Stateful operators keep a local, on-disk 'memory' (like a running total) next to each processing task, and also write every change to a durable, replayable backup log. If the task dies or moves machines, it rebuilds its memory by replaying that backup log.

open as a page

In windowed stream processing, what is a watermark, and how does it let a system decide when a time window is 'done' despite events arriving out of order?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A watermark is the stream processor's best guess of 'we won't see any more events older than this timestamp.' Once the watermark passes a window's end time, the system treats that window as closed and emits its result, accepting it might occasionally be wrong if a late straggler shows up after.

open as a page

A team adds a processed-message dedup table to every consumer in their event-driven system by default, even for consumers whose downstream operation is already a naturally idempotent upsert keyed by the entity's primary key. What is the cost of that blanket policy, and when should a team skip the dedup table entirely?

level: seniorimportance: should knowfreq 55%

basics

~20 s

A separate 'have I seen this before' table costs extra storage, a database write on every message, and upkeep to stop it growing forever. If the actual operation is already safe to repeat on its own, like setting a status to a fixed value, that extra table isn't needed at all.

open as a page

You're designing a system to send password-reset emails: an API call enqueues a 'send reset email' task, and a worker pool sends it exactly once soon after. Would you reach for a destructive-read queue like SQS or an event log like Kafka, and why?

level: seniorimportance: should knowfreq 55%

basics

~20 s

A queue - you want the task done once by one worker and then forgotten, with automatic retry if it fails. A log is overkill here because nobody else needs to replay or re-read 'send this email' later; you'd actually want to avoid accidentally resending it.

open as a page

A partitioned topic is keyed by user ID to preserve per-user event ordering, but one 'celebrity' user generates 100x the event volume of a typical user. What problem does this cause, and what are reasonable ways to mitigate it without abandoning per-user ordering entirely?

level: seniorimportance: should knowfreq 45%

basics

~20 s

That one heavy user's events all pile onto a single partition no matter how many partitions the topic has, so that one partition gets overloaded and lags behind while the rest stay fine. Fixes usually involve splitting that user's events across a few partitions and reordering them again on the consumer side, or treating that user as a special case.

open as a page

Some messaging systems have the broker push messages to consumers, while others have consumers pull (poll) messages from the broker. How do these two models differ in how they handle a consumer that's overwhelmed, and why do most high-throughput systems favor pull?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Push means the broker decides when to send you the next message, which risks flooding a slow consumer. Pull means the consumer asks for more only when it's ready, so it naturally never asks for more than it can handle.

open as a page

After a message has been successfully consumed, how does its fate typically differ between a traditional point-to-point queue and a pub-sub topic with multiple independent subscribers, and why does that difference matter for durability?

level: seniorimportance: should knowfreq 65%

basics

~20 s

In a queue, once one consumer finishes a message, it's normally deleted - gone for good. In pub-sub, each subscriber has its own independent copy and cursor, so one subscriber finishing doesn't remove the message for the others, and depending on the system it may still be replayable later.

open as a page

Beyond schema registry compatibility checks, what does 'contract testing' add when verifying that producers and consumers of asynchronous messages actually work together, and how does it differ from just relying on the registry's compatibility mode?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Schema compatibility checks only verify the message's structure is technically valid. Contract testing goes further and checks that consumers get the actual data and values they depend on, catching cases where a message is structurally fine but breaks business logic downstream.

open as a page

A platform team is debating whether to mandate FULL compatibility mode for every schema subject across the organization. What are the concrete costs of that policy, and in what situations would a looser mode, or no registry at all, be the better call?

level: seniorimportance: should knowfreq 45%

basics

~20 s

FULL mode is the safest but the most restrictive: it forces every schema change to be additive with defaults, which slows down teams that need to rename fields, tighten types, or make big changes. For internal, tightly-coordinated systems or truly transient data, a looser policy or skipping the registry entirely can be the more pragmatic choice.

open as a page

What operational and architectural trade-offs do you take on when you choose a brokerless messaging library (like ZeroMQ) that connects producers and consumers directly, instead of routing all messages through a central broker process?

level: principalimportance: should knowfreq 40%

basics

~30 s

A broker is a separate service that sits in the middle and stores messages for you. A brokerless library like ZeroMQ instead makes the sending and receiving programs talk straight to each other over the network, which is faster and has one less thing to run, but now each program has to handle things the broker used to handle, like remembering where everyone is and what to do if a message can't be delivered yet.

open as a page

As a principal engineer designing a payment pipeline that spans a Kafka topic, an external payment gateway, and a legacy non-transactional data warehouse sink, how would you decide where to enforce which delivery guarantee, and what's the risk of applying the same guarantee uniformly everywhere?

level: principalimportance: should knowfreq 40%

basics

~20 s

Different parts of the pipeline need different guarantees based on how costly a duplicate or a loss is there — you can't just pick one setting for the whole system. Use real exactly-once only where it's cheap, like Kafka-to-Kafka, and layer at-least-once plus deduplication everywhere the guarantee has to reach outside Kafka, matching the mechanism to each sink's actual risk and cost.

open as a page

You're designing a system where an OrderPlaced event needs to independently trigger inventory update, email notification, and analytics logging at different paces, while a separate CapturePayment step must be processed exactly once by exactly one worker with strict, ordered handling per account. Would you use a pub-sub topic or a point-to-point queue for each, and how would you combine them?

level: principalimportance: should knowfreq 55%

basics

~20 s

For OrderPlaced, use pub-sub - every interested service (inventory, email, analytics) needs its own full copy at its own speed. For CapturePayment, use a point-to-point queue so only one worker handles each payment exactly once, with per-account ordering.

open as a page

How would you design a consumer's error-handling logic to classify a failure as transient (worth retrying) versus permanent (should go straight to the dead-letter queue without wasting the full retry budget), and what happens if that classification is wrong in either direction?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

You look at what kind of error occurred — a 'try again later' error (like a timeout) should get retried, but a 'this will never work' error (like bad data) should skip straight to the DLQ instead of burning through retries pointlessly.

open as a page

At a system-design level, how do stream-processing frameworks achieve 'exactly-once' processing guarantees for stateful aggregations, tying together offset commits, state updates, and output writes - and what breaks that guarantee?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

The system bundles 'I read this event,' 'I updated my running total,' and 'I wrote the result' into one all-or-nothing operation using transactions, so a crash can never leave things half-done - either all three happened or none did.

open as a page

A high-throughput consumer group processes millions of events per hour and checks a shared relational 'processed_messages' table before applying each one. What architectural problems does this single shared table cause at that scale, and what alternative designs would you consider?

level: principalimportance: nice to knowfreq 35%

basics

~20 s

One shared 'have we seen this before' table gets hammered by every single message across a huge system, becomes a bottleneck, and keeps growing forever. At large scale, teams split it up, for example one store per partition, or use a fast key-value store with automatic expiry instead of one giant table everyone reads and writes.

open as a page

A single malformed message lands in position 500 of a Kafka partition that a consumer group is processing sequentially, and the consumer's deserializer throws every time it hits that offset. How does the resulting outage differ from what would happen if the same bad message had instead been delivered via an SQS queue, and why?

level: principalimportance: nice to knowfreq 35%

basics

~20 s

In Kafka, that one bad message can jam up the whole line behind it, because the consumer must go through messages in order - everything after gets stuck waiting. In SQS, a bad message just gets set aside (retried, then dead-lettered) while all the other messages keep moving normally.

open as a page

A financial ledger system needs every transaction across all accounts to be applied in one single, globally agreed-upon order (not just per-account order), but the team is using a partitioned event log for throughput. What architectural approaches let you get a global order guarantee out of a system whose native primitive only guarantees per-partition order?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

You either force everything through one lane (slow but simple), or you let events flow through many lanes and later merge them back into one order using timestamps or a sequence number handed out by a single central 'ticket booth,' accepting some extra delay or complexity to reconstruct that single true order.

open as a page

You're designing back-pressure policy for a shared event-streaming platform where dozens of teams' producers and consumers coexist on the same broker cluster. What strategies would you weigh for handling a consumer that lags badly, and what cascading risks does each carry across tenants?

level: principalimportance: nice to knowfreq 35%

basics

~20 s

When lots of teams share one messaging system, one team's slow consumer can hurt everyone else if it's allowed to fill up shared storage or overload shared brokers. The fix is isolating teams from each other (quotas, per-topic limits) plus giving each team clear tools (autoscaling, dead-letter queues, shedding) to fix their own lag before it spreads.

open as a page

When choosing between Avro, Protobuf, and JSON Schema for a message registry, what actually differs about how each handles schema evolution, and what should drive the choice?

level: principalimportance: nice to knowfreq 35%

basics

~30 s

All three can express similar rules for adding/removing fields safely, but they differ in how evolution is tracked: Avro relies on a full schema document and reader/writer resolution, Protobuf uses stable numbered field tags baked into the format itself, and JSON Schema is more flexible but has weaker native tooling for enforcing evolution rules. The right pick depends on your ecosystem, tooling, and how much schema strictness you actually want.

open as a page

showing 31–59 of 59