Concretely, how does a Kafka-style consumer track 'what have I already processed' using offsets, and how is that different from how an SQS or RabbitMQ consumer acknowledges a message?
answer
- offset = your bookmark in the log
- ack = tell broker delete this message
- visibility timeout / in-flight limit
- consumer owns offset commit timing
- rebalance can cause duplicate processing
basics
~20 sA log consumer remembers a number (an offset) marking its place in an ordered list of messages, and moves that number forward as it reads. A queue consumer instead tells the broker 'I'm done with this one specific message,' and the broker deletes just that message.
solid answer
~30 sIn a log, each partition is an ordered, immutable sequence, and every message has a monotonically increasing offset. A consumer tracks the last offset it has committed; to catch up it just requests messages after that offset, and reading never touches the stored data. In a queue, there's no shared ordered index - the broker hands out individual messages and expects a per-message ack (or nack) back; on ack the message is deleted, on timeout or nack it's redelivered. Offset-tracking is consumer-owned and log-preserving, while acking is broker-owned and message-destroying.
go deeper
Can state that queues delete on ack and logs use a position marker (offset) instead.
Should explain visibility timeout / redelivery and offset commit as mechanisms, and know that 'ack' destroys while 'offset' doesn't.
Should discuss commit-before-vs-after ordering and its delivery-guarantee implications, plus at least one production failure mode (in-flight limit or rebalance).
Should reason about achieving exactly-once via idempotency keys or transactional offset commits, and weigh operational cost (lag monitoring, rebalance tuning) as a design factor.
## The log side: a bookmark you own In a Kafka-style log, a partition is literally an **append-only array**, and an offset is just an integer index into it. - A consumer group's committed position is stored centrally (Kafka keeps it in an internal topic, `__consumer_offsets`), and reading via `poll()` never removes or modifies anything. - Only explicitly committing the offset (auto-commit on an interval, or manual sync/async commit) advances the consumer's bookmark. - Crucially, multiple consumer groups reading the same partition each hold a completely independent offset, so one group's progress has zero effect on another's. ## The queue side: state the broker keeps for you In a queue, the broker instead maintains real **per-message state**: available, in-flight, or deleted. When a consumer receives a message, it becomes invisible (SQS's visibility timeout) or unacked (RabbitMQ), and the client must send back an explicit acknowledgment referencing that specific message before the broker purges it. If no ack arrives within the window, the broker assumes failure and makes the message available again for redelivery, possibly to an entirely different consumer instance. ## Why the split exists This split exists because the two models optimize for different guarantees. - **The queue model** optimizes for 'this piece of work must be done by exactly one worker, and once done, it's off the books' - minimal broker-side state, self-cleaning. - **The log model** optimizes for 'this fact must be available to many independent readers, at their own pace' - since there's no shared notion of 'done,' each reader must independently bookmark its own position. ## The trade-offs The trade-offs cut both ways. Offset commits are cheap and batchable - you can commit once per thousand messages or once per few seconds - and replay is as simple as resetting the offset backward. But the ordering of commit versus processing matters: 1. **Committing before finishing work** risks silent message loss if the consumer crashes mid-batch (at-most-once). 2. **Committing after finishing work** risks reprocessing the same messages on restart (at-least-once, requiring idempotent handling). True exactly-once needs either idempotent side effects keyed by offset/event-ID, or a transactional API that ties the offset commit and the output write into one atomic operation. Acks give simpler single-message semantics with automatic redelivery and dead-lettering built in, but that per-message state doesn't scale to letting ten independent services each read the same message without physically duplicating it into ten queues. ## Production failure modes Production failure modes differ accordingly. - **Logs** suffer from consumer-group lag creeping unnoticed until someone checks a lag metric, and from 'rebalance storms' - when a consumer joins or leaves a group, partitions get reassigned, and the new owner resumes from the last committed offset, which can be behind where the old owner actually got to, causing a brief burst of duplicate processing. - **Queues** suffer from in-flight message limits (SQS caps how many unacked messages can be outstanding per queue at once, and once that cap is hit no new messages can be received until some are acked or expire) and from poison-pill loops where a message keeps failing and keeps getting redelivered until a `maxReceiveCount` setting finally routes it to a dead-letter queue. ## Two workers, two outcomes A concrete contrast: - **The queue side** - a checkout worker on SQS receives an order message with a 30-second visibility timeout; if the database write actually takes 45 seconds, the message reappears and a second worker may pick it up concurrently, causing double processing unless the write is idempotent. - **On the log side**, an analytics consumer on Kafka commits offsets every 5 seconds in batches of thousands of events; if it crashes mid-batch, it resumes from the last committed offset on restart and reprocesses the last few seconds of events - which is safe here because the downstream aggregation upserts by event ID rather than blindly incrementing counters.
- What's the difference between committing offsets before vs after processing a batch, in terms of delivery guarantees?Committing before processing gives at-most-once - if the consumer crashes mid-batch, those messages are considered done and skipped on restart, silently losing work. Committing after processing (the common default) gives at-least-once - a crash before commit means the same messages get redelivered and reprocessed, so consumers must be idempotent to avoid duplicate side effects.
- Why can SQS run out of capacity to deliver more messages even though the queue has plenty of items waiting?SQS caps the number of in-flight (received-but-not-yet-acked) messages per queue at a fixed limit. If consumers are slow to ack or crash without acking, messages pile up as in-flight until their visibility timeout expires, and no new messages can be received past that cap in the meantime.
- What causes a Kafka consumer group rebalance, and why can it cause brief duplicate processing?A rebalance is triggered when a consumer joins, leaves, or is deemed dead (missed heartbeats) in a group, causing partitions to be reassigned among remaining members. If the outgoing consumer had processed messages past its last committed offset, the next owner of that partition resumes from the older committed offset and reprocesses that gap.
Offsets are like a bookmark in a library book everyone shares - the book itself never changes, you just remember your page number, and different readers can have bookmarks at wildly different pages. Acking is like ripping a page out of a single shared notepad once you've copied down what's on it - it's gone for everyone.
saying these in an interview costs you the question
- Says acking and committing an offset are the same operation
- Thinks reading from a log deletes the message
- Doesn't know offset commits can happen before or after processing, with different guarantees
- Can't explain why a queue caps in-flight messages
- Assumes exactly-once is free with either offsets or acks