A consumer service has a bug that silently corrupts 3 days' worth of derived data before anyone notices. If the upstream system is a Kafka-style event log with 14-day retention, how would you recover, and why would the same recovery be much harder (or impossible) if the upstream were an SQS queue instead?
answer
- retention.ms independent of consumption
- seek/offset reset = replay
- replay needs idempotent writes
- SQS retention only helps unconsumed messages
- outbox/archive pattern backstops queues
basics
~20 sWith a log, you can rewind your reader to before the bug started and reprocess those 3 days from the original messages, since they're all still stored. With a queue, those messages are already deleted once the first, buggy pass consumed them, so there's nothing left to replay - you'd need another way to get that data back.
solid answer
~30 sBecause the log retains every message for the configured window regardless of consumption, recovery is: fix the bug, reset the consumer group's offset to just before the corruption started, and reprocess forward - the source data is untouched, so replay reproduces exactly what happened, assuming the derived-data writes are idempotent (upserts, not blind appends) so replaying doesn't double-count. With SQS, once a message is acked it's gone from the broker permanently; there's no source of truth to replay from unless the raw messages were separately archived elsewhere before processing - the destructive-read model has no built-in undo.
go deeper
Should grasp that a log lets you go back and reread old messages while a queue does not, once a message is acked.
Should be able to describe the mechanic (offset reset/seek, retention window) in concrete terms.
Should identify idempotency as the precondition for safe replay and know retention is a bounded, costed resource, not infinite.
Should propose architectural backstops (outbox pattern, cold archival, tiered storage) for systems that need replay guarantees beyond raw retention or a destructive queue.
## How replay actually works A log's retention is configured per topic by time (e.g., `retention.ms`) or size (`retention.bytes`), and log segments are only deleted once they age out or exceed that size - entirely independent of what any consumer group has read. Replaying means: - telling a consumer to `seek()` to an earlier offset (or a timestamp, via tooling that translates time to offset), - or standing up a brand-new consumer group and resetting its starting position backward. Because messages are immutable once written, replaying gives byte-identical input to the corrected processing logic - there's no ambiguity about what 'actually happened.' ## Why a log is built this way This exists because logs are meant to be a durable source of truth, analogous to a database's write-ahead log but externally readable by many parties. That's a deliberate design choice: it decouples 'when did I read this' from 'does this still exist,' which enables recovery from bugs, reprocessing for backfills, and late-arriving consumers, without requiring anyone to have planned for that specific failure in advance. ## What that durability costs That durability isn't free. - **Retention costs real disk** - 14 days of a high-volume topic, replicated three times for durability, can be many terabytes - so retention length is a genuine cost/capability trade-off, and if a bug isn't caught within the retention window, replay from the log alone becomes impossible, exactly like a queue. - **Teams that need 'replay anytime, no matter how late we notice'** typically add a separate cold archive (tiered storage, or a sink to a data lake) rather than simply extending retention indefinitely. - **Queues, in exchange for offering none of this,** stay operationally simpler - no retention tuning, no offset management - at the structural cost of zero recovery capability once a message is acknowledged. ## Failure modes Several failure modes show up in practice. 1. **First**, replaying into a non-idempotent downstream sink causes double-counting: replaying 3 days of 'payment succeeded' events into a naive running `SUM(amount)` inflates totals unless writes are upserts keyed by a stable event ID, or the aggregate is fully recomputed rather than incrementally added to. 2. **Second**, replaying past the retention boundary doesn't error - it silently returns whatever's oldest still available - so an on-call engineer who asks to seek 20 days back on a 14-day-retention topic may believe recovery succeeded while data from days 15 through 20 is actually missing. 3. **Third**, on the queue side, teams sometimes try to retrofit replay by raising SQS's message retention period (up to 14 days) - but that only helps for messages still sitting unconsumed in the queue; once a message has been acked and deleted, no retention setting brings it back, making an after-the-fact 'we need to reprocess last Tuesday' request a dead end unless raw payloads were persisted somewhere else first. ## A worked recovery A worked scenario: an e-commerce platform ingests `OrderPlaced` events into Kafka with 14-day retention. A newly deployed tax-calculation service misapplies a rate for 3 days before a finance audit catches it. - **Recovery**: patch the bug, spin up a repair consumer group that seeks to the offset just before the bad deploy, replays those 3 days of `OrderPlaced` events through the corrected logic, and upserts corrected tax records keyed by order ID, so the replay is safe even where it overlaps already-correct records. - **Contrast** this with the same `OrderPlaced` data flowing through SQS with no durable archive: the only recovery path becomes reconstructing affected orders from the orders database itself, if it happens to retain enough detail, and manually re-publishing synthetic messages - slower, riskier, and only possible because a different system happened to preserve the data, not because the messaging layer itself supported it. ## The pattern teams settle on Because of this asymmetry, teams that use SQS or RabbitMQ for task distribution often separately publish a durable copy of significant events to a log, or to a database via an outbox pattern, purely to retain replay capability - effectively layering a log underneath a destructive-read queue rather than relying on the queue for both delivery and history.
- Why must downstream writes be idempotent for replay to be safe, and what's a concrete technique to achieve that?Replay reprocesses the same events a second time, so any non-idempotent operation like total += amount would double the value on replay. The standard fix is to make writes upserts keyed by a stable identifier, such as the event's unique ID or the source offset, so reprocessing the same event produces the same end state rather than accumulating.
- If retention is 14 days and the bug wasn't caught for 20 days, what are your options?The log itself no longer has the data, so you must fall back to any secondary durable copy - an archived S3 dump of raw events, a data-lake sink connector's output, or reconstructing state from a downstream system that still holds enough detail - none of which are guaranteed to exist unless someone planned for exactly this.
- What's the outbox pattern and how does it relate to giving a destructive-read queue some replay capability?The outbox pattern has a service write events to a durable outbox table (or a log) as part of the same transaction as its business writes, then a separate process publishes from that outbox to the queue. Because the outbox itself persists, you retain a durable, replayable record even though the queue consumers still get standard destructive-read delivery.
Replaying a log is like rewinding a security camera's recorded footage to see what happened three days ago - the tape kept rolling regardless of who watched live. A queue is like a live guard who reports what they see once and then forgets it forever - if nobody wrote it down elsewhere, it's unrecoverable.
saying these in an interview costs you the question
- Says you can replay from SQS after messages are acked
- Doesn't mention idempotency as a requirement for safe replay
- Thinks Kafka retention is unlimited by default
- Assumes seeking past the retention window throws an error rather than silently returning less data
- Has no answer for how to add replay capability to a queue-based system