skip to content

You're designing a system to send password-reset emails: an API call enqueues a 'send reset email' task, and a worker pool sends it exactly once soon after. Would you reach for a destructive-read queue like SQS or an event log like Kafka, and why?

level: seniorimportance: should knowfreq 55%

answer

  1. command vs fact test
  2. one consumer vs many = queue vs log signal
  3. replay wanted or must-prevent
  4. DLQ maps to retry-then-alert
  5. PII retention cost of a log

basics

~20 s

A queue - you want the task done once by one worker and then forgotten, with automatic retry if it fails. A log is overkill here because nobody else needs to replay or re-read 'send this email' later; you'd actually want to avoid accidentally resending it.

solid answer

~40 s

A queue is the better fit: this is a discrete unit-of-work with a single consumer group (the worker pool), where 'exactly one worker does it, then it's done' is the desired semantic, and built-in features like visibility timeout, automatic redelivery on failure, and dead-lettering after N attempts map directly onto 'retry a failed send, then give up and alert.' A log would add unnecessary complexity - you'd have to build your own 'mark as done' bookkeeping to prevent an idle-then-resumed consumer from resending old emails, and you gain nothing from multi-consumer fan-out or replay, since replaying 'send email' tasks is actively undesirable.

go deeper

for a junior

Should pick the queue and give one reason (retry/DLQ fits a one-time task).

for a middle

Should articulate the command-vs-fact framing at a basic level and name at least one concrete queue feature (visibility timeout, DLQ) that fits.

for a senior

Should name the deciding question (single vs potentially-multiple independent consumers, replay wanted or not) and identify a counter-scenario where the same event would flip to log-shaped.

for a principal

Should weigh secondary costs like PII/retention exposure and design an idempotency safety net regardless of which primitive is chosen, since delivery guarantees alone don't guarantee exactly-once execution.

## The three questions that decide it The decision hinges on three questions: 1. Is this a **command** (do this exactly once) or a **fact** (this happened, and multiple parties may care about it now or later)? 2. Do you need multiple independent consumers, now or foreseeably? 3. And is replay something you want, or something you must actively prevent? Sending a password-reset email is a command with one logical consumer, and replay is undesirable - all three point toward a queue. ## Why the mechanism favors a queue The mechanism argument favors a queue directly: SQS or RabbitMQ's core primitives - ack-to-delete, visibility timeout with automatic redelivery, and a dead-letter queue after a configured `maxReceiveCount` - exist precisely to implement 'try this task, retry on failure, give up after N tries and alert a human' with minimal custom code. A worker pool of N instances naturally load-balances via competing consumers, with no partition-key design required. ## Why a log fits badly here A log would be awkward here for the opposite reasons. - To prevent double-sends on a log, you'd need to persist 'have I already sent for offset X' somewhere yourself - essentially reinventing acking on top of the log. - And you'd need to reason about consumer-group rebalances potentially reprocessing a task mid-send. - You'd gain none of the log's headline benefits, since nobody else should be reading 'send email' tasks, and replay here is a bug magnet rather than a feature. ## The counter-scenario that flips it A counter-scenario sharpens the boundary: if instead the event were 'PasswordResetRequested' published as a fact (not a task), and consumed independently by an email-sending service, a security anomaly-detection service, AND an audit-log archiver, that's now multiple independent consumers reacting to one occurrence - textbook log territory. The tell is whether new, unforeseen consumers might reasonably want the same occurrence in the future; a queue's 'do this job' framing assumes a fixed, known consumer, while a log's 'this happened' framing keeps the door open. ## What each choice costs There are real costs on both sides of this specific choice. - **Choosing a queue** means no free audit trail (you'd add application-level logging separately) and no easy path to adding a new service that reacts to reset requests later without extra plumbing. - **Choosing a log** instead for this case would cost extra engineering to prevent resends (a dedup store or offset-based idempotency key), harder-to-reason-about retry semantics (hand-rolling a DLQ equivalent), and paying for retention storage of data you'd rather not keep around - reset emails often carry sensitive tokens, and a log holding those for a multi-day retention window is a larger PII exposure surface than a queue that deletes the payload the moment it's sent. ## The resolution The resolution for this specific case: use SQS. - The API handler enqueues a task containing the user ID, reset token, and timestamp. - A worker pool of email senders additionally enforces idempotency by checking a short-TTL key in a fast store (e.g., Redis) keyed on the reset token before sending, defending against the rare double-delivery during a visibility-timeout race. - Failures retry up to a configured limit then land in a DLQ that pages on-call. - And the token itself expires quickly, so even a very late redelivery is naturally harmless rather than dangerous.

  • What single question would you ask in a requirements-gathering session to decide between a queue and a log for a new integration?
    Ask: 'Will more than one independent service ever need to react to this same occurrence, including services we haven't built yet?' If the honest answer is 'just this one worker pool, doing one job,' lean queue; if the answer is 'possibly, or definitely more than one today,' lean log.
  • How would you add basic idempotency protection to the SQS-based email-sending worker without switching to a log?
    Before sending, check a fast key-value store, such as Redis, for a short-TTL key derived from the reset token or message ID; if present, skip sending since it was already handled. Set the key right before or after a successful send - this guards against the rare double-delivery from a visibility-timeout race without needing log-style replay machinery.
  • If the password-reset feature later needs a compliance audit trail of every reset requested, does that change the queue-vs-log decision?
    It adds a second, independent consumer - an audit or archival service - reacting to the same 'PasswordResetRequested' fact, which is exactly the multi-consumer signal that favors a log. At that point, either publish the fact to a log in addition to enqueueing the send task, or route both the queue and an audit sink off a shared event published once.

Sending a password-reset email is like handing a single sealed task to one courier and needing confirmation of delivery - if they fail, you retry with another courier, but you never want the 'deliver this' instruction replayed to ten couriers later. A log is better suited to announcing news that many different departments might independently want to hear about, now or in the future.

saying these in an interview costs you the question

  • Reaches for Kafka by default for a single-worker task queue
  • Doesn't distinguish 'command, single consumer' from 'fact, many consumers' as the deciding question
  • Thinks a log is strictly 'better' or 'more modern' regardless of use case
  • Proposes replay as a feature for a send-email task
  • Ignores the PII/retention cost of storing sensitive tokens in a longer-retention log

context