skip to content

In an AWS design, why do teams subscribe an SQS queue to an Amazon SNS topic for each consumer instead of subscribing the consumers (Lambda functions, HTTPS endpoints) to the topic directly?

level: middleimportance: must knowfreq 74%

answer

  1. push once, no stored backlog
  2. each consumer needs its own buffer
  3. failure isolation between subscribers
  4. queue owns retries and the DLQ
  5. publish once, subscribe many

basics

~20 s

Amazon SNS pushes once and eventually drops a message it cannot deliver. Putting an SQS queue in front of each consumer stores that consumer's copy durably, so each one retries, backs off and drains at its own rate without losing events.

solid answer

~40 s

SNS is push-only: it makes delivery attempts and keeps nothing you can poll later, so if a subscriber is down, slow or throttled, its copy is retried for a while and then discarded. Putting an SQS queue between the topic and each consumer turns one push into durable per-consumer storage. The queue holds the message until the consumer deletes it, each consumer drains at its own rate, and each gets its own visibility timeout, redrive policy and dead-letter queue, so a broken invoicing service cannot slow down the search indexer. You still publish once, and adding a seventh consumer is a new subscription rather than a producer change. The cost is that standard topics and queues are at-least-once and unordered, so every consumer has to be idempotent.

go deeper

for a junior

Be able to say that SNS pushes a copy to every subscriber and does not hold messages, while SQS stores messages until a consumer deletes them. Naming the fanout shape — one topic, one queue per consumer — is enough at this level.

for a middle

Explain the mechanics: independent retries and dead-letter queues per queue, backpressure through queue depth, and the queue policy that lets SNS write to each queue. Expect to be asked what happens when one consumer is down.

for a senior

Show production judgment about isolation and blast radius: which consumers may lag without harm, what the backlog alarm is, how you make handlers idempotent, and when direct delivery is honestly good enough. Mention the encrypted-queue trap.

for a principal

Own the tradeoff between fanout and routing across a platform: who may subscribe to whose topics, whether teams get a shared bus or per-domain topics, how contracts and payload size are governed, and what the cost and operational load of N queues per event actually is.

## What SNS actually promises Amazon SNS is a push-based publish/subscribe service. A publisher calls `Publish` (or `PublishBatch`) against a topic ARN, and SNS attempts delivery to every subscription on that topic: SQS queues, Lambda functions, HTTPS endpoints, email, SMS, and Firehose delivery streams. Two properties drive every design decision around it: - **There is no storage you can read.** A standard SNS topic has no cursor, no offset, no "give me the last hour" API. Once SNS has attempted delivery to a subscription and exhausted its retries, that copy is gone unless you configured somewhere for it to go. - **Delivery is per subscription and at-least-once.** Each subscriber's fate is independent; a subscriber can legitimately see the same notification twice. ## Why a queue per subscriber The "fanout sandwich" — one topic, N queues, one consumer per queue — exists because the queue supplies exactly what SNS does not: 1. **Durability and backpressure.** The message sits in the queue until the consumer deletes it. A consumer that is deployed, scaled to zero, throttled, or crash-looping simply builds a backlog instead of losing traffic. `ApproximateNumberOfMessagesVisible` becomes your lag signal. 2. **Failure isolation.** Each subscriber owns its own queue, so one slow consumer cannot influence the others. With direct subscriptions, a consistently failing HTTPS endpoint just burns retries and drops messages. 3. **Consumer-owned retry semantics.** The queue gives you a visibility timeout you control, a `maxReceiveCount` and a dead-letter queue per consumer, plus redrive when the bug is fixed. Direct SNS delivery gives you a delivery policy and, at best, a subscription dead-letter queue that nothing redrives automatically. 4. **Rate decoupling.** A burst of 50,000 publishes does not have to become 50,000 concurrent Lambda invocations. The queue lets you batch, cap concurrency, and smooth spikes. 5. **Independent evolution.** Adding a new consumer is one `Subscribe` call plus a queue policy — the publisher never learns about it, which is the whole point of pub/sub. ## Wiring it correctly Two details break this pattern in practice. First, the **queue's own resource policy** must allow the SNS service principal to `sqs:SendMessage`, normally scoped with a condition on `aws:SourceArn` equal to the topic ARN. Subscribing the queue does not by itself grant SNS permission to write to it. Second, most teams enable **raw message delivery** on the subscription so consumers receive the published payload rather than the SNS JSON envelope. ```bash aws sns subscribe \ --topic-arn arn:aws:sns:eu-west-1:111122223333:orders \ --protocol sqs \ --notification-endpoint arn:aws:sqs:eu-west-1:111122223333:orders-billing \ --attributes RawMessageDelivery=true ``` If the queue is encrypted, note that SNS must be able to use the key: an SQS queue encrypted with the AWS managed key for SQS cannot receive SNS deliveries, because you cannot edit that key's policy. Use a customer managed KMS key whose policy grants the SNS service principal `kms:GenerateDataKey*` and `kms:Decrypt`. ## What the sandwich does not fix - **Duplicates.** Standard topics and standard queues are at-least-once end to end. Consumers must be idempotent — typically an idempotency key stored per business event. - **Ordering.** Standard delivery does not preserve order, and two subscribers can see different orders. Ordering requires a FIFO topic feeding FIFO queues. - **Large payloads.** SNS messages are small (256 KB as of 2025). Big payloads use the claim-check pattern: write the object to S3, publish the pointer. - **Unnecessary work.** Every queue receives every message unless you attach a filter policy to the subscription, so subscribers that only care about a slice should filter server-side rather than receive-and-discard. ## When to skip the queue Direct subscriptions are reasonable for genuinely fire-and-forget, low-volume paths — an operational alert to an HTTPS webhook or a chat integration, where losing one notification after retries is acceptable and a queue is ceremony. A direct Lambda subscription is also fine when the function is trivially fast, the volume is low, and you have configured Lambda's own failure destination. The moment the consumer does real work, calls a database, or must not lose events, put the queue in. Finally, if the requirement is not "everyone gets a copy" but "route this event to the right target based on its content, including events emitted by AWS services themselves," that is a different primitive and a different service choice.

  • If the queue absorbs the retries, why keep SNS at all instead of having the producer write to each queue?
    Because that recouples the producer to its consumers. With SNS the producer makes one `Publish` call and never learns who listens; adding or removing a consumer is a subscription change in the consumer's own account and code. Direct multi-queue writes also make partial failure the producer's problem — two queues written, one failed, now what.
  • What guarantees does this pattern give up, and how do you compensate?
    Standard topics and queues are at-least-once and unordered, so a consumer can see duplicates and can see events out of sequence relative to another consumer. Compensate with idempotent handlers keyed on a business identifier, and by designing state transitions that tolerate replay. If ordering is genuinely required, move to a FIFO topic feeding FIFO queues and accept much lower throughput.
  • How would you tell, from CloudWatch alone, that a subscriber is not receiving what you think it is?
    Compare the topic's `NumberOfMessagesPublished` against `NumberOfNotificationsDelivered` and `NumberOfNotificationsFailed`, and check `NumberOfNotificationsFilteredOut` — a high filtered-out count usually means a filter policy is stricter than intended or the publisher stopped sending the attribute. On the queue side, a flat `NumberOfMessagesSent` while publishes continue points at the queue policy or the KMS key.

saying these in an interview costs you the question

  • SNS stores messages until a consumer is ready to read
  • SNS guarantees exactly-once delivery to each subscriber
  • You can replay a topic by re-reading past messages
  • Subscribing the queue automatically grants SNS permission to write
  • SNS preserves publish order across all subscribers

context