An HTTPS endpoint subscribed to an Amazon SNS topic is unavailable for an hour. What does SNS do with the notifications during that time, and how do you stop them from being lost?
answer
- non-2xx or timeout is a failed attempt
- retries are bounded and configurable
- exhausted means discarded, not stored
- redrive policy lives on the subscription
- nothing replays the DLQ for you
basics
~20 sSNS retries each failed delivery according to that subscription's delivery policy and then discards the message. To preserve it, attach a redrive policy naming a dead-letter SQS queue on the subscription, and enable delivery-status logging so failures are visible.
solid answer
~50 sSNS treats a timeout or a non-2xx response as a failed delivery attempt and retries per the subscription's delivery policy, which for HTTP/S you can tune — number of retries, minimum and maximum delay, and the backoff function. When those retries are exhausted the notification is dropped silently; nothing in the publish path ever noticed. The fix is a **subscription-level redrive policy** pointing at an SQS dead-letter queue: undeliverable notifications land there instead of disappearing, and you can inspect and re-drive them once the endpoint recovers. Alongside it, turn on delivery-status logging and alarm on `NumberOfNotificationsFailed`, because otherwise an hour of loss looks exactly like an hour of quiet. For anything that must not be lost, the stronger answer is not to subscribe the endpoint directly at all — put an SQS queue in the middle and let a worker call the endpoint.
go deeper
Know that SNS retries a failed delivery a limited number of times and then gives up, and that a dead-letter queue can be attached so the notification is not lost. Do not claim SNS holds messages until the endpoint returns.
Explain the delivery policy and its retry/backoff knobs, that the redrive policy is a subscription attribute pointing at an SQS queue, and that the queue policy must permit SNS to write to it.
Show the operational path end to end: alarm on failed notifications and DLQ depth, use delivery-status logging to see the actual status codes, plan the manual replay, and argue for a queue in front of any endpoint whose messages matter.
Own the policy across the estate: which classes of event may use direct HTTPS delivery at all, who owns DLQ drain runbooks, what the acceptable loss budget is, and how partner endpoints are onboarded so one slow consumer cannot quietly shed a platform's events.
## What counts as a failure For an HTTP/S subscription, Amazon SNS considers a delivery attempt failed when the endpoint does not respond within the timeout, refuses the connection, or answers with something outside the 2xx range. There is no negotiation and no half-success: the notification either got a good response or it is queued for retry inside SNS. ## Retries are governed by the delivery policy Each subscription carries a `DeliveryPolicy` (and the topic an `EffectiveDeliveryPolicy`) whose `healthyRetryPolicy` block controls the retry schedule: how many attempts, the minimum and maximum delay between them, and the backoff function (linear, geometric, exponential). The important properties to internalise rather than memorise: - Retries are **bounded**. They are generous enough to ride out a deploy or a brief blip, not an outage of arbitrary length. - Retries are **per subscription**. A failing HTTPS endpoint does not affect the SQS subscriber next to it. - Retry schedules are **protocol-specific**. AWS-internal targets such as SQS and Lambda behave differently from customer-owned HTTP/S endpoints, which are the ones you can actually tune. - When retries run out, the notification is **discarded**. `Publish` returned success long ago; nothing propagates back to the producer. ## The subscription dead-letter queue The mechanism that turns "discarded" into "kept" is a redrive policy set on the *subscription* (not the topic): ```bash aws sns set-subscription-attributes \ --subscription-arn arn:aws:sns:eu-west-1:111122223333:orders:8f1c2d3e-4a5b-6c7d-8e9f-0a1b2c3d4e5f \ --attribute-name RedrivePolicy \ --attribute-value '{"deadLetterTargetArn":"arn:aws:sqs:eu-west-1:111122223333:orders-https-dlq"}' ``` Points candidates commonly miss: - The DLQ is an **SQS queue you own**, and its queue policy must allow the SNS service principal to send to it, exactly like any other SNS-to-SQS delivery. - It is **per subscription**. Five subscriptions on a topic can have five different DLQs, and a subscription with no redrive policy still loses its messages while its neighbour's are preserved. - Messages land there in the shape that subscription would have received, so you are storing notifications, not raw business events, unless raw delivery was on. - **Nothing redrives automatically.** Getting the messages back into the flow is your job: a small consumer that replays them to the endpoint, or a re-publish once the cause is fixed. ## Seeing the failure at all The worst property of dropped notifications is silence. Two mechanisms make it visible: 1. **CloudWatch metrics on the topic** — `NumberOfNotificationsFailed` is the alarm you want, ideally compared against `NumberOfMessagesPublished` and `NumberOfNotificationsDelivered`. A DLQ that is filling should also alarm on its own depth. 2. **Delivery status logging** — configuring the success/failure feedback role attributes on the topic (for example `HTTPSuccessFeedbackRoleArn` and `HTTPFailureFeedbackRoleArn`, with a sample rate) writes per-delivery outcomes to CloudWatch Logs, including the status code the endpoint returned. That is what turns "deliveries are failing" into "the endpoint is returning 502 on this path". ## The design-level answer A senior answer does not stop at "add a DLQ". If the notification matters, a direct HTTPS subscription is the wrong shape: SNS's retry budget is fixed, the endpoint's availability is not yours to control, and a DLQ is a manual recovery path. Subscribe an SQS queue instead and have a worker call the endpoint. Now the retry policy is yours, the backlog is durable for the queue's retention period, backpressure is visible as queue depth, and poison messages have a real redrive story. Reserve direct HTTPS subscriptions for genuinely fire-and-forget notifications — chat messages, low-value alerts — where a bounded retry and a DLQ are proportionate. Also worth saying: an endpoint that is slow rather than down is the more dangerous case. Timeouts consume the retry budget without ever succeeding, so a degraded downstream can burn through the whole schedule and dump a large batch into the DLQ in one go. Make the endpoint acknowledge fast and do its work asynchronously — accepting into local storage and returning 200 promptly is precisely why a queue in front is the better architecture.
- Where exactly does the redrive policy live, and what does it need on the other side?On the subscription, as the `RedrivePolicy` attribute naming a `deadLetterTargetArn` that points at an SQS queue you own. That queue's resource policy must allow the SNS service principal to `sqs:SendMessage`, and if it is encrypted with a customer managed KMS key, the key policy must let SNS generate a data key. Without those, the DLQ silently stays empty.
- How would you get the dead-lettered notifications back into normal processing?Manually, because SNS does not replay them. Fix the endpoint, then run a consumer over the dead-letter queue that re-posts each notification to the endpoint (or re-publishes it to the topic if all subscribers should see it again), deleting each message only after success. Deduplicate on a business key, since some of them may already have been processed.
- An endpoint responds in 25 seconds instead of failing outright. Why is that worse than being down?Slow responses consume the retry budget through timeouts while also holding resources, so the subscription can exhaust its schedule and dump a burst into the dead-letter queue without a single clean failure signal. It also masks the problem in metrics. The fix is an endpoint that acknowledges quickly and does its work asynchronously — which is an argument for a queue in front of it.
saying these in an interview costs you the question
- SNS retries indefinitely until the endpoint recovers
- A failed delivery surfaces as an error to the publisher
- The dead-letter queue is configured on the topic
- Messages in the subscription DLQ are automatically redriven
- A 500 response means the message is dropped immediately