skip to content

Dead-Letter Queues & Redrive

You will learn how AWS implements poison-message handling: a redrive policy with a receive counter, a dead-letter queue that must match the source queue type, and the redrive task that replays messages after a fix. Interviewers ask what happens to a message that keeps failing, and expect the operational answer.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

6

In an Amazon SQS redrive policy, what does maxReceiveCount actually count, and what has to happen before a message lands in the dead-letter queue?

level: middleimportance: must knowfreq 78%

answer

  1. not an error counter
  2. the counter rides on the message
  3. a killed consumer still spends one
  4. evaluated on the next delivery attempt
  5. ApproximateReceiveCount versus maxReceiveCount

basics

~20 s

maxReceiveCount caps how many times SQS may deliver one message. Every ReceiveMessage delivery increments that message's ApproximateReceiveCount — whether the consumer failed, crashed, or never answered — and once the count passes the limit, SQS routes the message to the dead-letter queue.

solid answer

~50 s

`maxReceiveCount` is part of the source queue's `RedrivePolicy`, alongside `deadLetterTargetArn`. It is a **delivery** counter, not an error counter. Each message carries a system attribute, `ApproximateReceiveCount`, that SQS increments every time the message is handed to a consumer by `ReceiveMessage`. If the consumer calls `DeleteMessage`, the message is gone and the count is irrelevant. If it does not — because the handler threw, the process was killed, or the visibility timeout simply elapsed while work was still running — the message becomes visible again and the next receive bumps the count. When the count exceeds `maxReceiveCount`, SQS moves the message to the DLQ instead of redelivering it. Nothing in your consumer code sends it there; the move is done by the service, so a consumer that silently swallows an exception and deletes the message will never dead-letter anything.

code

json · 4 lines
json
{
  "deadLetterTargetArn": "arn:aws:sqs:us-east-1:111122223333:orders-dlq",
  "maxReceiveCount": 5
}

go deeper

for a junior

Know that a redrive policy has two parts — the target DLQ ARN and maxReceiveCount — and that SQS, not your code, moves the message once the limit is passed.

for a middle

Be ready to explain that ApproximateReceiveCount rises on every ReceiveMessage delivery, so a timeout or a killed process counts the same as a thrown exception, and that the move is checked on the next delivery.

for a senior

Show that you would diagnose unexpected dead-lettering by comparing handler duration against the visibility timeout before touching maxReceiveCount, and that an always-empty DLQ suggests swallowed errors.

for a principal

Own the guidance for the fleet: how maxReceiveCount times the visibility timeout sets the quarantine window, and why the number should follow from the handler's own retry behaviour rather than being copied across every queue.

## The two attributes that define the behaviour A dead-letter queue in Amazon SQS is not a special kind of queue. It is an ordinary queue that some *other* queue names in its `RedrivePolicy` attribute. That attribute is a JSON document with exactly two fields: ```json { "deadLetterTargetArn": "arn:aws:sqs:us-east-1:111122223333:orders-dlq", "maxReceiveCount": 5 } ``` `deadLetterTargetArn` says where failing messages go. `maxReceiveCount` says how many deliveries a message is allowed before it goes there. The valid range runs from 1 to 1000. ## What increments the counter Every message in a queue carries a system attribute called `ApproximateReceiveCount`. SQS increments it each time the message is returned by a `ReceiveMessage` call. That is the whole rule, and it is the part candidates most often get wrong: **the counter tracks deliveries, not failures.** Concretely, all of these burn one receive: - The handler throws and the code deliberately does not delete the message. - The consumer process is killed mid-processing (a container eviction, an OOM kill, a scale-in). - Processing succeeds but takes longer than the queue's `VisibilityTimeout`, so the message reappears and is picked up by a second consumer while the first is still working. - A consumer receives the message, decides it is not its concern, and drops it on the floor. And this one does *not*: the handler catches an exception, logs it, and calls `DeleteMessage` anyway. The message is deleted, so it will never reach the DLQ no matter what `maxReceiveCount` says. "Our DLQ is always empty" is far more often a symptom of swallowed errors than of a healthy system. You can read the counter yourself: ```bash aws sqs receive-message \ --queue-url https://sqs.us-east-1.amazonaws.com/111122223333/orders \ --attribute-names ApproximateReceiveCount ``` The value is *approximate* on standard queues for the same reason queue depth is: SQS stores messages redundantly across servers, and the count can occasionally be higher than the number of distinct processing attempts. ## When the move actually happens The move is lazy, not eager. SQS does not run a background sweeper that notices "this message has now failed five times" the instant the fifth attempt goes bad. The check happens on the *next* receive: when a message whose count has passed the threshold is about to be delivered again, SQS sends it to the dead-letter queue instead. This has a practical consequence — if no consumer is polling the queue, nothing dead-letters, however broken the messages are. A stopped consumer fleet produces a growing backlog on the source queue and an empty DLQ. When the move happens, the message keeps its body, its message attributes, and its original message ID. It gets a new receipt handle, because receipt handles are per-receive. ## Choosing a value The number encodes how many transient failures you are willing to absorb before declaring a message poisonous. Setting it to 1 means the first blip — a downstream 503, a brief database failover, a rolling deploy that kills a pod mid-message — permanently dead-letters otherwise-fine traffic. Setting it to 100 means a genuinely malformed message ties up consumer capacity a hundred times over and the DLQ tells you nothing useful for hours. A common starting point is somewhere in the range of 3 to 5 for a handler that already retries transient errors in-process, on the reasoning that anything surviving several independent deliveries is probably a data problem rather than a weather problem. The interaction with `VisibilityTimeout` matters at least as much as the number itself. The gap between consecutive receives of a failing message is roughly the visibility timeout, so `maxReceiveCount` multiplied by the visibility timeout is roughly how long a bad message churns before it is quarantined. A 30-second timeout with `maxReceiveCount` of 5 quarantines in a couple of minutes; a 12-hour timeout with the same count would take days. ## The symptom to recognise in an interview "Messages dead-letter under load but process fine in isolation" is almost never a poison-message problem. It is a timing problem: the handler is slower than the visibility timeout under load, so each message is delivered repeatedly to different consumers, the receive count climbs on messages that are actually being processed successfully, and they get quarantined anyway — usually after being processed several times over. The fix is on the timeout and handler side, not on `maxReceiveCount`.

  • A team reports that messages dead-letter under peak load but replay perfectly from the DLQ. What is your first hypothesis?
    That the handler is slower than the queue's `VisibilityTimeout` under load. Each message reappears mid-processing, gets delivered to another consumer, and its `ApproximateReceiveCount` climbs on traffic that is actually succeeding — often several times over. I would compare p99 handler duration to the timeout, then either raise the timeout, call `ChangeMessageVisibility` to extend it, or shrink the batch.
  • Your DLQ has been empty for months. Why might that not be good news?
    Because the only way a message avoids the DLQ is being deleted. A handler that catches an exception, logs it, and deletes anyway destroys the evidence — the receive count never climbs and nothing dead-letters. I would check whether failures are being swallowed, and confirm the consumers are actually polling: with nothing calling `ReceiveMessage`, the move is never evaluated at all.
  • Does the message body or message ID change when SQS moves a message to the dead-letter queue?
    No. The body, the message attributes, and the message ID are preserved, which is what makes a DLQ useful for forensics. The receipt handle is new, because handles are issued per receive, and the receive count restarts relative to the DLQ's own consumers if the DLQ itself has a redrive policy.

saying these in an interview costs you the question

  • Says maxReceiveCount counts consumer exceptions rather than deliveries
  • Thinks the consumer explicitly publishes the message to the DLQ
  • Believes the move happens the instant the handler throws
  • Assumes a crashed or evicted consumer does not increment the count
  • Sets maxReceiveCount to 1 and calls transient failures poison messages

context

open as a page

You set an Amazon SQS FIFO queue's redrive policy to point at a standard queue as its dead-letter queue. What happens, and what are SQS's rules for what a DLQ may be?

level: juniorimportance: should knowfreq 42%

basics

~20 s

SQS rejects the configuration. A dead-letter queue must be the same type as its source queue — FIFO for FIFO, standard for standard — and must live in the same AWS account and Region. Otherwise it is an ordinary queue you create yourself.

open as a page

Which CloudWatch metrics tell you that an Amazon SQS dead-letter queue is filling up, and what does ApproximateAgeOfOldestMessage measure on a DLQ specifically?

level: middleimportance: should knowfreq 44%

basics

~20 s

Alarm on the DLQ's ApproximateNumberOfMessagesVisible at 1 or more — any message there is abnormal. ApproximateAgeOfOldestMessage on a DLQ measures time since the message was moved in, not since it was originally sent, so it tracks triage delay rather than true message age.

open as a page

After deploying a fix, how do you replay the messages sitting in an Amazon SQS dead-letter queue back to the queue they came from, and what should you verify before you start?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Use the SQS message move task: call StartMessageMoveTask with the DLQ's ARN as the source, optionally throttling with MaxNumberOfMessagesPerSecond. Before starting, confirm the fix is actually deployed and consumers are healthy, or the messages will fail again and return straight to the DLQ.

open as a page

Failed orders had been landing in an Amazon SQS dead-letter queue for days, but when the team finally opened it, most were gone — and nothing was consuming that DLQ. What explains it, and how do you prevent it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

They expired. A message's SQS retention clock starts at its original enqueue time and is not reset when the message moves to a dead-letter queue, so time already spent failing on the source queue is deducted from its life in the DLQ. Raise the DLQ's MessageRetentionPeriod and alarm on arrivals.

open as a page

What is the RedriveAllowPolicy attribute on an Amazon SQS queue, and what does its redrivePermission setting control?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

RedriveAllowPolicy is set on a queue acting as a dead-letter queue and declares which source queues may name it as their DLQ. Its redrivePermission field takes allowAll, denyAll, or byQueue with an explicit list of source queue ARNs.

open as a page