In an Amazon SQS redrive policy, what does maxReceiveCount actually count, and what has to happen before a message lands in the dead-letter queue?
answer
- not an error counter
- the counter rides on the message
- a killed consumer still spends one
- evaluated on the next delivery attempt
- ApproximateReceiveCount versus maxReceiveCount
basics
~20 smaxReceiveCount caps how many times SQS may deliver one message. Every ReceiveMessage delivery increments that message's ApproximateReceiveCount — whether the consumer failed, crashed, or never answered — and once the count passes the limit, SQS routes the message to the dead-letter queue.
solid answer
~50 s`maxReceiveCount` is part of the source queue's `RedrivePolicy`, alongside `deadLetterTargetArn`. It is a **delivery** counter, not an error counter. Each message carries a system attribute, `ApproximateReceiveCount`, that SQS increments every time the message is handed to a consumer by `ReceiveMessage`. If the consumer calls `DeleteMessage`, the message is gone and the count is irrelevant. If it does not — because the handler threw, the process was killed, or the visibility timeout simply elapsed while work was still running — the message becomes visible again and the next receive bumps the count. When the count exceeds `maxReceiveCount`, SQS moves the message to the DLQ instead of redelivering it. Nothing in your consumer code sends it there; the move is done by the service, so a consumer that silently swallows an exception and deletes the message will never dead-letter anything.
code
json · 4 lines{
"deadLetterTargetArn": "arn:aws:sqs:us-east-1:111122223333:orders-dlq",
"maxReceiveCount": 5
}go deeper
Know that a redrive policy has two parts — the target DLQ ARN and maxReceiveCount — and that SQS, not your code, moves the message once the limit is passed.
Be ready to explain that ApproximateReceiveCount rises on every ReceiveMessage delivery, so a timeout or a killed process counts the same as a thrown exception, and that the move is checked on the next delivery.
Show that you would diagnose unexpected dead-lettering by comparing handler duration against the visibility timeout before touching maxReceiveCount, and that an always-empty DLQ suggests swallowed errors.
Own the guidance for the fleet: how maxReceiveCount times the visibility timeout sets the quarantine window, and why the number should follow from the handler's own retry behaviour rather than being copied across every queue.
## The two attributes that define the behaviour A dead-letter queue in Amazon SQS is not a special kind of queue. It is an ordinary queue that some *other* queue names in its `RedrivePolicy` attribute. That attribute is a JSON document with exactly two fields: ```json { "deadLetterTargetArn": "arn:aws:sqs:us-east-1:111122223333:orders-dlq", "maxReceiveCount": 5 } ``` `deadLetterTargetArn` says where failing messages go. `maxReceiveCount` says how many deliveries a message is allowed before it goes there. The valid range runs from 1 to 1000. ## What increments the counter Every message in a queue carries a system attribute called `ApproximateReceiveCount`. SQS increments it each time the message is returned by a `ReceiveMessage` call. That is the whole rule, and it is the part candidates most often get wrong: **the counter tracks deliveries, not failures.** Concretely, all of these burn one receive: - The handler throws and the code deliberately does not delete the message. - The consumer process is killed mid-processing (a container eviction, an OOM kill, a scale-in). - Processing succeeds but takes longer than the queue's `VisibilityTimeout`, so the message reappears and is picked up by a second consumer while the first is still working. - A consumer receives the message, decides it is not its concern, and drops it on the floor. And this one does *not*: the handler catches an exception, logs it, and calls `DeleteMessage` anyway. The message is deleted, so it will never reach the DLQ no matter what `maxReceiveCount` says. "Our DLQ is always empty" is far more often a symptom of swallowed errors than of a healthy system. You can read the counter yourself: ```bash aws sqs receive-message \ --queue-url https://sqs.us-east-1.amazonaws.com/111122223333/orders \ --attribute-names ApproximateReceiveCount ``` The value is *approximate* on standard queues for the same reason queue depth is: SQS stores messages redundantly across servers, and the count can occasionally be higher than the number of distinct processing attempts. ## When the move actually happens The move is lazy, not eager. SQS does not run a background sweeper that notices "this message has now failed five times" the instant the fifth attempt goes bad. The check happens on the *next* receive: when a message whose count has passed the threshold is about to be delivered again, SQS sends it to the dead-letter queue instead. This has a practical consequence — if no consumer is polling the queue, nothing dead-letters, however broken the messages are. A stopped consumer fleet produces a growing backlog on the source queue and an empty DLQ. When the move happens, the message keeps its body, its message attributes, and its original message ID. It gets a new receipt handle, because receipt handles are per-receive. ## Choosing a value The number encodes how many transient failures you are willing to absorb before declaring a message poisonous. Setting it to 1 means the first blip — a downstream 503, a brief database failover, a rolling deploy that kills a pod mid-message — permanently dead-letters otherwise-fine traffic. Setting it to 100 means a genuinely malformed message ties up consumer capacity a hundred times over and the DLQ tells you nothing useful for hours. A common starting point is somewhere in the range of 3 to 5 for a handler that already retries transient errors in-process, on the reasoning that anything surviving several independent deliveries is probably a data problem rather than a weather problem. The interaction with `VisibilityTimeout` matters at least as much as the number itself. The gap between consecutive receives of a failing message is roughly the visibility timeout, so `maxReceiveCount` multiplied by the visibility timeout is roughly how long a bad message churns before it is quarantined. A 30-second timeout with `maxReceiveCount` of 5 quarantines in a couple of minutes; a 12-hour timeout with the same count would take days. ## The symptom to recognise in an interview "Messages dead-letter under load but process fine in isolation" is almost never a poison-message problem. It is a timing problem: the handler is slower than the visibility timeout under load, so each message is delivered repeatedly to different consumers, the receive count climbs on messages that are actually being processed successfully, and they get quarantined anyway — usually after being processed several times over. The fix is on the timeout and handler side, not on `maxReceiveCount`.
- A team reports that messages dead-letter under peak load but replay perfectly from the DLQ. What is your first hypothesis?That the handler is slower than the queue's `VisibilityTimeout` under load. Each message reappears mid-processing, gets delivered to another consumer, and its `ApproximateReceiveCount` climbs on traffic that is actually succeeding — often several times over. I would compare p99 handler duration to the timeout, then either raise the timeout, call `ChangeMessageVisibility` to extend it, or shrink the batch.
- Your DLQ has been empty for months. Why might that not be good news?Because the only way a message avoids the DLQ is being deleted. A handler that catches an exception, logs it, and deletes anyway destroys the evidence — the receive count never climbs and nothing dead-letters. I would check whether failures are being swallowed, and confirm the consumers are actually polling: with nothing calling `ReceiveMessage`, the move is never evaluated at all.
- Does the message body or message ID change when SQS moves a message to the dead-letter queue?No. The body, the message attributes, and the message ID are preserved, which is what makes a DLQ useful for forensics. The receipt handle is new, because handles are issued per receive, and the receive count restarts relative to the DLQ's own consumers if the DLQ itself has a redrive policy.
saying these in an interview costs you the question
- Says maxReceiveCount counts consumer exceptions rather than deliveries
- Thinks the consumer explicitly publishes the message to the DLQ
- Believes the move happens the instant the handler throws
- Assumes a crashed or evicted consumer does not increment the count
- Sets maxReceiveCount to 1 and calls transient failures poison messages