skip to content

Failed orders had been landing in an Amazon SQS dead-letter queue for days, but when the team finally opened it, most were gone — and nothing was consuming that DLQ. What explains it, and how do you prevent it?

level: seniorimportance: should knowfreq 45%

answer

  1. nothing consumed them, yet they left
  2. the clock did not restart on the move
  3. the source queue already spent part of the budget
  4. default four days, maximum fourteen
  5. retention is a deadline, not a workflow

basics

~20 s

They expired. A message's SQS retention clock starts at its original enqueue time and is not reset when the message moves to a dead-letter queue, so time already spent failing on the source queue is deducted from its life in the DLQ. Raise the DLQ's MessageRetentionPeriod and alarm on arrivals.

solid answer

~50 s

Nothing consumed them — they aged out. SQS computes expiry from a message's **original** enqueue timestamp, and moving it to a dead-letter queue does not reset that timestamp. So if a message spent a day being retried on the source queue and the DLQ keeps the default `MessageRetentionPeriod` of four days, it is deleted after three days in the DLQ, not four. Nothing logs it and no metric spikes; the messages simply stop existing. The fixes are layered: set the DLQ's `MessageRetentionPeriod` to the maximum of 14 days, since the whole point of the queue is forensics; keep the source queue's retention short relative to that so less of the budget is consumed before the move; and alarm on the DLQ having any messages at all so triage happens in minutes rather than days. The retention period is the deadline, not the workflow.

code

bash · 3 lines
bash
aws sqs set-queue-attributes \
  --queue-url https://sqs.us-east-1.amazonaws.com/111122223333/orders-dlq \
  --attributes MessageRetentionPeriod=1209600

go deeper

for a junior

Know that every SQS message expires after the queue's MessageRetentionPeriod — four days by default, 14 days at most — and that expiry is silent.

for a middle

Explain that expiry is computed from the original enqueue timestamp and is not reset by the move, so time spent failing on the source queue is deducted from the message's life in the DLQ.

for a senior

Diagnose the disappearance without guessing: relate VisibilityTimeout times maxReceiveCount to how much budget the source queue consumed, and separate the retention ceiling from the alarm that actually forces timely triage.

for a principal

Set the estate-wide policy: 14-day retention on every DLQ, an arrival alarm with a named owner, and a durable sink for message classes whose loss is unacceptable — because a queue is a deadline, not an archive.

## The rule that surprises people In Amazon SQS, `MessageRetentionPeriod` is a per-queue attribute: minimum 60 seconds, default 4 days (345,600 seconds), maximum 14 days (1,209,600 seconds). When it elapses, SQS deletes the message. There is no event, no CloudWatch metric for expiry, no archive — it is gone. The subtlety is *when the clock starts*. AWS's rule is that expiration is always based on the message's original enqueue timestamp, and moving a message to a dead-letter queue does **not** change that timestamp. The consequence is arithmetic: > A message spends 1 day on the source queue being retried. The DLQ's retention is 4 days. The message is deleted 3 days after it arrives in the DLQ. The DLQ's retention period is not a fresh budget granted on arrival. It is the total lifetime of the message, measured from the moment your producer first called `SendMessage`, and the source queue has already spent part of it. ## Why the delay before the move can be large How much budget the source queue consumes is a function of two settings you probably chose for unrelated reasons. A failing message is redelivered roughly once per `VisibilityTimeout`, and it takes `maxReceiveCount` deliveries to quarantine it. So the time to dead-letter is approximately: ``` timeToDeadLetter ≈ VisibilityTimeout × maxReceiveCount ``` A 30-second timeout with `maxReceiveCount` of 5 burns a couple of minutes — negligible. But a long-running job with a 1-hour visibility timeout and `maxReceiveCount` of 10 burns ten hours before the message even reaches the DLQ, and that comes straight off its shelf life there. Teams that raise the visibility timeout to accommodate slow handlers often do not notice they have shortened their forensic window. ## The second trap: the age metric lies about this The natural instinct is "I will alarm on `ApproximateAgeOfOldestMessage` on the DLQ and catch it before expiry." That metric does not measure what you need here. On a dead-letter queue, `ApproximateAgeOfOldestMessage` reflects **when the message moved into the DLQ**, not when it was originally sent. So the metric reads 3 days while the message has only 1 day of actual life left. An alarm calibrated against the retention period using that metric will fire after the evidence is already gone. The usable signal is different: alarm on the DLQ having any messages at all. ## The layered fix **1. Maximise the DLQ's retention.** A dead-letter queue exists so a human can look at what failed. Set it to 14 days at creation and treat that as the default for every DLQ in the estate: ```bash aws sqs set-queue-attributes \ --queue-url https://sqs.us-east-1.amazonaws.com/111122223333/orders-dlq \ --attributes MessageRetentionPeriod=1209600 ``` There is no extra charge for retention in SQS — you pay per request, not per message-day — so there is no reason to leave a DLQ at the default. **2. Shorten the runway to the DLQ.** Keep `VisibilityTimeout × maxReceiveCount` small enough that a poison message is quarantined in minutes, not hours. If the handler genuinely needs a long visibility timeout, use `ChangeMessageVisibility` to extend it heartbeat-style while working rather than setting a permanently huge timeout on the queue. **3. Alarm on arrival, not on age.** A CloudWatch alarm on the DLQ's `ApproximateNumberOfMessagesVisible` greater than or equal to 1 turns a silent seven-day countdown into a page within minutes. If retention is the safety net, the alarm is the actual control. **4. Persist what you cannot afford to lose.** If the messages represent business events that must never be dropped — payments, regulatory records — retention is the wrong durability story regardless of how long you set it. Drain the DLQ into durable storage: a consumer that writes each message to S3 or a table, or an EventBridge Pipe fed from the DLQ. Fourteen days is a deadline, and a queue is not an archive. ## What a strong answer sounds like Name the mechanism (retention measured from original enqueue), do the arithmetic out loud (source-queue time is deducted), and then separate the two controls: retention is the outer bound, the alarm is what actually saves you. Candidates who only say "increase the retention period" have fixed the symptom and left the seven-day silent countdown in place.

  • Would alarming on ApproximateAgeOfOldestMessage on the DLQ have caught this in time?
    No, and that is the trap. On a dead-letter queue that metric measures time since the message *moved in*, not since it was originally sent, so it under-reports the message's true age by however long the source queue spent retrying. An alarm tuned against the retention period using it fires late. Alarm on `ApproximateNumberOfMessagesVisible` being at least 1 instead.
  • The messages represent payments that legally cannot be dropped. Is a 14-day retention enough?
    No. Fourteen days is a hard ceiling with no extension, and any queue-based retention is a deadline rather than durability. Drain the DLQ continuously into real storage — a consumer or an EventBridge Pipe writing each message to S3 or a database — so the queue is a transport for failures, not the system of record for them.
  • How does a long VisibilityTimeout shorten the useful life of a message in the DLQ?
    Because time to dead-letter is roughly VisibilityTimeout multiplied by maxReceiveCount, and all of it is deducted from the message's single retention budget. A one-hour timeout with maxReceiveCount of 10 burns about ten hours before the message even arrives in the DLQ. Prefer a short queue timeout extended per message with `ChangeMessageVisibility` when a handler genuinely runs long.

saying these in an interview costs you the question

  • Assumes the retention clock restarts when a message enters the DLQ
  • Thinks expired messages are logged or emit a CloudWatch event
  • Believes the DLQ inherits the source queue's retention period
  • Says ApproximateAgeOfOldestMessage on the DLQ tracks the original send time
  • Treats a dead-letter queue as long-term archival storage

context