skip to content

Which CloudWatch metrics tell you that an Amazon SQS dead-letter queue is filling up, and what does ApproximateAgeOfOldestMessage measure on a DLQ specifically?

level: middleimportance: should knowfreq 44%

answer

  1. on a DLQ, presence is the threshold
  2. depth means capacity here, incident there
  3. the age metric starts at the move
  4. an idle queue publishes nothing
  5. Maximum, not Average, over the period

basics

~20 s

Alarm on the DLQ's ApproximateNumberOfMessagesVisible at 1 or more — any message there is abnormal. ApproximateAgeOfOldestMessage on a DLQ measures time since the message was moved in, not since it was originally sent, so it tracks triage delay rather than true message age.

solid answer

~50 s

The primary signal is `ApproximateNumberOfMessagesVisible` on the dead-letter queue itself, alarmed at 1 or more. Unlike a work queue, where depth is a capacity question, any message in a DLQ means something failed repeatedly, so the threshold is presence rather than a tuned level. `ApproximateAgeOfOldestMessage` on the DLQ is useful but often misread: on a dead-letter queue it counts from when the message *moved into* the DLQ, not from when the producer sent it. That makes it a good measure of how long failures have gone untriaged, and a poor proxy for how much retention the message has left. On the *source* queue the same metric is the backlog and latency signal — rising age there means consumers are falling behind. One operational detail: CloudWatch treats a queue as active only for a period after its last message or API call, so a long-idle DLQ stops publishing datapoints; set the alarm's missing-data handling to `notBreaching` so it does not sit in `INSUFFICIENT_DATA`.

code

bash · 8 lines
bash
aws cloudwatch put-metric-alarm \
  --alarm-name orders-dlq-not-empty \
  --namespace AWS/SQS \
  --metric-name ApproximateNumberOfMessagesVisible \
  --dimensions Name=QueueName,Value=orders-dlq \
  --statistic Maximum --period 60 --evaluation-periods 1 \
  --threshold 1 --comparison-operator GreaterThanOrEqualToThreshold \
  --treat-missing-data notBreaching

go deeper

for a junior

Know the two SQS metrics by name — ApproximateNumberOfMessagesVisible for queue depth and ApproximateAgeOfOldestMessage for how long the oldest message has waited.

for a middle

Explain why a DLQ alarms on presence rather than a tuned level, and that on a DLQ the age metric counts from the move into the queue rather than from the original send.

for a senior

Make the alarm actually work in production: Maximum over a short period, threshold of one, missing data treated as notBreaching so an idle DLQ reads healthy instead of unknown.

for a principal

Own what the signal obliges. An arrival alarm on every DLQ is only meaningful if each queue has one owner and a defined triage path, and if the age metric's triage-delay reading is not mistaken for a durability guarantee.

## Why a DLQ alarm is different from a queue-depth alarm On a working queue, `ApproximateNumberOfMessagesVisible` is a capacity signal. Some depth is normal; you alarm on a level that means consumers are losing ground, and picking that level takes tuning. On a dead-letter queue the semantics invert. A message arrives there only after being delivered `maxReceiveCount` times without ever being deleted — it is, by construction, an abnormal event. So the useful threshold is not a level but a presence check: > `ApproximateNumberOfMessagesVisible >= 1` on the DLQ, over one period. That single alarm is the control that makes the whole dead-letter mechanism worth having. Without it, the DLQ is a bucket nobody looks in, and its messages quietly expire when retention elapses — silently, with no metric and no event marking their deletion. ## What the age metric actually measures here `ApproximateAgeOfOldestMessage` reports the age of the oldest non-deleted message in a queue, and its meaning depends on which queue you point it at. On the **source** queue it is the classic backlog and latency signal. Rising age means messages are waiting longer than they should — consumers are down, throttled, or under-scaled — and it is a far better SLO signal than raw depth, because depth conflates a brief burst with a genuine stall. On the **dead-letter** queue it means something narrower: AWS documents that for a DLQ the metric reflects when the message *moved into* the dead-letter queue, not when it was originally sent. Two consequences follow. First, it is a good triage-delay metric. "Oldest untriaged failure is six hours old" is a real, useful statement about your operational responsiveness, and a secondary alarm on it catches the case where the arrival alarm was acknowledged and then forgotten. Second, it is a bad proxy for remaining lifetime. A message's SQS retention is computed from its *original* enqueue timestamp and is not reset by the move, so the age metric systematically under-reports how close a message is to expiry — by exactly however long the source queue spent retrying it. An alarm that tries to catch messages before they expire by comparing this metric to the retention period will fire late. ## Supporting metrics A few others are worth naming, because interviewers often probe whether you know the difference: - `ApproximateNumberOfMessagesNotVisible` — messages currently in flight, received but not yet deleted. On a DLQ, a non-zero value means someone or something is actively consuming it. - `NumberOfMessagesSent` and `NumberOfMessagesDeleted` — throughput counters. Note that messages arriving in a DLQ by redrive are not producer sends, so do not treat the DLQ's send count as your dead-letter arrival rate; the visible-count and age metrics are the reliable indicators. - `NumberOfEmptyReceives` — receive calls that returned nothing, a cost and long-polling signal on the source queue. ## The idle-queue gotcha SQS publishes queue metrics to CloudWatch continuously while a queue is considered active, and CloudWatch treats a queue as active for a period after its last message or API activity. A healthy DLQ is, by definition, empty and untouched for long stretches — so it stops emitting datapoints, and an alarm evaluating against missing data lands in `INSUFFICIENT_DATA` rather than `OK`. The fix is to set the alarm's missing-data treatment to `notBreaching`, so an idle DLQ reads as healthy instead of unknown. Teams that skip this either get a permanently ambiguous alarm state or, worse, disable the alarm because it "never works". ```bash aws cloudwatch put-metric-alarm \ --alarm-name orders-dlq-not-empty \ --namespace AWS/SQS \ --metric-name ApproximateNumberOfMessagesVisible \ --dimensions Name=QueueName,Value=orders-dlq \ --statistic Maximum --period 60 --evaluation-periods 1 \ --threshold 1 --comparison-operator GreaterThanOrEqualToThreshold \ --treat-missing-data notBreaching ``` Use `Maximum` rather than `Average` for the statistic: a single message appearing within the period should breach, and averaging can dilute a brief spike below the threshold. ## Why "approximate" Standard queues store messages redundantly across many servers, so the counts are eventually consistent — a message just sent may not be reflected immediately, and a just-deleted one may still be counted. For a threshold of one over a one-minute period this is immaterial; it matters if you try to build precise reconciliation logic on top of these numbers, which you should not. For an exact count of what is in a DLQ, drain it and count what you receive.

  • Why use the Maximum statistic rather than Average for a DLQ presence alarm?
    Because you want a single message inside the evaluation period to breach. `Average` blends datapoints across the period, so one message appearing briefly can average below a threshold of one and never trigger. `Maximum` asks the question you actually mean: did the depth reach one at any point?
  • On a healthy source queue, what does a rising ApproximateAgeOfOldestMessage tell you that depth does not?
    That messages are genuinely waiting longer, not merely that a burst arrived. Depth spikes and recovers under normal traffic, but sustained rising age means consumers are not keeping up — down, throttled, under-scaled, or slowed by a dependency. It maps far more directly to an end-to-end latency objective than a count of queued messages does.
  • Your DLQ alarm has been in INSUFFICIENT_DATA for weeks. What is happening?
    The DLQ is empty and untouched, so CloudWatch no longer receives datapoints for it — SQS publishes metrics only while a queue is considered active. The alarm has no data to evaluate. Set the alarm's missing-data handling to `notBreaching` so an idle DLQ reads as OK, and you get a real state transition when a message finally arrives.

saying these in an interview costs you the question

  • Tunes a DLQ depth threshold as if it were a capacity signal
  • Reads ApproximateAgeOfOldestMessage on a DLQ as time since the original send
  • Uses the age metric to predict when messages will expire
  • Leaves the alarm in INSUFFICIENT_DATA because an idle queue emits nothing
  • Treats the approximate counts as exact for reconciliation

context