skip to content

After deploying a fix, how do you replay the messages sitting in an Amazon SQS dead-letter queue back to the queue they came from, and what should you verify before you start?

level: seniorimportance: should knowfreq 52%

answer

  1. SQS does the shuttling for you now
  2. the DLQ is the source, not the destination
  3. omit the destination and it goes home
  4. choose the rate on purpose
  5. have the abort command ready first

basics

~20 s

Use the SQS message move task: call StartMessageMoveTask with the DLQ's ARN as the source, optionally throttling with MaxNumberOfMessagesPerSecond. Before starting, confirm the fix is actually deployed and consumers are healthy, or the messages will fail again and return straight to the DLQ.

solid answer

~40 s

SQS has a first-class redrive operation, so you do not write a shuttle consumer. `StartMessageMoveTask` takes the dead-letter queue as `SourceArn`; if you omit `DestinationArn` the messages go back to the queues they originally came from. `MaxNumberOfMessagesPerSecond` throttles the replay so you do not re-flood a downstream that is still fragile. You track progress with `ListMessageMoveTasks` and stop a bad replay with `CancelMessageMoveTask`; only one task can be active per source queue at a time. The pre-flight matters more than the API. Confirm the fixed code is deployed and receiving traffic, confirm consumers are healthy and scaled, and confirm the failure was actually transient — replaying genuinely malformed messages just walks them back through `maxReceiveCount` into the DLQ, burning retention budget. When unsure, receive a few messages manually and inspect them first.

code

bash · 6 lines
bash
aws sqs start-message-move-task \
  --source-arn arn:aws:sqs:us-east-1:111122223333:orders-dlq \
  --max-number-of-messages-per-second 10

aws sqs list-message-move-tasks \
  --source-arn arn:aws:sqs:us-east-1:111122223333:orders-dlq

go deeper

for a junior

Know that SQS can move messages off a dead-letter queue for you with StartMessageMoveTask, and that by default they return to the queue they came from.

for a middle

Explain the parameters — SourceArn must be the DLQ, DestinationArn is optional, MaxNumberOfMessagesPerSecond throttles — and that the move deletes each message from the DLQ as it is sent.

for a senior

Demonstrate the pre-flight: fix deployed and serving, failure genuinely transient, consumers scaled, handler idempotent, and a live signal plus CancelMessageMoveTask ready before you start.

for a principal

Own replay as a controlled production change: who authorises it, what rate is safe for shared downstreams, when messages should be diverted to a repair path instead, and when stale work should be dropped rather than reprocessed.

## The operation SQS gives you Before 2023 the standard answer to "how do you replay a DLQ" was "write a Lambda that receives from the DLQ and sends to the source queue, and be careful with idempotency". SQS now ships the operation natively as a **message move task**, exposed through three APIs: - `StartMessageMoveTask` — begins moving messages off a dead-letter queue. - `ListMessageMoveTasks` — reports status and progress for a source queue's tasks. - `CancelMessageMoveTask` — stops one in flight. The parameters that matter: ```bash aws sqs start-message-move-task \ --source-arn arn:aws:sqs:us-east-1:111122223333:orders-dlq \ --max-number-of-messages-per-second 10 ``` `SourceArn` is required and must be a dead-letter queue — specifically a DLQ whose sources are other SQS queues. A queue that only collects failures from other services is not a valid source for the task. `DestinationArn` is optional. Omit it and messages return to the queue or queues they originally came from, which is the common case. Supply it to divert the replay somewhere else — a quarantine queue, or a separate queue consumed by a repair worker — when you want the messages inspected or transformed rather than reprocessed as-is. `MaxNumberOfMessagesPerSecond` is optional and is the parameter experienced operators reach for. Without it the task moves as fast as it can, and if the DLQ holds a large backlog you have just converted a quiet incident into a thundering herd against whatever downstream was already struggling. Throttling to a rate your consumers and downstream dependencies can absorb turns the replay into a controlled drain. One task may be active per source queue at a time, so replay is a serialised operation you can reason about, not something a script can accidentally start five copies of. ## The move is destructive to the DLQ Messages are moved, not copied. As each one is sent to the destination it is removed from the dead-letter queue. If the replay fails downstream, your safety net is now the source queue's own redrive policy sending them back — which works, but each round trip consumes retention budget measured from the original enqueue time. This is why the pre-flight is the substance of the answer. ## The pre-flight **Is the fix actually live?** Not merged, not built — deployed, serving, and receiving traffic. Replaying into the old code produces exactly the same failures plus a second pass through `maxReceiveCount`. **Was the failure transient or structural?** Retry-safe causes — a downstream outage, a throttled dependency, a bad deploy since rolled back — replay cleanly. Malformed payloads, references to entities that no longer exist, or messages a schema change made unparseable will fail again no matter how healthy the consumer is. Receiving a handful with `ReceiveMessage` and reading them is a two-minute check that saves an hour. **Are the consumers healthy and scaled?** A replay lands as a burst. If the consumer fleet is sized for steady state, throttle the task or scale out first. **Is reprocessing safe?** Some of these messages may have been partially processed before failing — a row written, an email sent, a charge attempted. Whether replay is safe depends on the handler's idempotency, and if it is not idempotent, the honest options are to make it so, to route the replay to a repair consumer via `DestinationArn`, or to triage by hand. **Is the message still meaningful?** A message that has been sitting for a week may reference a state that no longer exists, or represent a decision that has since been superseded. Replaying stale work can be worse than dropping it. ## Watching it and stopping it ```bash aws sqs list-message-move-tasks \ --source-arn arn:aws:sqs:us-east-1:111122223333:orders-dlq ``` The result reports the task's status along with how many messages have moved and how many failed. Watch the source queue's own error rate and the DLQ depth at the same time: if messages start arriving back in the DLQ while the task runs, the fix did not work, and `CancelMessageMoveTask` stops the bleeding immediately. ## The framing that reads as senior The API is the easy half. What distinguishes a strong answer is treating replay as a change to production: verify the fix, choose a rate, watch a signal that tells you it is working, and know the abort command before you start. Candidates who answer only "there's a redrive button in the console" have described the mechanism and skipped the operation.

  • What does omitting DestinationArn do, and when would you set it explicitly?
    Omitting it returns messages to the queue they originally came from, which is what you want after a straightforward fix. Set it explicitly when the messages should not be reprocessed as-is — divert them to a quarantine queue for inspection, or to a repair consumer that transforms them before they re-enter the normal path. It is also how you replay into a different queue after a migration.
  • Why is MaxNumberOfMessagesPerSecond more than a nicety?
    Because a DLQ can hold a large backlog and an unthrottled move delivers it as a burst. That is a thundering herd aimed at a downstream that was, minutes ago, the thing that failed. Setting the rate to something the consumers and their dependencies absorb turns the replay into a drain you can watch, and keeps a failed replay from causing a second incident.
  • Halfway through a replay, messages start reappearing in the DLQ. What do you do?
    Cancel immediately with `CancelMessageMoveTask` — the fix did not work, and every further message moved is one that will burn a second pass through maxReceiveCount and more of its retention budget. Then pull a sample from the DLQ and read the failures: the cause is usually either that the deploy did not reach the consumers or that the failure was structural rather than transient.

saying these in an interview costs you the question

  • Says you must write a custom Lambda to shuttle messages back
  • Replays without confirming the fix is deployed and serving
  • Ignores throttling and floods a still-fragile downstream
  • Assumes replay copies messages and leaves the DLQ intact
  • Treats replay as safe regardless of handler idempotency

context