skip to content

SQS

You will learn Amazon SQS as a work queue: which queue type to pick, how a message is leased and acknowledged, and what happens to poison messages. Interviewers treat SQS as the default decoupling answer, so they expect you to reason about visibility timeout and redrive rather than just 'put a queue in front of it'.

part ofAWSoverview, primer and where to startread it →
on this pageshow

explore

questions

16

In Amazon SQS, what does the visibility timeout do after a consumer calls ReceiveMessage, and why must the consumer still call DeleteMessage?

level: juniorimportance: must knowfreq 82%

answer

  1. hidden, not removed
  2. a lease, not a lock
  3. expiry means redelivery
  4. the delete is the acknowledgement
  5. receipt handle, not message ID

basics

~20 s

ReceiveMessage hides a message for the visibility timeout instead of removing it, giving one consumer a temporary exclusive lease. Only DeleteMessage removes it; if the lease expires first, the message becomes visible again and is delivered to another consumer.

solid answer

~50 s

SQS is a pull system with an explicit acknowledgement. `ReceiveMessage` returns the message plus a receipt handle, and for the length of the visibility timeout — 30 seconds by default, up to 12 hours — that message is hidden from every other `ReceiveMessage` call on the queue. It is not removed: the timeout is a lease, not a lock. The consumer does the work and then calls `DeleteMessage` with that receipt handle, which is the only thing that takes the message off the queue. If the consumer crashes, hangs, or simply forgets to delete, the lease expires, the message becomes visible again, and another consumer picks it up. That is precisely how SQS survives a dead worker without losing work — and also why redelivery is a normal event rather than an error. If the work legitimately needs longer, extend the lease with `ChangeMessageVisibility`.

code

python · 18 lines
python
import boto3

sqs = boto3.client("sqs")
queue_url = sqs.get_queue_url(QueueName="orders")["QueueUrl"]

resp = sqs.receive_message(
    QueueUrl=queue_url,
    MaxNumberOfMessages=1,
    VisibilityTimeout=120,  # per-receive override of the queue default
)

for msg in resp.get("Messages", []):
    print("processing", msg["MessageId"])
    # ... do the work here; the message is invisible for 120 seconds ...
    sqs.delete_message(
        QueueUrl=queue_url,
        ReceiptHandle=msg["ReceiptHandle"],  # not MessageId
    )

go deeper

for a junior

Be able to name the three calls in order — ReceiveMessage, do the work, DeleteMessage — and say plainly that receiving only hides a message and deleting is what removes it.

for a middle

Explain the lease mechanics: the 30-second default, the per-receive override, the receipt handle versus the message ID, and why deleting after processing gives at-least-once while deleting first loses messages.

for a senior

Show you operate this: point at ApproximateNumberOfMessagesNotVisible to spot consumers that receive but never delete, and size the timeout from real handler latency rather than accepting the default.

for a principal

Own the tradeoff the lease encodes — SQS chose duplicate work over lost work. Be ready to argue where that default belongs in a platform, and what teams must build downstream because redelivery is normal.

## The queue never pushes, and it never assumes success Amazon SQS does not deliver messages to consumers; consumers ask for them. The whole consumption model is three calls — `ReceiveMessage`, then your work, then `DeleteMessage` — and the visibility timeout is the safety mechanism that holds the middle step together. When you call `ReceiveMessage`, SQS does **not** remove the message. It returns a copy of the body along with a **receipt handle** — an opaque token identifying *this particular receive* of *this particular message* — and starts a clock. Until that clock expires, the message is *invisible*: no other `ReceiveMessage` call against the queue, from any consumer, will return it. That window is the **visibility timeout**. ## A lease, not a lock The distinction matters because it explains every behaviour that follows. A lock is held until released; a lease expires on its own. SQS deliberately chose a lease, because the queue has no way to know whether a silent consumer is busy or dead. If the consumer never comes back, the safe assumption is that the work did not happen, so the message becomes visible again and is handed to whoever asks next. That is why SQS gives you *at-least-once* delivery on standard queues: a message may be processed more than once whenever a lease expires before the delete lands. The queue is trading duplicate work for never silently losing work. ## The delete is the acknowledgement Nothing about a successful `ReceiveMessage` implies success of processing, so nothing about it deletes the message. Only `DeleteMessage` does — and it takes the *receipt handle*, not the message ID: ```python import boto3 sqs = boto3.client("sqs") resp = sqs.receive_message(QueueUrl=QUEUE_URL, MaxNumberOfMessages=1) for msg in resp.get("Messages", []): handle_work(msg["Body"]) # only after this succeeds sqs.delete_message(QueueUrl=QUEUE_URL, ReceiptHandle=msg["ReceiptHandle"]) ``` A message ID is stable for the life of the message; a receipt handle is issued fresh on every receive. If a message is received twice, you get two different handles, and you should always delete with the most recently received one — an older handle is not guaranteed to delete the message. Ordering matters too: deleting *before* the work finishes converts the queue to at-most-once and loses messages on a crash. Delete after. ## Defaults, ranges and overrides The queue attribute `VisibilityTimeout` sets the default for every receive, is 30 seconds on a new queue, and can be set from 0 seconds to 12 hours. `ReceiveMessage` also accepts a per-call `VisibilityTimeout` that overrides the queue default for just those messages, which is useful when one consumer handles a much slower class of work than the others. And `ChangeMessageVisibility` re-sets the timeout for an in-flight message from *now*, so a long job can renew its lease periodically; passing `0` does the opposite and releases the message immediately, which is the standard way to hand back work you have decided not to do. ## Watching it in production Two CloudWatch metrics tell the story: `ApproximateNumberOfMessagesVisible` is the backlog waiting to be received, and `ApproximateNumberOfMessagesNotVisible` is the count currently leased — messages received but neither deleted nor expired. A steadily climbing not-visible count with a flat visible count usually means consumers are receiving work and failing to delete it, which is a bug, not a backlog. ## The failure modes this design creates - **Timeout shorter than the handler.** The lease expires mid-flight and a second consumer starts the same work while the first is still running. This is the single most common SQS bug. - **Delete never called.** Every message is redelivered forever until it ages out or is routed away by the queue's redrive configuration. - **Deleting first, processing second.** A crash now loses the message permanently. - **Assuming a huge timeout is free.** A 12-hour lease means a crashed consumer's messages sit untouched for 12 hours before anyone retries them. The mental model to carry into an interview: SQS hands out a temporary, revocable lease and waits for an explicit acknowledgement. Everything else — heartbeating, batching, redelivery — follows from that one sentence.

  • What exactly is a receipt handle, and why can't you delete a message by its MessageId?
    A receipt handle identifies one specific *receive* of a message, so SQS can tell which lease you are acting on; the MessageId identifies the message for its whole life. Every receive issues a new handle, and `DeleteMessage`/`ChangeMessageVisibility` only accept handles. Always use the most recently received one — an older handle may fail to delete.
  • Which CloudWatch metric tells you how many messages are currently leased, and what does a rising value mean?
    `ApproximateNumberOfMessagesNotVisible` counts in-flight messages — received but not yet deleted or expired. A rising in-flight count alongside a flat visible count normally means consumers are receiving work and failing to delete it: handlers erroring out, crashing, or running far longer than expected.
  • What happens if you call ChangeMessageVisibility with a VisibilityTimeout of 0?
    The message becomes visible immediately and the next `ReceiveMessage` can return it. That is the clean way to hand back work you have decided not to process — a shutting-down worker, or a message you want retried right away — instead of holding the lease until it expires.

saying these in an interview costs you the question

  • Thinks ReceiveMessage removes the message from the queue
  • Believes SQS deletes automatically once processing returns
  • Deletes using the MessageId instead of the receipt handle
  • Deletes the message before doing the work
  • Calls the visibility timeout a lock other consumers cannot break

context

open as a page

In Amazon SQS, what are the differences between a standard queue and a FIFO queue, and when would you choose each?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Standard SQS queues give near-unlimited throughput with best-effort ordering and possible duplicate deliveries. FIFO queues preserve order within each MessageGroupId and deduplicate messages, but cap throughput. Choose FIFO only when ordering or deduplication is genuinely required.

open as a page

In an Amazon SQS redrive policy, what does maxReceiveCount actually count, and what has to happen before a message lands in the dead-letter queue?

level: middleimportance: must knowfreq 78%

basics

~20 s

maxReceiveCount caps how many times SQS may deliver one message. Every ReceiveMessage delivery increments that message's ApproximateReceiveCount — whether the consumer failed, crashed, or never answered — and once the count passes the limit, SQS routes the message to the dead-letter queue.

open as a page

In an SQS FIFO queue, what does MessageGroupId control, and how does the choice of group key affect consumer parallelism?

level: middleimportance: must knowfreq 64%

basics

~20 s

MessageGroupId is the ordering scope of a FIFO queue: SQS delivers messages sharing a group id strictly in order, one in flight at a time, while different groups proceed independently. The number of distinct group ids therefore sets the maximum consumer parallelism.

open as a page

A worker reading from an Amazon SQS queue takes about 90 seconds per message, and operators notice the same message being processed by several workers at once. What is happening, and how do you fix it?

level: seniorimportance: must knowfreq 62%

basics

~20 s

The visibility timeout is shorter than the handler — with the 30-second default, the lease expires while the first worker is still running, so SQS makes the message visible again and hands it to another worker. Raise the timeout above worst-case processing time, or heartbeat with ChangeMessageVisibility.

open as a page

You set an Amazon SQS FIFO queue's redrive policy to point at a standard queue as its dead-letter queue. What happens, and what are SQS's rules for what a DLQ may be?

level: juniorimportance: should knowfreq 42%

basics

~20 s

SQS rejects the configuration. A dead-letter queue must be the same type as its source queue — FIFO for FIFO, standard for standard — and must live in the same AWS account and Region. Otherwise it is an ordinary queue you create yourself.

open as a page

A Lambda function triggered by an Amazon SQS queue receives a batch of 10 messages and one of them fails. By default what happens to the other nine, and how does the ReportBatchItemFailures setting change it?

level: middleimportance: should knowfreq 46%

basics

~20 s

By default a failing Lambda invocation deletes nothing, so all ten messages become visible again after the visibility timeout and the nine successes are reprocessed. Enabling ReportBatchItemFailures lets the handler return only the failed message IDs, and Lambda deletes the rest.

open as a page

In Amazon SQS, what is the difference between short polling and long polling on ReceiveMessage, and which settings control which one you get?

level: middleimportance: should knowfreq 66%

basics

~20 s

Short polling samples a subset of SQS's servers and returns immediately, often empty even when messages exist. Long polling waits up to WaitTimeSeconds (max 20) for a message to arrive, cutting empty responses, request cost, and latency. Zero means short polling.

open as a page

Which CloudWatch metrics tell you that an Amazon SQS dead-letter queue is filling up, and what does ApproximateAgeOfOldestMessage measure on a DLQ specifically?

level: middleimportance: should knowfreq 44%

basics

~20 s

Alarm on the DLQ's ApproximateNumberOfMessagesVisible at 1 or more — any message there is abnormal. ApproximateAgeOfOldestMessage on a DLQ measures time since the message was moved in, not since it was originally sent, so it tracks triage delay rather than true message age.

open as a page

How does deduplication work on an SQS FIFO queue, and why does it not give you end-to-end exactly-once processing?

level: middleimportance: should knowfreq 52%

basics

~20 s

An SQS FIFO queue ignores a send whose MessageDeduplicationId it has already accepted within the previous five minutes, using either an explicit id or a SHA-256 hash of the body when ContentBasedDeduplication is on. It deduplicates sends only, so a consumer crash after processing still causes redelivery.

open as a page

After deploying a fix, how do you replay the messages sitting in an Amazon SQS dead-letter queue back to the queue they came from, and what should you verify before you start?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Use the SQS message move task: call StartMessageMoveTask with the DLQ's ARN as the source, optionally throttling with MaxNumberOfMessagesPerSecond. Before starting, confirm the fix is actually deployed and consumers are healthy, or the messages will fail again and return straight to the DLQ.

open as a page

Failed orders had been landing in an Amazon SQS dead-letter queue for days, but when the team finally opened it, most were gone — and nothing was consuming that DLQ. What explains it, and how do you prevent it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

They expired. A message's SQS retention clock starts at its original enqueue time and is not reset when the message moves to a dead-letter queue, so time already spent failing on the source queue is deducted from its life in the DLQ. Raise the DLQ's MessageRetentionPeriod and alarm on arrivals.

open as a page

A team asks to convert a high-volume standard SQS queue to FIFO because production shows duplicate and out-of-order messages. How do you evaluate that request?

level: principalimportance: should knowfreq 38%

basics

~20 s

Separate the two complaints first: duplicates are inherent to at-least-once delivery and are answered by idempotent consumers, not by FIFO. Only genuine order dependence justifies FIFO, and it costs a new queue, a throughput ceiling, and head-of-line blocking per group.

open as a page

Amazon SQS's DeleteMessageBatch can return HTTP 200 while some entries failed. Why does it work that way, and what must a consumer do about it?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

SQS batch APIs are per-entry, not atomic: the call succeeds while individual entries can fail, so the response splits into Successful and Failed lists. A consumer must inspect Failed and retry those deletes, or the messages reappear after the visibility timeout and get processed again.

open as a page

What is the RedriveAllowPolicy attribute on an Amazon SQS queue, and what does its redrivePermission setting control?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

RedriveAllowPolicy is set on a queue acting as a dead-letter queue and declares which source queues may name it as their DLQ. Its redrivePermission field takes allowAll, denyAll, or byQueue with an explicit list of source queue ARNs.

open as a page

What limits throughput on a default SQS FIFO queue, and what do the DeduplicationScope and FifoThroughputLimit attributes change?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

A default FIFO queue is quota-limited to 300 API calls per second per action, roughly 3,000 messages per second when batching ten per call. Setting DeduplicationScope to messageGroup and FifoThroughputLimit to perMessageGroupId enables high throughput mode, which lifts the ceiling but narrows deduplication to each group.

open as a page