skip to content

In Amazon SQS, what does the visibility timeout do after a consumer calls ReceiveMessage, and why must the consumer still call DeleteMessage?

level: juniorimportance: must knowfreq 82%

answer

  1. hidden, not removed
  2. a lease, not a lock
  3. expiry means redelivery
  4. the delete is the acknowledgement
  5. receipt handle, not message ID

basics

~20 s

ReceiveMessage hides a message for the visibility timeout instead of removing it, giving one consumer a temporary exclusive lease. Only DeleteMessage removes it; if the lease expires first, the message becomes visible again and is delivered to another consumer.

solid answer

~50 s

SQS is a pull system with an explicit acknowledgement. `ReceiveMessage` returns the message plus a receipt handle, and for the length of the visibility timeout — 30 seconds by default, up to 12 hours — that message is hidden from every other `ReceiveMessage` call on the queue. It is not removed: the timeout is a lease, not a lock. The consumer does the work and then calls `DeleteMessage` with that receipt handle, which is the only thing that takes the message off the queue. If the consumer crashes, hangs, or simply forgets to delete, the lease expires, the message becomes visible again, and another consumer picks it up. That is precisely how SQS survives a dead worker without losing work — and also why redelivery is a normal event rather than an error. If the work legitimately needs longer, extend the lease with `ChangeMessageVisibility`.

code

python · 18 lines
python
import boto3

sqs = boto3.client("sqs")
queue_url = sqs.get_queue_url(QueueName="orders")["QueueUrl"]

resp = sqs.receive_message(
    QueueUrl=queue_url,
    MaxNumberOfMessages=1,
    VisibilityTimeout=120,  # per-receive override of the queue default
)

for msg in resp.get("Messages", []):
    print("processing", msg["MessageId"])
    # ... do the work here; the message is invisible for 120 seconds ...
    sqs.delete_message(
        QueueUrl=queue_url,
        ReceiptHandle=msg["ReceiptHandle"],  # not MessageId
    )

go deeper

for a junior

Be able to name the three calls in order — ReceiveMessage, do the work, DeleteMessage — and say plainly that receiving only hides a message and deleting is what removes it.

for a middle

Explain the lease mechanics: the 30-second default, the per-receive override, the receipt handle versus the message ID, and why deleting after processing gives at-least-once while deleting first loses messages.

for a senior

Show you operate this: point at ApproximateNumberOfMessagesNotVisible to spot consumers that receive but never delete, and size the timeout from real handler latency rather than accepting the default.

for a principal

Own the tradeoff the lease encodes — SQS chose duplicate work over lost work. Be ready to argue where that default belongs in a platform, and what teams must build downstream because redelivery is normal.

## The queue never pushes, and it never assumes success Amazon SQS does not deliver messages to consumers; consumers ask for them. The whole consumption model is three calls — `ReceiveMessage`, then your work, then `DeleteMessage` — and the visibility timeout is the safety mechanism that holds the middle step together. When you call `ReceiveMessage`, SQS does **not** remove the message. It returns a copy of the body along with a **receipt handle** — an opaque token identifying *this particular receive* of *this particular message* — and starts a clock. Until that clock expires, the message is *invisible*: no other `ReceiveMessage` call against the queue, from any consumer, will return it. That window is the **visibility timeout**. ## A lease, not a lock The distinction matters because it explains every behaviour that follows. A lock is held until released; a lease expires on its own. SQS deliberately chose a lease, because the queue has no way to know whether a silent consumer is busy or dead. If the consumer never comes back, the safe assumption is that the work did not happen, so the message becomes visible again and is handed to whoever asks next. That is why SQS gives you *at-least-once* delivery on standard queues: a message may be processed more than once whenever a lease expires before the delete lands. The queue is trading duplicate work for never silently losing work. ## The delete is the acknowledgement Nothing about a successful `ReceiveMessage` implies success of processing, so nothing about it deletes the message. Only `DeleteMessage` does — and it takes the *receipt handle*, not the message ID: ```python import boto3 sqs = boto3.client("sqs") resp = sqs.receive_message(QueueUrl=QUEUE_URL, MaxNumberOfMessages=1) for msg in resp.get("Messages", []): handle_work(msg["Body"]) # only after this succeeds sqs.delete_message(QueueUrl=QUEUE_URL, ReceiptHandle=msg["ReceiptHandle"]) ``` A message ID is stable for the life of the message; a receipt handle is issued fresh on every receive. If a message is received twice, you get two different handles, and you should always delete with the most recently received one — an older handle is not guaranteed to delete the message. Ordering matters too: deleting *before* the work finishes converts the queue to at-most-once and loses messages on a crash. Delete after. ## Defaults, ranges and overrides The queue attribute `VisibilityTimeout` sets the default for every receive, is 30 seconds on a new queue, and can be set from 0 seconds to 12 hours. `ReceiveMessage` also accepts a per-call `VisibilityTimeout` that overrides the queue default for just those messages, which is useful when one consumer handles a much slower class of work than the others. And `ChangeMessageVisibility` re-sets the timeout for an in-flight message from *now*, so a long job can renew its lease periodically; passing `0` does the opposite and releases the message immediately, which is the standard way to hand back work you have decided not to do. ## Watching it in production Two CloudWatch metrics tell the story: `ApproximateNumberOfMessagesVisible` is the backlog waiting to be received, and `ApproximateNumberOfMessagesNotVisible` is the count currently leased — messages received but neither deleted nor expired. A steadily climbing not-visible count with a flat visible count usually means consumers are receiving work and failing to delete it, which is a bug, not a backlog. ## The failure modes this design creates - **Timeout shorter than the handler.** The lease expires mid-flight and a second consumer starts the same work while the first is still running. This is the single most common SQS bug. - **Delete never called.** Every message is redelivered forever until it ages out or is routed away by the queue's redrive configuration. - **Deleting first, processing second.** A crash now loses the message permanently. - **Assuming a huge timeout is free.** A 12-hour lease means a crashed consumer's messages sit untouched for 12 hours before anyone retries them. The mental model to carry into an interview: SQS hands out a temporary, revocable lease and waits for an explicit acknowledgement. Everything else — heartbeating, batching, redelivery — follows from that one sentence.

  • What exactly is a receipt handle, and why can't you delete a message by its MessageId?
    A receipt handle identifies one specific *receive* of a message, so SQS can tell which lease you are acting on; the MessageId identifies the message for its whole life. Every receive issues a new handle, and `DeleteMessage`/`ChangeMessageVisibility` only accept handles. Always use the most recently received one — an older handle may fail to delete.
  • Which CloudWatch metric tells you how many messages are currently leased, and what does a rising value mean?
    `ApproximateNumberOfMessagesNotVisible` counts in-flight messages — received but not yet deleted or expired. A rising in-flight count alongside a flat visible count normally means consumers are receiving work and failing to delete it: handlers erroring out, crashing, or running far longer than expected.
  • What happens if you call ChangeMessageVisibility with a VisibilityTimeout of 0?
    The message becomes visible immediately and the next `ReceiveMessage` can return it. That is the clean way to hand back work you have decided not to process — a shutting-down worker, or a message you want retried right away — instead of holding the lease until it expires.

saying these in an interview costs you the question

  • Thinks ReceiveMessage removes the message from the queue
  • Believes SQS deletes automatically once processing returns
  • Deletes using the MessageId instead of the receipt handle
  • Deletes the message before doing the work
  • Calls the visibility timeout a lock other consumers cannot break

context