skip to content

A worker reading from an Amazon SQS queue takes about 90 seconds per message, and operators notice the same message being processed by several workers at once. What is happening, and how do you fix it?

level: seniorimportance: must knowfreq 62%

answer

  1. the lease expired mid-flight
  2. default is 30 seconds
  3. compare p99 handler time, not the mean
  4. renew instead of guessing
  5. the ceiling is 12 hours

basics

~20 s

The visibility timeout is shorter than the handler — with the 30-second default, the lease expires while the first worker is still running, so SQS makes the message visible again and hands it to another worker. Raise the timeout above worst-case processing time, or heartbeat with ChangeMessageVisibility.

solid answer

~50 s

This is the classic SQS bug. The visibility timeout is a lease, and a new queue leases for 30 seconds by default. A 90-second handler blows through that lease at the 30-second mark, at which point SQS considers the message unprocessed, makes it visible, and a second worker receives it — while the first is still working. Under load you get a pile-up, since each redelivery starts another 30-second clock. The fixes, in order of preference: size the queue's `VisibilityTimeout` off the p99 handler duration with headroom, not the average; for genuinely variable work, heartbeat by calling `ChangeMessageVisibility` periodically from the handler to renew the lease while progress continues; and break very long jobs into smaller units so no single lease has to cover hours. Also confirm the delete actually happens after processing — a handler that throws before `DeleteMessage` produces exactly the same symptom.

code

python · 25 lines
python
import threading
import boto3

sqs = boto3.client("sqs")

def process_with_heartbeat(queue_url, msg, work_fn, extend_to=120, every=30):
    done = threading.Event()

    def heartbeat():
        # renew the lease from *now*, but only while work is still running
        while not done.wait(timeout=every):
            sqs.change_message_visibility(
                QueueUrl=queue_url,
                ReceiptHandle=msg["ReceiptHandle"],
                VisibilityTimeout=extend_to,
            )

    t = threading.Thread(target=heartbeat, daemon=True)
    t.start()
    try:
        work_fn(msg["Body"])
        sqs.delete_message(QueueUrl=queue_url, ReceiptHandle=msg["ReceiptHandle"])
    finally:
        done.set()
        t.join()

go deeper

for a junior

Recall that the default lease is 30 seconds and that a handler running longer than the lease causes the message to be handed to another worker. Say that the timeout must exceed processing time.

for a middle

Explain the timeline of expiry and re-receive, name ChangeMessageVisibility as the renewal call, and describe how ApproximateReceiveCount confirms redelivery.

for a senior

Demonstrate the diagnosis: correlate p99 handler latency against the configured timeout, spot the amplification as each expiry adds another concurrent processor, and choose between resizing, heartbeating, and shrinking the unit of work.

for a principal

Own the tradeoff that the timeout is simultaneously the duplicate guard and the crash-recovery delay. Be ready to argue where long-running work belongs at all, and what the platform owes teams so rare redelivery stays harmless.

## The mechanism behind the symptom SQS hands a received message to one consumer under a lease called the **visibility timeout**. During the lease the message is invisible to other `ReceiveMessage` calls. When the lease expires without a `DeleteMessage`, SQS assumes the consumer died and makes the message visible again. There is no liveness check, no heartbeat by default, and no way for the queue to distinguish "still working" from "crashed": *elapsed time is the only signal it has*. So a 90-second handler against the default 30-second timeout produces this timeline: ``` t=0 worker A receives msg, lease until t=30 t=30 lease expires -> msg visible again (A is still running) t=31 worker B receives msg, lease until t=61 <-- duplicate #1 t=61 lease expires -> visible again t=62 worker C receives msg <-- duplicate #2 t=90 A finishes, calls DeleteMessage ``` The amplification is the nasty part: each expiry spawns another concurrent processor, so a queue that is merely slow can convert into a self-inflicted thundering herd. Meanwhile `ApproximateNumberOfMessagesNotVisible` looks healthy, because the message genuinely is in flight — just in flight three times. ## Diagnosing it quickly Three signals confirm it within minutes: - **Duplicate spacing matches the timeout.** Log the `MessageId` and `ApproximateReceiveCount` (request it via `AttributeNames`) on every receive. Duplicates arriving one visibility-timeout apart, with a climbing receive count, is diagnostic. - **Handler latency crosses the timeout.** Compare the p99 of your processing duration against the queue's `VisibilityTimeout`. If p99 > timeout, you have this bug for the tail of your traffic even if the mean looks fine. - **Deletes lag receives.** If the delete count trails the receive count by more than the duplicate rate you expect, work is finishing after its lease. ## Fix one: size the timeout from real latency Set the queue's `VisibilityTimeout` above the **worst-case** handler duration, not the average — the tail is exactly what redelivers. Include everything inside the lease: downstream API calls, retries, and their timeouts. If your HTTP client will spend up to 60 seconds retrying a flaky dependency, that 60 seconds is inside the lease. Don't overcorrect to the 12-hour maximum, though. The timeout is also your recovery time: if a worker really does die, its messages sit untouched for the whole remaining lease before anyone retries them. A 12-hour lease means a 12-hour stall on real failures. Pick the smallest value that comfortably covers the tail. ## Fix two: heartbeat the lease When processing time is genuinely unpredictable — a video transcode, a large report — a fixed timeout can't be right for both the fast and slow cases. Renew instead. `ChangeMessageVisibility` re-sets the lease *from now* for an in-flight message: ```python while not done.wait(timeout=30): # every 30s while working sqs.change_message_visibility( QueueUrl=queue_url, ReceiptHandle=receipt_handle, VisibilityTimeout=120, # another 2 minutes from now ) ``` Keep a modest base timeout and extend in steps. The critical property is that the heartbeat must stop when the work stops — if you renew from a background thread that outlives a hung handler, you have rebuilt the crash-invisible failure you were trying to avoid. Note also that total invisibility for one message is capped at 12 hours from the original receive; renewals cannot exceed that ceiling. ## Fix three: shrink the unit of work If a single message means an hour of work, the queue is being used as a job scheduler. Split it: enqueue coarse work as a small number of fine-grained messages, or hand the long-running part to a workflow orchestrator and keep the queue message as the trigger. Short leases are easier to reason about, fail faster, and retry more cheaply. ## The thing that is *not* the fix Watch for the neighbouring bug that presents identically: a handler that throws — or exits — before reaching `DeleteMessage`. Then the message legitimately never got acknowledged, and redelivery is SQS doing its job. Distinguish them by checking whether the first processor *completed*. If it did and the message still came back, it is a timeout problem; if it did not, it is an error-handling problem. Finally, treat redelivery as normal rather than eradicable. Even a perfectly sized timeout can redeliver — a worker can die a millisecond before its delete. Sizing and heartbeating reduce the rate to something rare; making the effect of reprocessing harmless is a separate design concern that consumers must handle regardless.

  • Why not just set the visibility timeout to its 12-hour maximum and stop worrying?
    Because the timeout is also your recovery time. If a worker dies holding a 12-hour lease, that message is invisible for the rest of the 12 hours before anyone retries it — a small crash becomes a half-day stall. Pick the smallest value that covers your latency tail, and heartbeat when the tail is unpredictable.
  • Which received message attribute helps you confirm redelivery, and how do you get it?
    `ApproximateReceiveCount` — request it on `ReceiveMessage` via `AttributeNames` and log it with the MessageId. A value climbing past 1 on messages your workers are completing successfully points straight at lease expiry rather than handler errors.
  • What is the danger of running the heartbeat from a thread that is independent of the work?
    It can keep renewing the lease after the work has hung or died, so SQS never redelivers and the message is silently stuck until the 12-hour ceiling. The heartbeat must be tied to observable progress and must stop the moment the handler stops — otherwise you have traded duplicates for a stall.
  • Two workers report they both processed the same message. How do you tell a timeout problem from an error-handling problem?
    Check whether the first processor completed. If it finished the work and the message still came back, the lease expired mid-flight — a timeout problem. If it threw or exited before calling DeleteMessage, redelivery is correct behaviour and the bug is in the handler's error path.

saying these in an interview costs you the question

  • Sizing the visibility timeout from average rather than tail latency
  • Setting the maximum 12-hour timeout as a blanket fix
  • Believing a correct timeout eliminates redelivery entirely
  • Heartbeating from a thread that outlives the hung handler
  • Assuming duplicates mean SQS is broken rather than the lease lapsed

context