In Amazon SQS, what is the difference between short polling and long polling on ReceiveMessage, and which settings control which one you get?
answer
- empty response from a non-empty queue
- sampling a subset of servers
- waits for arrival, up to 20 seconds
- zero means short polling
- fewer requests, and lower latency too
basics
~20 sShort polling samples a subset of SQS's servers and returns immediately, often empty even when messages exist. Long polling waits up to WaitTimeSeconds (max 20) for a message to arrive, cutting empty responses, request cost, and latency. Zero means short polling.
solid answer
~50 sSQS stores a queue across many servers. A short poll queries only a **subset** of them and answers straight away, so it can return zero messages while messages are sitting on the servers it did not ask — you then poll again in a tight loop, paying for every empty request. A long poll waits: SQS holds the connection open until at least one message is available or `WaitTimeSeconds` elapses, and it queries all servers, so an empty response really does mean an empty queue. You get long polling either per-call, by passing `WaitTimeSeconds` on `ReceiveMessage` (1–20 seconds), or per-queue, by setting the `ReceiveMessageWaitTimeSeconds` queue attribute so every receive inherits it. The default is 0, which is short polling. Long polling is the sane default for almost every worker: fewer API calls, lower bill, and *lower* latency, because a message arriving mid-wait is returned immediately rather than on your next poll.
code
bash · 12 linesQUEUE_URL=$(aws sqs get-queue-url --queue-name orders --query QueueUrl --output text)
# Make long polling the queue-wide default (max 20 seconds)
aws sqs set-queue-attributes \
--queue-url "$QUEUE_URL" \
--attributes ReceiveMessageWaitTimeSeconds=20
# Per-call override; returns as soon as one message is available
aws sqs receive-message \
--queue-url "$QUEUE_URL" \
--max-number-of-messages 10 \
--wait-time-seconds 20go deeper
Know that the default receive can come back empty even when the queue has messages, and that setting a wait time of up to 20 seconds fixes it.
Explain the sampling versus full-fleet difference, name both controls — the WaitTimeSeconds parameter and the ReceiveMessageWaitTimeSeconds queue attribute — and state that 0 means short polling.
Argue the operational case: empty receives are billable requests, long polling lowers rather than raises latency, and enforcing it at the queue prevents a hot-looping consumer from inflating the bill.
Frame it as a platform default. Decide whether polling workers belong in your architecture at all versus event-driven invocation, and how you would detect and stop cost-burning poll loops across many teams.
## Why short polling can lie to you An SQS queue is not one box. Messages are stored redundantly across many servers, and a `ReceiveMessage` request has to decide how much of that fleet to consult before answering. **Short polling** (the default: `WaitTimeSeconds = 0`) consults a *sample* of the servers and returns whatever it finds, immediately. If your queue holds five messages and the sampled servers happen not to hold any of them, you get an empty response — from a queue that is demonstrably not empty. Poll again a moment later and you probably get one. On a low-traffic queue, this is normal and surprising the first time you see it. **Long polling** (`WaitTimeSeconds` between 1 and 20) changes both halves of that behaviour: SQS queries all the servers, and rather than answering instantly it holds the request open until either at least one message becomes available or the wait time runs out. An empty long-poll response is therefore meaningful — for that window, the queue really had nothing. ## Turning it on, in two places ```bash # Per call: this receive waits up to 20s aws sqs receive-message \ --queue-url "$QUEUE_URL" \ --wait-time-seconds 20 \ --max-number-of-messages 10 # Per queue: every receive inherits long polling aws sqs set-queue-attributes \ --queue-url "$QUEUE_URL" \ --attributes ReceiveMessageWaitTimeSeconds=20 ``` The per-call `WaitTimeSeconds` parameter overrides the queue attribute, including overriding it *down* to 0 to force a short poll. Twenty seconds is the maximum for both. A common production pattern is to set `ReceiveMessageWaitTimeSeconds` on the queue so nothing can accidentally hot-loop, and let application code stay ignorant of it. ## The three things you actually gain **Cost.** SQS bills per request, and empty receives are requests. A worker short-polling in a loop against an idle queue can generate millions of billable, useless calls a month. Long polling with a 20-second wait caps an idle consumer at roughly three requests a minute. **Latency.** This is the counter-intuitive one, and the one interviewers like. Waiting does not make you slower — it makes you faster. With short polling plus a sleep between attempts, a message that arrives one millisecond after your poll returns waits for your next sleep to finish. With long polling, that same message is returned to the already-open request as soon as it lands. **Correctness of signal.** Code that treats "empty response" as "queue drained" — a batch job that exits, a scale-in decision — is simply wrong under short polling. Under long polling it is defensible. ## The gotchas worth naming - **Client read timeouts.** Your HTTP/SDK socket read timeout must exceed `WaitTimeSeconds`, or the client aborts the connection mid-wait and you see spurious timeout errors. The AWS SDKs generally account for this, but hand-rolled clients and aggressive custom timeouts trip on it. - **Long polling does not fill the batch.** `ReceiveMessage` returns as soon as *at least one* message is available. Asking for `MaxNumberOfMessages=10` with a 20-second wait will still commonly return one or two messages — it does not linger to accumulate ten. - **It is not a subscription.** Long polling is still polling: one request, one response, up to 20 seconds. A consumer loop re-issues the call. If you need genuine push semantics, that is a different integration choice, not a polling tune-up. - **Zero is a real value.** Because 0 means short polling, an SDK wrapper that defaults an unset field to 0 silently turns long polling off. Set it explicitly on the queue if you care. ## The one-line answer to carry in Short polling asks some of the servers and answers now; long polling asks all of them and answers when there is something to say. Outside of a deliberate low-latency drain loop with constant traffic, long polling is what you want, and setting it on the queue is how you make sure you keep it.
- If you request MaxNumberOfMessages=10 with WaitTimeSeconds=20, will SQS wait the full 20 seconds to gather ten messages?No. The call returns as soon as at least one message is available, so a quiet queue typically yields one or two messages long before the wait elapses. The wait bounds how long SQS will hold an *empty* request open, not how long it accumulates a batch.
- Why do some teams see client-side timeout errors right after enabling long polling?Their HTTP client's socket read timeout is shorter than WaitTimeSeconds, so it aborts the connection while SQS is still legitimately waiting. Raise the client read timeout above the wait value — the AWS SDKs normally handle this, hand-rolled clients and tightened custom timeouts often do not.
- Is there any case where short polling is the right choice?Rarely, and only with continuously busy queues where you want the lowest possible per-call latency and empty responses are effectively impossible — for example a drain loop that always has backlog. Even then the saving is marginal, and the cost of empty receives on any lull is real.
saying these in an interview costs you the question
- Thinks an empty short-poll response proves the queue is empty
- Says long polling adds latency because it waits
- Believes long polling holds the call open to fill the whole batch
- Sets WaitTimeSeconds above 20 and expects it to work
- Confuses long polling with push delivery or a subscription