Amazon SQS's DeleteMessageBatch can return HTTP 200 while some entries failed. Why does it work that way, and what must a consumer do about it?
answer
- per-entry, not atomic
- 200 does not mean all ten landed
- Successful and Failed lists
- a missed delete means reprocessing later
- SenderFault tells you who to blame
basics
~20 sSQS batch APIs are per-entry, not atomic: the call succeeds while individual entries can fail, so the response splits into Successful and Failed lists. A consumer must inspect Failed and retry those deletes, or the messages reappear after the visibility timeout and get processed again.
solid answer
~50 s`DeleteMessageBatch`, `SendMessageBatch` and `ChangeMessageVisibilityBatch` take up to ten entries and each entry is handled independently — there is no transaction across a batch. A 200 means "the request was accepted and here is what happened to each entry", so the response carries a `Successful` list and a `Failed` list of error entries, keyed by the per-entry `Id` you supplied. Most SDKs raise nothing in this case, which is exactly the trap: code that calls the batch API and moves on has silently skipped some deletes. Those messages are still leased, and when the visibility timeout expires they become visible again and get reprocessed. The consumer must read `Failed`, map each entry back to its receipt handle via the `Id`, and retry — throttling and transient sender faults are the usual causes and normally succeed on a retry. Batching is still worth it: ten deletes cost one billable request instead of ten.
code
python · 18 linesimport boto3
sqs = boto3.client("sqs")
def delete_batch(queue_url, handles):
# local Ids only correlate request entries with results
by_id = {str(i): h for i, h in enumerate(handles)}
entries = [{"Id": i, "ReceiptHandle": h} for i, h in by_id.items()]
resp = sqs.delete_message_batch(QueueUrl=queue_url, Entries=entries)
retryable = []
for err in resp.get("Failed", []):
if err["SenderFault"]:
print("bug in our entry", err["Id"], err["Code"], err["Message"])
else:
retryable.append(by_id[err["Id"]])
return retryable # feed these back in with backoffgo deeper
Know that SQS batch calls handle each entry separately, so you must look at the response's Failed list rather than assuming a successful call deleted everything.
Explain that batches are non-atomic, that the per-entry Id is your correlation key back to the receipt handle, and that a missed delete resurfaces as a reprocessed message after the visibility timeout.
Show the operational instinct: use SenderFault to split bugs from throttling, retry the transient subset with backoff, and alarm on a persistent delete-failure rate before it shows up as duplicate work downstream.
Push this into shared consumer tooling rather than per-team loops — a house library that batches, correlates, retries and reports failed acknowledgements removes an entire class of quiet duplicate-processing incidents.
## Batch means ten independent operations, not one transaction SQS offers batch variants of the three per-message calls: `SendMessageBatch`, `DeleteMessageBatch` and `ChangeMessageVisibilityBatch`. Each accepts up to ten entries in a single request, and — this is the whole point of the question — each entry is evaluated on its own. There is no all-or-nothing semantic. A batch can be entirely successful, entirely failed, or anything in between, and the HTTP status reflects only whether the *request* was well-formed and authorized. So the response is shaped as two lists: ```python resp = sqs.delete_message_batch(QueueUrl=queue_url, Entries=entries) resp["Successful"] # [{"Id": "0"}, {"Id": "2"}, ...] resp["Failed"] # [{"Id": "1", "SenderFault": False, # "Code": "...", "Message": "..."}] ``` Each entry you send carries an `Id` that you choose, unique within that request. It is not the message ID and it means nothing to SQS beyond correlation — its only job is to let you match a result back to the entry you submitted. Keep a local map from `Id` to receipt handle so a failed entry can be retried without guessing. ## Why a delete failure is a correctness problem, not a cosmetic one A failed delete leaves the message on the queue. It is still leased for the remainder of its visibility timeout, so nothing looks wrong immediately — and then the lease expires, the message becomes visible, another consumer receives it, and the work runs a second time. The symptom shows up minutes later and far from the cause, which is why this bug survives code review so often. The failure modes are ordinary: request throttling when a consumer fleet hammers a queue, transient service-side faults, an expired or malformed receipt handle. The `SenderFault` boolean in each error entry tells you which side to blame — `true` means your entry was bad (a stale handle, a duplicate `Id`) and retrying it unchanged will fail again; `false` means it was a service-side or throttling condition and a retry with backoff is the right response. ## What correct handling looks like 1. Build the entries with stable local IDs and keep the `Id` → receipt-handle map. 2. Call the batch API. 3. Iterate `Failed`. For `SenderFault=false`, retry those entries with exponential backoff. For `SenderFault=true`, log loudly with the error `Code` — that is a bug in your code, not a blip. 4. Never treat the absence of an exception as proof the deletes landed. A short retry of the failed subset is usually enough, and it is far cheaper than reprocessing the work. ## Why batch at all SQS bills per request. Batching deletes ten-to-one cuts both the bill and the round-trip overhead on a high-throughput consumer, and the same applies to sends and to visibility changes. The pairing that matters in practice is `ReceiveMessage` with `MaxNumberOfMessages` up to ten, followed by one `DeleteMessageBatch` covering the messages you actually finished — note the *actually finished* part: batching does not license deleting messages whose work failed, and the entries you include should be exactly the ones you processed successfully. ## Two related traps - **Duplicate entry IDs.** The `Id` values must be unique inside a single request; reusing one is a sender fault for that entry. - **Deleting the whole received batch reflexively.** Convenient loops that receive ten and delete ten regardless of per-message outcome turn a partial processing failure into silent data loss. Track outcomes per message and delete only the successes. The rule to carry away: with SQS batch APIs, a 200 is a receipt for the *request*, not for the work. Read the `Failed` list every time.
- What is the Id field on each batch entry, and does SQS use it for anything else?It is a correlation token you choose, unique within that one request. SQS echoes it back in the Successful and Failed lists so you can map results to entries, and that is its entire purpose — it is not the MessageId, it is not persisted, and reusing a value inside a request is a sender fault.
- How does the SenderFault flag change your retry decision?SenderFault=true means the entry itself was invalid — a stale receipt handle, a duplicate Id — so retrying it unchanged fails again and it should be logged as a bug. SenderFault=false points at throttling or a transient service condition, where retrying that subset with exponential backoff is the correct response.
- A consumer receives ten messages and deletes all ten in one batch regardless of outcome. What is wrong with that?It acknowledges work that may have failed. Deleting is the acknowledgement, so a message whose handler threw is now gone with the work undone — silent loss rather than a retry. Track per-message outcomes and include only the successes in the delete batch.
saying these in an interview costs you the question
- Treats an HTTP 200 as proof every entry succeeded
- Never inspects the Failed list at all
- Confuses the batch entry Id with the MessageId
- Deletes the whole received batch regardless of per-message outcome
- Retries sender-fault entries unchanged in a tight loop