A function is triggered by an SQS queue with a batch size of 10. If 7 of the 10 messages in a batch process successfully but 3 throw errors, what happens to each message by default, and how does 'partial batch response' (reporting individual item failures) change that behavior?
answer
- batch all-or-nothing by default
- no partial ack -> whole batch redelivered
- ReportBatchItemFailures = per-message ack
- must return itemIdentifier list
- protects against duplicate side effects
basics
~20 sNormally, if any message in the batch fails, the whole batch of 10 gets retried later — including the 7 that already succeeded, so they run twice. Partial batch response lets the function tell the queue exactly which 3 failed, so only those get retried and the other 7 aren't repeated.
solid answer
~50 sBy default, an SQS-triggered function's poller treats the batch as all-or-nothing: unless every message in the invocation succeeds, the entire batch is considered failed and none of the 10 messages are deleted from the queue, meaning all 10 — including the 7 that already succeeded — become visible again after the visibility timeout and get redelivered, causing duplicate processing of the successful ones. Enabling partial batch response (AWS calls this ReportBatchItemFailures) changes the contract: the function returns a list of message IDs that specifically failed, and the poller deletes only the successful messages from the queue while leaving just the 3 failed ones for redelivery. This requires the handler code to track per-message success/failure explicitly and return that list, but it eliminates redundant reprocessing and lets failure-prone messages get retried and eventually routed to a DLQ independently of their batch-mates.
go deeper
Should know that a batch failure by default can cause already-successful messages to be reprocessed, and have a rough sense that this can cause duplicate work.
Should describe the delete-vs-leave-alone mechanics driving the all-or-nothing default and know that a feature exists to acknowledge messages individually within a batch.
Should explain exactly how partial batch response changes the poller's delete behavior, know it requires explicit handler code compliance with a specific response shape, and reason about idempotency as the real mitigation when partial batch response isn't available.
Should weigh batch-size and visibility-timeout tuning against blast radius, DLQ threshold precision, and downstream idempotency guarantees as a system design decision, and recognize partial batch response as one layer of defense that doesn't replace idempotent handler design.
## What a batch is, and why batching exists SQS-triggered serverless functions receive not a single message but a **batch** — a poller-configured group of up to some maximum count (10 for standard AWS Lambda SQS integrations, more with certain configurations) pulled from the queue in one long-poll cycle and delivered to a single function invocation as an array. This batching exists to amortize the fixed overhead of invoking a function across multiple units of work: - cold-start risk, - per-invocation billing granularity, - connection setup, and to give the consumer visibility into a chunk of the backlog at once rather than one message at a time. ## The default all-or-nothing contract The default completion contract for a batch is coarse: a poll-based invocation either returns successfully as a whole or throws/errors as a whole, and the poller only has that single binary signal to act on. - **If the invocation returns successfully**, the poller deletes every message in the batch from the queue, and they're gone for good, marked processed. - **If the invocation throws an unhandled error** — or times out — the poller treats the entire batch as unprocessed and does not delete anything. Practically, that means it does nothing explicit at that moment; instead, the visibility timeout that was already applied when the batch was received simply expires naturally, and every message in the batch — including any that were, from the handler's own internal perspective, successfully processed just before the error was thrown — becomes visible to other consumers again and gets redelivered on a subsequent poll. ## Why the default is coarse This all-or-nothing default exists because SQS's core contract only understands two outcomes at the API level: delete a message (success) or leave it alone (implicit failure, handled entirely through visibility timeout expiry). Without extra signaling from the function, the poller has no way to know that only 3 of 10 messages actually failed; it just sees 'this whole batch invocation ended in an error' and defers to the visibility timeout mechanism uniformly across everything it delivered. ## The trade-off of accepting it The trade-off of accepting this default is simplicity at the cost of duplicate work: a handler doesn't need any special bookkeeping, but any batch with even one failing message causes full-batch redelivery, meaning already-successful side effects (a database write, an email sent, a payment charged) can run again. - **For operations that are naturally idempotent** — an upsert keyed by message ID, for instance — this is a tolerable inefficiency. - **For operations with real side effects that aren't idempotent**, it's a correctness bug waiting to happen: a redelivered 'charge customer' message can double-charge unless the handler explicitly deduplicates. ## Partial batch response Partial batch response (AWS's ReportBatchItemFailures feature) closes this gap by changing the return contract. Instead of the function needing to succeed or fail as a monolith, the handler: 1. processes each message individually, 2. catches per-message errors internally, 3. and returns a structured list identifying exactly which message IDs failed (typically leaving the successful ones out of that list entirely). The poller then deletes every message not named in the failure list and leaves only the named failures in the queue for redelivery once their visibility timeout expires. This is not automatic — it must be explicitly enabled on the event source mapping and the handler must be written to comply with the exact response shape (an object listing itemIdentifier entries for failures) or the platform falls back to treating the whole batch as failed regardless of intent, a common source of confusion when teams enable the feature but don't update their handler code correctly. | Behavior | Default | Partial batch response | |---|---|---| | Signal the poller acts on | a single binary signal | a structured list identifying which message IDs failed | | Messages redelivered after one failure | every message in the batch | only the named failures | | Duplicate side effects | full-batch redelivery | only the genuinely bad message keeps cycling | ## The operational failure mode it fixes The operational failure mode this feature exists to fix is duplicate side effects at scale: without it, a single consistently-failing message (a malformed payload, a downstream dependency that's down for one particular record) will keep dragging its 9 well-behaved batch-mates through repeated redelivery and reprocessing every cycle until the bad message finally exhausts its maxReceiveCount and gets diverted to a dead-letter queue — multiplying load on downstream systems by the batch size for no reason. With partial batch response enabled, only the genuinely bad message keeps cycling, and the DLQ threshold is reached faster and more precisely because the retry count only increments for the message that's actually failing, not artificially for its neighbors. ## A concrete scenario A concrete scenario: an order-fulfillment worker processes a batch of 10 SQS messages, each representing 'ship this order.' One order references a product ID that was deleted from the catalog, causing a lookup exception. Without partial batch response, all 10 shipping jobs — nine of which already succeeded and, say, called a shipping-label API — get redelivered and re-run, generating nine duplicate shipping labels and duplicate carrier charges. With partial batch response enabled and the handler correctly catching the one bad order's exception and returning only its message ID as failed, the nine good orders are deleted and never touched again, while only the one broken order retries and eventually lands in the DLQ for manual investigation.
- What happens if a team enables partial batch response but the handler code has a bug and always returns an empty failure list, even for messages it internally caught errors on?The poller sees an empty failure list and interprets that as every message in the batch succeeding, so it deletes all of them — including the ones the handler internally knows failed. Those failures are silently lost forever with no redelivery and no DLQ routing, since from the poller's perspective nothing went wrong.
- How does the batch-size setting interact with the visibility timeout when using partial batch response?The visibility timeout still needs to comfortably exceed the time it takes to process the full batch, not just one message, because all messages in the batch share the same received-at clock even though they finish being processed at different points within the handler's loop. Undersizing it can cause even successfully-processed messages near the end of a large batch to become visible again mid-processing.
- Why might a team deliberately choose a small batch size (like 1 or 2) instead of the maximum, even knowing it increases invocation overhead?A small batch size limits the blast radius of a full-batch retry when partial batch response isn't used or isn't reliable, and it reduces how much work is at risk if the function crashes mid-batch. It's a reasonable trade for handlers with expensive, non-idempotent side effects where duplicate work is costly, even though it means paying for more invocations and losing some batching efficiency.
Default batch handling is like a delivery truck that, if even one of ten packages can't be dropped off, brings ALL ten back to the depot to retry tomorrow — even the nine that were delivered fine. Partial batch response is telling the depot exactly which one package failed, so the other nine stay delivered.
saying these in an interview costs you the question
- Believes SQS automatically knows which individual message in a batch failed without any extra configuration
- Assumes a failed batch only redelivers the specific message that errored
- Doesn't realize successfully-processed messages can be re-run if the batch overall reports failure
- Thinks enabling partial batch response requires no handler code changes
- Confuses visibility timeout with a queue-level retry limit (maxReceiveCount)