skip to content

A Lambda function triggered by an Amazon SQS queue receives a batch of 10 messages and one of them fails. By default what happens to the other nine, and how does the ReportBatchItemFailures setting change it?

level: middleimportance: should knowfreq 46%

answer

  1. success acknowledges the whole batch
  2. one bad record replays all ten
  3. report the failed IDs, keep the rest
  4. return, never throw
  5. empty list means everything succeeded

basics

~20 s

By default a failing Lambda invocation deletes nothing, so all ten messages become visible again after the visibility timeout and the nine successes are reprocessed. Enabling ReportBatchItemFailures lets the handler return only the failed message IDs, and Lambda deletes the rest.

solid answer

~50 s

When Lambda polls SQS it deletes the batch's messages only if the invocation succeeds. So by default, one bad record fails the whole invocation, nothing is deleted, and all ten messages reappear when the visibility timeout expires — the nine that worked get processed again, and again on every retry until the poison message is routed away by the queue's redrive configuration. The fix is a **partial batch response**: add `ReportBatchItemFailures` to the event source mapping's `FunctionResponseTypes`, and have the handler return `{"batchItemFailures": [{"itemIdentifier": "<messageId>"}]}` listing only the records it could not process. Lambda deletes every message not named there and leaves the named ones to become visible again. Two rules make or break it: the handler must *return* rather than throw — an uncaught exception still fails the whole batch — and an empty `batchItemFailures` list means complete success.

code

python · 17 lines
python
def handler(event, context):
    batch_item_failures = []

    for record in event["Records"]:
        try:
            process(record["body"])
        except Exception:
            # identify the record by its SQS messageId
            batch_item_failures.append({"itemIdentifier": record["messageId"]})

    # returning (not raising) is what lets Lambda delete the successes
    return {"batchItemFailures": batch_item_failures}


def process(body):
    if body == "poison":
        raise ValueError("cannot process")

go deeper

for a junior

Know that a Lambda invocation triggered by SQS acknowledges the whole batch at once, so a single failing record sends every message in that batch back to the queue.

for a middle

Explain the ReportBatchItemFailures contract precisely: the FunctionResponseTypes setting on the event source mapping, the batchItemFailures list keyed by messageId, and why returning beats throwing.

for a senior

Bring the operational consequences: repeated side effects on the records that already succeeded, wasted invocations on every retry round, and sizing the queue's visibility timeout against the function timeout so leases outlive invocations.

for a principal

Frame batch size and failure isolation as a design decision — larger batches buy throughput and cost efficiency but widen the replay blast radius, and per-record reporting only pays off when records are genuinely independent.

## How Lambda acknowledges an SQS batch With an SQS event source mapping, Lambda does the polling for you: it calls `ReceiveMessage`, invokes your function with the messages as the event, and — on a **successful** invocation — deletes them. That last clause is the whole story. Success means the whole batch is acknowledged; failure means none of it is. So the default behaviour with one bad record in ten is: 1. The handler throws (or times out). 2. Lambda deletes nothing. 3. All ten messages stay leased until the visibility timeout expires, then become visible again. 4. Lambda receives them again and reruns all ten — including the nine that already succeeded. 5. This repeats until the queue's redrive configuration routes the poison message away. The cost is not just wasted invocations: any non-idempotent side effect in those nine successes happens once per retry round. ## Partial batch responses The event source mapping has a `FunctionResponseTypes` field. Set it to `["ReportBatchItemFailures"]` and Lambda changes its contract with your function: instead of reading only success or failure, it reads the function's **return value** to decide which messages to delete. The handler returns a list of the records it could not process, identified by SQS `MessageId`: ```python def handler(event, context): failures = [] for record in event["Records"]: try: process(record["body"]) except Exception: failures.append({"itemIdentifier": record["messageId"]}) return {"batchItemFailures": failures} ``` Lambda deletes every message *not* listed and leaves the listed ones alone, so only they return to the queue after the visibility timeout. The nine successes are gone for good; the one failure retries on its own. ## The rules that trip people up - **Return, don't throw.** If the function raises, Lambda never sees a response body and falls back to whole-batch failure. All the per-record `try`/`except` work is pointless if an exception escapes the loop. - **An empty list means total success.** Returning `{"batchItemFailures": []}` — or a response Lambda reads as empty — tells Lambda the entire batch is done and everything is deleted. A handler that swallows errors and returns an empty list has built a message-losing machine. - **`itemIdentifier` is the `messageId`, not the receipt handle.** If an identifier is empty, null or does not correspond to a message in the batch, Lambda treats the whole batch as failed — safe, but not what you intended. - **Visibility timeout must accommodate the function.** AWS requires the queue's visibility timeout to be at least the function timeout, and recommends around six times it, so a slow invocation cannot have its messages redelivered while it is still running. ## When to reach for it Partial batch responses matter most when batches are large, records are independent, and reprocessing is expensive or visible — sending notifications, writing to a downstream API, charging something. If your batch size is 1, none of this applies. If the records in a batch are interdependent, per-record failure reporting may be the wrong shape entirely and a smaller batch is simpler. It is also worth being honest about what it does *not* solve. It reduces redundant reprocessing; it does not make processing exactly-once. A message can still be redelivered if the invocation dies after the side effect but before Lambda's delete, so the handler's effects still need to tolerate repetition. Partial batch responses are an efficiency and blast-radius control, not a delivery guarantee. ## The interview-ready summary Default: the batch is acknowledged as a unit, so one failure replays all ten. With `ReportBatchItemFailures` on the event source mapping, the function returns the failed `messageId`s and Lambda acknowledges the rest — provided the handler returns instead of throwing, and never reports an empty failure list it did not earn.

  • What happens if the handler catches every error but still lets an exception escape the loop?
    Lambda never receives a response body, so it falls back to whole-batch failure: nothing is deleted and all the messages return to the queue when the visibility timeout expires. Partial batch responses only work when the function *returns* its failure list; an uncaught exception discards it entirely.
  • Your handler returns an empty batchItemFailures list on every invocation. What is the consequence?
    Lambda reads that as complete success and deletes the whole batch every time — including records that failed. Errors are swallowed and messages disappear without the work being done, which is silent data loss rather than the retry the code was trying to express.
  • Why does AWS want the queue's visibility timeout set well above the function timeout for an SQS event source mapping?
    Because the lease has to cover the entire invocation plus Lambda's bookkeeping. If the timeout lapses first, messages become visible while the function is still running and a second invocation picks them up — the same duplicate-processing bug you get with any under-sized visibility timeout. AWS requires at least the function timeout and recommends roughly six times it.
  • Does enabling ReportBatchItemFailures make SQS-triggered processing exactly-once?
    No. It narrows what gets replayed, not whether replay happens. An invocation can still die after a side effect but before Lambda's delete, so any given record may be processed more than once. It is a blast-radius and efficiency control; the handler's effects still have to tolerate being repeated.

saying these in an interview costs you the question

  • Thinks Lambda deletes each record as it is processed
  • Reports failures but still throws from the handler
  • Returns an empty batchItemFailures list while errors occurred
  • Uses the receipt handle as itemIdentifier instead of the messageId
  • Claims partial batch responses give exactly-once processing

context