skip to content

An SQS-triggered Lambda processes batches of ten. When one message fails, the other nine are delivered and processed all over again. Why does that happen, and how do you make only the failing message be retried?

level: seniorimportance: must knowfreq 54%

answer

  1. the poller only learns pass or fail
  2. opt in on the mapping, report in the response
  3. itemIdentifier: messageId
  4. empty list means everything succeeded
  5. on streams it checkpoints, it does not skip

basics

~20 s

The event source mapping deletes messages only when the invocation succeeds, so any thrown error redelivers the whole batch. Enable ReportBatchItemFailures on the mapping and return a batchItemFailures list of the failed message IDs so only those are retried.

solid answer

~50 s

By default the poller has exactly one bit of information: did the invocation succeed. Success deletes every message in the batch; failure deletes none, so all ten come back and the nine good ones are processed twice. To get finer resolution you opt into **partial batch responses**: set `FunctionResponseTypes` to `["ReportBatchItemFailures"]` on the event source mapping, catch errors **per record** so the handler returns normally, and return `{"batchItemFailures": [{"itemIdentifier": "<messageId>"}]}` listing only what failed. Lambda deletes everything not named and redelivers the rest. Two traps matter in production: an empty list means "all succeeded", so a handler that swallows an exception and returns an empty list silently discards work; and if the response is missing, malformed or names an ID that was not in the batch, Lambda falls back to treating the whole batch as failed. None of this removes the need for idempotency — it shrinks duplicate processing, it does not eliminate it.

go deeper

for a junior

Understand that a failure anywhere in the batch means the whole batch is delivered again, so your handler will sometimes see the same message twice and must tolerate it.

for a middle

Be able to name the three moving parts — the FunctionResponseTypes setting, catching per record so you return rather than throw, and the batchItemFailures response shape — and what an empty list means.

for a senior

Demonstrate the production judgment: the silent-success trap, the malformed-response fallback to whole-batch failure, and how reporting stops healthy messages from being dead-lettered for sharing a batch.

for a principal

Own the delivery-semantics contract across the platform — where idempotency keys live, how duplicate work is bounded and detected, and what batch size the organisation should default to given that failure isolation costs code discipline.

## Why the default is all-or-nothing An event source mapping's poller invokes your function synchronously and then makes one decision: acknowledge the batch or not. With an ordinary handler, the only signal it gets is whether the invocation returned or threw. There is no per-record channel, so there is no per-record outcome — a thrown exception means *nothing* is deleted, and every message in the batch becomes visible again for redelivery. At a batch size of 10 with a 1% failure rate, that is roughly a 10% chance per batch of re-processing nine perfectly good messages. At a batch size of 500 it is close to a guarantee, which is why large batches and partial failure reporting always travel together. ## Turning on partial batch responses There are three parts, and all three are required. **1. Configure the mapping.** Set `FunctionResponseTypes` to `["ReportBatchItemFailures"]`. Until you do, Lambda ignores whatever your function returns. ```bash aws lambda update-event-source-mapping \ --uuid 1a2b3c4d-5e6f-7a8b-9c0d-1e2f3a4b5c6d \ --function-response-types ReportBatchItemFailures ``` **2. Stop throwing.** The handler must return normally, which means catching per record. If it throws, the response never arrives and the whole batch fails — the setting has no effect on an exception. **3. Return the right shape.** ```javascript export const handler = async (event) => { const batchItemFailures = []; for (const record of event.Records) { try { await process(JSON.parse(record.body)); } catch (err) { console.error({ messageId: record.messageId, err }); batchItemFailures.push({ itemIdentifier: record.messageId }); } } return { batchItemFailures }; }; ``` For SQS, `itemIdentifier` is the record's `messageId`. For Kinesis and DynamoDB Streams it is the record's **sequence number**, and the semantics change accordingly — see below. ## The failure modes people meet in production **The silent-success bug.** `batchItemFailures: []` means every record succeeded, and Lambda deletes the entire batch. A handler with a broad `catch` that logs and forgets, then returns an empty list, is deleting failed work with no dead-letter trail and no error metric. If you catch, you must report. **The malformed-response fallback.** If the response is absent, has the wrong shape, or names an identifier that was not in this batch, Lambda cannot trust it and treats the batch as a complete failure. This is a safe default, but it looks like the feature is "not working" — the usual culprit is a typo in the key names or a handler that returns the array directly instead of the wrapping object. **Ordering on streams.** For Kinesis and DynamoDB Streams, records within a shard are ordered, so the report is not a set — it is a checkpoint. Lambda takes the **lowest** reported sequence number and retries the batch from there, meaning records after that point are re-delivered even if you did not name them. You cannot skip a failure and keep going on an ordered shard; that would break ordering. **Duplicates remain possible.** A timeout, an out-of-memory kill, or a crash after the work but before the response still redelivers everything. Partial batch reporting is an optimisation on top of at-least-once delivery, not a replacement for idempotent handlers. ## Interaction with the retry chain For SQS, a reported failure simply means the message is not deleted. It reappears, its receive count increments, and once it exceeds the queue's redrive threshold it lands in the dead-letter queue. That is the behaviour you want: a poison message reaches the dead-letter queue at the same rate as before, while its batch-mates go through cleanly the first time. One consequence worth stating out loud in an interview: without partial reporting, the good messages' receive counts also increment on every failed batch. A queue with an aggressive redrive threshold can send perfectly healthy messages to the dead-letter queue purely because they kept sharing a batch with a bad one. Enabling `ReportBatchItemFailures` fixes that class of false positive as well. ## How to demonstrate this well Name the three parts (mapping setting, catch per record, response shape), state the empty-list semantics unprompted, and mention that the handler still has to be idempotent. Then add the stream nuance about the lowest sequence number — it is the detail that separates someone who has read the documentation from someone who has run this.

  • Your handler catches every error, logs it, and returns an empty batchItemFailures array. What is the consequence?
    Every message in the batch is deleted, including the ones that failed. An empty list is a positive assertion that the whole batch succeeded, so the failures leave no trace beyond a log line: no redelivery, no dead-letter queue entry, no error metric on the function. Catching without reporting is strictly worse than not catching at all.
  • How does the identifier and its meaning differ between SQS and Kinesis when reporting partial failures?
    For SQS it is the message ID and the report is a set — exactly the named messages come back. For Kinesis and DynamoDB Streams it is the sequence number, and because a shard is ordered, Lambda retries from the lowest reported sequence number onward. Records after that point are redelivered whether or not you named them; you cannot skip past a failure and keep the rest.
  • Does enabling partial batch responses change when a message reaches the dead-letter queue?
    It changes which messages get there. A reported message is left in the queue, so its receive count climbs and it reaches the dead-letter queue on the normal schedule. What stops happening is the false positives: without reporting, the healthy messages sharing a batch with a poison one also accumulate receives and can be dead-lettered despite never having failed.
  • Why is idempotency still required after you enable this?
    Because reporting narrows redelivery, it does not remove it. A timeout, a memory kill or a crash means no response reaches the poller at all, so the whole batch is redelivered. Delivery is at-least-once by design, and any record can be handed to your function more than once regardless of how carefully you report failures.

saying these in an interview costs you the question

  • Thinks returning the failure list works without setting FunctionResponseTypes
  • Throws from the handler and expects only that record to be retried
  • Returns an empty failure list after swallowing errors
  • Believes partial reporting makes processing exactly-once
  • Assumes streams can skip a failed record and continue

context