skip to content

A Kinesis-triggered Lambda has stopped making progress on one shard: the same batch keeps failing and the iterator age climbs for hours. Why does one bad record stall the shard, and which event source mapping settings get it moving again?

level: seniorimportance: should knowfreq 42%

answer

  1. order is the constraint
  2. the default retry count is -1
  3. IteratorAge climbing in a straight line
  4. split the batch, find the one record
  5. the destination carries metadata, not payloads

basics

~20 s

A shard is processed in order, so the mapping retries the failing batch rather than skipping ahead, and by default it retries until the records expire from the stream. Bound it with MaximumRetryAttempts and MaximumRecordAgeInSeconds, isolate the record with BisectBatchOnFunctionError, and capture it via the mapping's on-failure destination.

solid answer

~50 s

Stream event source mappings preserve order within a shard, so a failed batch cannot be set aside — the poller retries it, and the checkpoint does not advance. With the defaults, `MaximumRetryAttempts` is unlimited, so it keeps retrying until the records age out of the stream's retention, which is why `IteratorAge` climbs for hours while everything behind the poison record waits. Four settings fix it. `MaximumRetryAttempts` and `MaximumRecordAgeInSeconds` bound how long a batch may block. `BisectBatchOnFunctionError` splits a failing batch in half and retries each half, converging on the individual bad record so only it is discarded rather than the whole batch. `DestinationConfig` with an `OnFailure` SQS queue or SNS topic captures **metadata** about the discarded batch — the shard, sequence numbers and error — so you can investigate. And `ReportBatchItemFailures` lets a good handler checkpoint past the records it did complete. Alarm on `IteratorAge`; it is the metric that makes this visible at all.

go deeper

for a junior

Know that stream records are processed in order per shard, so a record your function cannot handle holds up everything behind it until it is retried successfully or given up on.

for a middle

Explain the mapping's retry controls — maximum attempts, maximum record age, bisecting on error, the on-failure destination — and what the default of unlimited retries implies.

for a senior

Show the diagnosis: alarm on iterator age, recognise the linear climb, and design the handler to quarantine terminal failures itself rather than relying on the mapping's backstop.

for a principal

Own the data-loss contract: what retention buys you, what an expired shard costs the business, and how error budgets and quarantine paths are standardised across every stream consumer you run.

## Why a stream stalls where a queue would not An SQS queue has no order to protect, so a message the consumer cannot handle is simply retried on its own and eventually dead-lettered while everything else flows. A Kinesis shard or DynamoDB stream shard is different: records are ordered, and the event source mapping's contract is to deliver them in order. If batch N fails, the mapping cannot deliver batch N+1 — that would reorder the stream. So it retries batch N, and the shard's checkpoint stays where it is. One record your handler cannot parse blocks every record behind it on that shard. The default makes it worse: `MaximumRetryAttempts` defaults to **-1**, meaning retry indefinitely. "Indefinitely" ends only when the records fall out of the stream's retention window. Until then, the shard produces nothing but repeated failures, and everything behind the poison record is silently delayed and then, once retention expires, lost. ## The metric that tells you `IteratorAge` (published by Lambda as `IteratorAge` for the function, and by Kinesis as `GetRecords.IteratorAgeMilliseconds`) is the age of the oldest record in the last batch read. A healthy consumer holds it near zero. A stalled shard shows it climbing linearly and never recovering — a straight diagonal line is the signature of a blocked shard, as opposed to the sawtooth of a consumer that is merely behind and catching up. An alarm on iterator age is the single highest-value alarm on any stream-triggered Lambda, because nothing else surfaces the problem: invocations continue, errors are logged, and the function looks busy while delivering zero progress. ## The four settings that unblock it **`MaximumRetryAttempts`** caps retries for a batch. After the cap, the batch is discarded and the checkpoint advances. Set it to a number you can justify — enough to ride out a transient downstream outage, not so many that a genuinely undeliverable record holds the shard for hours. **`MaximumRecordAgeInSeconds`** bounds the same thing in time rather than attempts, and it also drops records that are already older than the limit when read. For a stream whose value is real-time, this is often the more natural knob: data that is an hour late may be worthless anyway. **`BisectBatchOnFunctionError`** is the precision tool. When a batch fails, the mapping splits it in two and retries each half, recursing until it isolates the individual record that fails. Without it, exhausting the retry cap throws away the entire batch — potentially hundreds of good records — to get rid of one bad one. With it, you lose the one. The cost is extra invocations while bisecting, which is a trivial price. **`DestinationConfig` / `OnFailure`** points at an SQS queue or SNS topic that receives a record when a batch is finally discarded. Note precisely what it contains: **metadata**, not your data — the stream ARN, the shard ID, the starting and ending sequence numbers, the approximate arrival timestamp and the error. To recover the actual payload you re-read the stream at those sequence numbers, while it is still within retention. Candidates who claim the failed records themselves are delivered have not used it. ```bash aws lambda update-event-source-mapping \ --uuid 1a2b3c4d-5e6f-7a8b-9c0d-1e2f3a4b5c6d \ --maximum-retry-attempts 5 \ --maximum-record-age-in-seconds 3600 \ --bisect-batch-on-function-error \ --destination-config '{"OnFailure":{"Destination":"arn:aws:sqs:eu-west-1:111122223333:stream-failures"}}' ``` **`ReportBatchItemFailures`** complements all of this: a handler that reports the sequence number where it stopped lets the mapping checkpoint past the records it did complete, so a retry does not redo work — though on an ordered shard it still cannot skip past the failure itself. ## What does not help `ParallelizationFactor` runs multiple concurrent batches per shard, which increases throughput, but it does not rescue a poison record — the failing partition key's batches still block. Adding shards does not help either: the bad record lives on one shard and stays there. Raising the function timeout only makes each failed attempt take longer. ## The design lesson Bound retries on every stream mapping, because the default is unbounded and unbounded is never the right answer for data that expires. Then make the handler itself distinguish **retryable** failures (a downstream 503 that will pass) from **terminal** ones (a record that will never parse). Terminal failures belong in your own quarantine — write the record to a dead-letter store and return success — so the checkpoint moves immediately. The mapping's retry and destination settings then exist as the backstop for the failures you did not anticipate, not as your primary error-handling strategy.

  • What exactly arrives in the on-failure destination when a stream batch is discarded?
    Metadata about the batch, not the records. You get the stream ARN, the shard ID, the starting and ending sequence numbers, the approximate arrival timestamp of the first record, and the error context. Recovering the payload means re-reading the shard at those sequence numbers, so it only works while the records are still inside the stream's retention window.
  • Would raising ParallelizationFactor help this stalled shard?
    No. It runs several concurrent batches per shard, split by partition key, so it raises throughput on a shard that is merely behind. A poison record still fails its own batch and still blocks the records ordered behind it, so the stall persists while your invocation count rises. The fix is bounded retries plus bisecting, not more parallelism.
  • Why is IteratorAge the alarm you set, rather than the function's error count?
    Because errors alone cannot distinguish a stalled shard from a noisy but progressing one — the function is invoked and fails repeatedly in both cases. Iterator age measures actual progress: it is the age of the oldest record you just read. A steady linear climb means the checkpoint has stopped moving, which is the condition you care about and the one that ends in data loss at retention.
  • How should the handler itself distinguish retryable from terminal failures?
    Retryable means a downstream problem that will pass — throttling, a 5xx, a timeout — and should fail the batch so the mapping retries. Terminal means the record will never succeed, such as an unparseable payload or a schema violation. Those should be written to your own quarantine store and reported as success, so the checkpoint advances immediately instead of burning the retry budget.

saying these in an interview costs you the question

  • Thinks the mapping skips a failing batch and moves on
  • Leaves MaximumRetryAttempts at its unlimited default
  • Expects the failed records themselves in the on-failure destination
  • Suggests adding shards or parallelism to clear a poison record
  • Watches only the error metric and never iterator age

context