skip to content

On a Lambda event source mapping, what do BatchSize and MaximumBatchingWindowInSeconds control, and what actually makes the poller stop collecting and invoke your function?

level: middleimportance: should knowfreq 56%

answer

  1. a maximum, not a promise
  2. three racers, first one wins
  3. the forgotten third: payload size
  4. window zero means small batches
  5. big batch, big blast radius

basics

~20 s

BatchSize is the maximum number of records per invocation and MaximumBatchingWindowInSeconds is how long the poller may keep gathering them. The invoke fires on whichever comes first: the batch is full, the window expires, or the payload reaches Lambda's synchronous 6 MB limit.

solid answer

~50 s

`BatchSize` sets the ceiling on records per invocation; `MaximumBatchingWindowInSeconds` sets how long the poller will wait to reach that ceiling. Three conditions race, and the first one to fire triggers the invocation: the batch reaches `BatchSize`, the batching window elapses, or the accumulated payload hits the 6 MB synchronous invocation limit. With the window at its default of zero, the poller invokes with whatever a read returned, so batches are often smaller than `BatchSize`. Raising an SQS mapping's batch size above 10 requires setting a window of at least one second. The tradeoff is latency against efficiency: bigger batches amortise cold starts, downstream round-trips and per-invocation cost, but they add up to the window's worth of delay, they force the function timeout to cover the whole batch, and they enlarge the blast radius when one poison record fails the invocation.

go deeper

for a junior

Know that records arrive in batches whose size you configure, and that the number you set is a maximum — your handler may well be invoked with fewer.

for a middle

Be ready to name all three conditions that end a batch, including the payload size limit, and to explain the latency-versus-efficiency trade the batching window makes.

for a senior

Show the second-order effects: timeout and redelivery-window arithmetic, cold-start amortisation, and why a large batch without partial failure reporting multiplies duplicate work.

for a principal

Frame batching as a system-level lever — where you spend your latency budget, how burst shape hits downstream capacity, and what per-invocation cost reduction is actually worth at your volume.

## Two knobs, three exit conditions An event source mapping's poller is in a small loop: read records, decide whether it has enough, invoke. Two settings shape that decision. **`BatchSize`** is a maximum, never a guarantee. It is the largest number of records the poller will put into one invocation. For SQS the default is 10; for Kinesis and DynamoDB Streams the default is 100. **`MaximumBatchingWindowInSeconds`** is how long the poller is allowed to keep accumulating before it gives up on filling the batch. It defaults to 0 seconds, and its maximum is 300 seconds. The invocation fires on the *first* of three conditions: 1. the batch has reached `BatchSize` records; 2. the batching window has elapsed; 3. the assembled payload has reached the **6 MB** synchronous invocation payload limit. That third condition is the one people forget, and it is why a mapping configured for 1,000 large records may reliably deliver forty. The size ceiling is enforced by the invoke API, not by your configuration, so no batch-size number can override it. ## Why the window is not "just latency" With the window at 0, an SQS poller invokes with whatever a single `ReceiveMessage` returned — often one or two messages on a lightly loaded queue, even though `BatchSize` is 10. Setting a window of, say, 5 seconds tells the poller to keep reading for up to five more seconds and hand over a fuller batch. That changes economics, not just timing. Lambda bills per invocation and per GB-second, so a function that opens a database connection, resolves a secret or makes an HTTP call once per invocation does that work once per *batch*. Turning 100 single-record invocations into 10 ten-record invocations removes 90 of those fixed costs, and removes 90 chances of paying a cold start. There is also a hard coupling worth memorising: an SQS mapping can only use a `BatchSize` above 10 if `MaximumBatchingWindowInSeconds` is at least 1. AWS will reject the configuration otherwise, because without a window there is nothing to make the poller wait long enough to gather more. ## The costs on the other side **Latency.** In the worst case a record waits the full window before its invocation even starts. For a queue feeding a user-visible workflow, a 30-second window is a 30-second tail you have chosen. **Timeout budget.** The invoke is synchronous and the function timeout applies to the whole batch. Ten records at four seconds each need a timeout above forty seconds — and, for SQS, the source's redelivery window must exceed that timeout, or the messages will be handed to another poller while you are still working. **Blast radius.** By default, one failing record fails the invocation, and the entire batch is redelivered or retried. A batch of 500 turns one poison record into 499 redundant re-processings. Large batches and partial-batch failure reporting belong together; large batches without it are a duplicate-processing machine. **Downstream burst shape.** A batch is a burst. Ten records that each write to a database arrive as ten writes back-to-back. If you use large batches to be gentle on concurrency, remember that you have concentrated, not reduced, the downstream load per invocation. ## Streams differ in one important way For Kinesis and DynamoDB Streams the poller reads from a shard iterator, and records within a shard are ordered. Batch size interacts with `ParallelizationFactor` (how many concurrent batches per shard) and with the `IteratorAge` metric: if iterator age is climbing, you are consuming slower than you are producing, and a bigger batch size is often a cheaper fix than more shards, because it does more work per `GetRecords` round-trip. ```bash aws lambda update-event-source-mapping \ --uuid 1a2b3c4d-5e6f-7a8b-9c0d-1e2f3a4b5c6d \ --batch-size 100 \ --maximum-batching-window-in-seconds 5 ``` ## How to choose Start from the latency budget: the window can never exceed the delay the consumer of the result will tolerate. Then size the batch so that `BatchSize x per-record duration` sits comfortably inside the function timeout, with headroom for a slow downstream day. Then check the payload arithmetic against 6 MB. Then, before you ship anything above a batch of ten, make sure the function reports partial batch failures — otherwise you have traded a small, expensive problem for a large, duplicated one.

  • You set BatchSize to 1000 on an SQS mapping but invocations consistently arrive with far fewer records. Give two reasons.
    Either the batch is being cut off early or there is nothing to fill it. The payload can reach the 6 MB synchronous invocation limit before the record count does, which caps the batch by bytes. Or the batching window is expiring first — on a queue that is not backed up, the poller simply has not seen 1000 messages in that time. Both are normal; batch size is a ceiling.
  • How does raising the batch size change the timeout you need on the function?
    The invoke is synchronous and covers the whole batch, so the timeout has to exceed batch size multiplied by worst-case per-record duration, not average. For SQS, the queue's redelivery window then has to exceed that timeout, or a batch you are still processing is handed out again and you get concurrent duplicate work.
  • When would you deliberately keep the batching window at zero?
    When latency dominates — an interactive or near-real-time workflow where waiting seconds to save invocations is not a trade you can make. It is also the right default while you are still measuring, since a zero window gives the smallest, fastest-failing batches, and it is the only setting compatible with an SQS batch size of 10 or less if you want no added delay at all.

saying these in an interview costs you the question

  • Treats BatchSize as a guaranteed number of records
  • Forgets the 6 MB payload limit can end a batch early
  • Raises batch size without raising the function timeout
  • Thinks a bigger batch reduces total downstream load
  • Uses large batches without partial batch failure reporting

context