skip to content

Event Source Mappings

For queues and streams, Lambda does not get pushed to — a managed poller reads the source and calls my function in batches. I learn batch sizing, filtering and partial-batch failure reporting, because that is where poison messages and stuck shards get diagnosed.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

6

You attach an SQS queue to a Lambda function, yet SQS never pushes anything anywhere. What is the event source mapping, and how do queue messages actually end up in your handler?

level: juniorimportance: must knowfreq 72%

answer

  1. nothing pushes; something polls
  2. a separate resource, not a function setting
  3. batches, not single messages
  4. success is what deletes the message
  5. the execution role reads the source

basics

~20 s

An event source mapping is a separate Lambda resource that runs an AWS-managed poller. It reads the queue for you, groups messages into a batch, and invokes your function synchronously with that batch in the event's Records array.

solid answer

~50 s

SQS is a pull service — nothing pushes to Lambda. When you "add an SQS trigger", Lambda creates an **event source mapping** (ESM), a resource with its own ARN and its own configuration, separate from the function. AWS runs a fleet of pollers for that mapping: they call `ReceiveMessage` on the queue with long polling, collect up to `BatchSize` messages, and then invoke your function **synchronously** with an event whose `Records` array holds those messages. If the invocation returns successfully, the poller deletes those messages from the queue; if the handler throws, it deletes nothing and the whole batch becomes visible again later. Two consequences matter immediately: the permission to read the queue lives in the function's **execution role**, not in a queue policy pointing at Lambda, and your handler must loop over `Records` rather than assume one message per invocation.

go deeper

for a junior

Be able to say plainly that a managed poller reads the queue and calls your function with a batch, and that your handler must loop over the Records array rather than expect a single message.

for a middle

Explain the acknowledge rule — success deletes, failure redelivers the whole batch — and where the read permissions sit. Expect to be asked how the batch interacts with the function timeout.

for a senior

Show that you treat the mapping as its own resource with its own tuning knobs and failure policy, and that you design handlers to be idempotent because whole-batch redelivery is the default behaviour.

for a principal

Own the choice of poll-based versus push-based integration itself: what backpressure you gain from a queue in front of compute, and what it costs you in end-to-end latency and duplicate handling across the platform.

## Two ways things reach a Lambda function Lambda has two fundamentally different paths to your code. The **push** path is a caller invoking the `Invoke` API — API Gateway, an SDK client, S3 event notifications, SNS. Something outside Lambda decides when to call you and needs permission to do it, which is why those integrations require a resource-based policy on the function. The **poll** path exists because some sources cannot push at all. SQS, Kinesis Data Streams, DynamoDB Streams and managed Kafka are all *pull* services: a consumer must ask for records. Lambda bridges that gap with an **event source mapping**. ## What the mapping actually is An ESM is a first-class AWS resource, created by `CreateEventSourceMapping`, with its own UUID, its own ARN and its own settings (`BatchSize`, `MaximumBatchingWindowInSeconds`, `FilterCriteria`, `Enabled`, and more). It is *not* a property of the function, which is why you can disable a trigger without touching your code, and why deleting a function does not automatically tidy up everything you configured. Behind the mapping AWS runs a managed poller fleet you never see, do not pay for as compute, and cannot log into. Its loop is simple: 1. Long-poll the source (`ReceiveMessage` for SQS, `GetRecords` on a shard iterator for streams). 2. Assemble a batch, up to `BatchSize` records or the invocation payload limit, whichever comes first. 3. Invoke the function **synchronously** (`RequestResponse`) with the batch as the event. 4. On success, acknowledge — delete the SQS messages, or advance the stream checkpoint. 5. On failure, do not acknowledge, and retry according to the mapping's failure settings. That step 4/5 split is the whole model: **the poller, not your code, decides what counts as processed**, and it decides on the basis of whether the invocation returned or threw. ## The event your handler receives For SQS the event is a JSON object with a `Records` array; each element carries `messageId`, `receiptHandle`, `body` (a string), `attributes` and `messageAttributes`. A handler that reads `event.Records[0]` and stops is a real, common bug — the batch is however many messages the poller happened to gather. ```javascript export const handler = async (event) => { for (const record of event.Records) { const payload = JSON.parse(record.body); await process(payload); } }; ``` Because the invoke is synchronous, the function's timeout bounds the *whole batch*, not one message. Ten messages that each take four seconds need a timeout above forty seconds, and the source's own redelivery window has to be longer than that timeout or the batch will be handed out again while you are still working on it. ## Permissions run the other way round With a push integration, the *caller* needs `lambda:InvokeFunction`. With an ESM, Lambda is the caller of the *source*, so the function's **execution role** needs the read permissions: `sqs:ReceiveMessage`, `sqs:DeleteMessage` and `sqs:GetQueueAttributes` for a queue; `kinesis:GetRecords`, `kinesis:GetShardIterator`, `kinesis:DescribeStream` and `kinesis:ListShards` for a stream. The AWS-managed policies `AWSLambdaSQSQueueExecutionRole` and `AWSLambdaKinesisExecutionRole` exist precisely to package these. A newly wired trigger that sits in state `Disabled` or reports an error almost always means the role is missing one of them. ## Scaling comes from the mapping, not from your traffic You never see a request rate here. For a queue, Lambda adds pollers while the backlog grows and removes them when it shrinks, so concurrency tracks *depth and duration*. For a stream, the shape is fixed by the source: by default one concurrent invocation per shard, so a five-shard stream gives at most five parallel executions of your function no matter how much data arrives. Understanding that difference is the start of every capacity conversation about poll-based Lambdas. ## Where beginners go wrong The three recurring mistakes are assuming SQS "triggers" the function directly (it does not — you can point twenty consumers at a queue and Lambda is just one of them), assuming one message per invocation, and calling `DeleteMessage` by hand inside the handler. That last one is not merely redundant: it breaks partial-failure reporting later, because you have already destroyed the evidence the poller uses to decide what to retry.

  • Where do the permissions to read the queue have to live, and why is that surprising to people used to S3 or SNS triggers?
    They live in the function's execution role, because with an event source mapping Lambda is the one calling SQS. Push sources are the opposite: S3 or SNS invokes the function, so they need `lambda:InvokeFunction` granted by a resource-based policy on the function. A queue-triggered function that never fires is very often an execution role missing `sqs:ReceiveMessage` or `sqs:DeleteMessage`.
  • If your handler throws halfway through a batch of ten, what has Lambda done with the messages it had already handled?
    Nothing — the poller deletes messages only when the invocation returns successfully, so all ten go back to the queue and all ten are delivered again. Your code will re-process the first few, which is why handlers must be idempotent by default. Reporting partial batch failures is the mechanism that narrows redelivery down to the messages that actually failed.
  • How is the invocation type here different from an S3 event notification calling the same function?
    An event source mapping invokes synchronously — the poller waits for the response and uses it to decide whether to acknowledge. S3 notifications invoke asynchronously: the caller gets an accepted response immediately and Lambda's own internal queue handles retries. So with an ESM, retry behaviour is governed by the mapping and the source; with async invocation it is governed by Lambda.

saying these in an interview costs you the question

  • Says SQS pushes messages into Lambda
  • Reads only event.Records[0] and ignores the rest
  • Puts queue read permissions in a resource-based policy on the function
  • Calls sqs:DeleteMessage manually inside the handler
  • Thinks the function timeout applies per message, not per batch

context

open as a page

An SQS-triggered Lambda processes batches of ten. When one message fails, the other nine are delivered and processed all over again. Why does that happen, and how do you make only the failing message be retried?

level: seniorimportance: must knowfreq 54%

basics

~20 s

The event source mapping deletes messages only when the invocation succeeds, so any thrown error redelivers the whole batch. Enable ReportBatchItemFailures on the mapping and return a batchItemFailures list of the failed message IDs so only those are retried.

open as a page

On a Lambda event source mapping, what do BatchSize and MaximumBatchingWindowInSeconds control, and what actually makes the poller stop collecting and invoke your function?

level: middleimportance: should knowfreq 56%

basics

~20 s

BatchSize is the maximum number of records per invocation and MaximumBatchingWindowInSeconds is how long the poller may keep gathering them. The invoke fires on whichever comes first: the batch is full, the window expires, or the payload reaches Lambda's synchronous 6 MB limit.

open as a page

A Kinesis-triggered Lambda has stopped making progress on one shard: the same batch keeps failing and the iterator age climbs for hours. Why does one bad record stall the shard, and which event source mapping settings get it moving again?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A shard is processed in order, so the mapping retries the failing batch rather than skipping ahead, and by default it retries until the records expire from the stream. Bound it with MaximumRetryAttempts and MaximumRecordAgeInSeconds, isolate the record with BisectBatchOnFunctionError, and capture it via the mapping's on-failure destination.

open as a page

An SQS-triggered Lambda scales out under load and exhausts the connection limit of the RDS database behind it. How do you cap how much of that queue is processed at once, and why is putting reserved concurrency on the function the wrong lever?

level: principalimportance: should knowfreq 44%

basics

~20 s

Set ScalingConfig MaximumConcurrency on the event source mapping, which tells the poller itself not to exceed that many concurrent invocations. Reserved concurrency instead lets the poller keep receiving messages and be throttled, inflating receive counts and pushing healthy messages toward the dead-letter queue.

open as a page

A Lambda function is invoked for every message on a busy SQS queue but discards about 90% of them immediately. How does FilterCriteria on the event source mapping change that, and what happens to the messages that are filtered out?

level: middleimportance: nice to knowfreq 40%

basics

~20 s

FilterCriteria puts event-pattern matching in the AWS-managed poller, so non-matching records never reach the function and are never billed as invocations. For SQS, filtered-out messages are deleted from the queue; for streams, the poller simply advances past them.

open as a page