skip to content

Event Triggers

What can invoke a function: HTTP through an API gateway, queues, streams, object-storage events, and schedules. You will learn how push and poll-based event source mappings differ in batching, retries and concurrency behaviour.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

6

A serverless platform can invoke a function two different ways depending on the event source: some sources 'push' an event straight to the function, while others are 'polled' by the platform on the function's behalf. Using an HTTP API Gateway trigger and a message-queue trigger as examples, what is the basic difference between these two invocation models?

level: juniorimportance: must knowfreq 55%

answer

  1. push = source calls you
  2. poll = platform calls source for you
  3. API Gateway=push, SQS/Kinesis=poll
  4. poller does batching+checkpointing
  5. backpressure lives with the poller

basics

~20 s

Push: the source (like an API Gateway) calls the function immediately for each request, so it runs right away. Poll: the platform itself checks a source (like a queue) on a timer, gathers waiting messages, then invokes the function with a batch.

solid answer

~50 s

Event source mappings fall into two invocation models. Push sources — API Gateway, S3, SNS — send the event directly to the function's invoke API the moment it occurs; the platform doesn't control pacing, so it must throttle or queue internally if the function can't keep up. Poll sources — SQS, Kinesis, Kafka, DynamoDB Streams — have no way to call a function themselves; instead the serverless platform runs a poller that continuously reads from the source, accumulates a batch, and invokes the function synchronously with that batch. Poll gives the platform natural backpressure control (it won't fetch faster than it can process) and enables batching for throughput, at the cost of added latency and polling infrastructure. Push gives lower latency and a simpler mental model, but the invocation rate is dictated entirely by the caller, so protection against overload has to live downstream.

go deeper

for a junior

Should correctly state which model each named source (HTTP vs queue) uses and describe the basic behavior difference (immediate call vs batched call) without needing to explain backpressure or checkpointing internals.

for a middle

Should additionally explain why the distinction exists (queues can't call out, HTTP has a waiting client) and name at least one operational implication, such as concurrency throttling on push or batching on poll.

for a senior

Should discuss backpressure, checkpointing/visibility-timeout mechanics, and failure modes like poison-pill blocking or DLQ-less event loss, and be able to map a given source (Kinesis, S3, EventBridge schedule) to the correct model unprompted.

for a principal

Should reason about system-level design consequences — where to place push vs poll boundaries in a pipeline, how to size concurrency and batch settings against downstream capacity, and how the choice affects end-to-end latency and failure-isolation guarantees across a multi-hop architecture.

## What an event source mapping is An **event source mapping** is the configuration that connects an event-producing service to a serverless function so that an event automatically triggers an invocation, without the developer writing any polling or subscription code themselves. Under the hood, every such mapping falls into one of two invocation mechanics: **push** or **poll**, and the mechanic is chosen by the source type, not by the developer. | Mechanic | Who calls whom | Named sources | |---|---|---| | Push | the event source calls the function | `API Gateway`, `S3`, `SNS` | | Poll | the platform reads the source | `SQS`, `Kinesis`, `Kafka`, `DynamoDB Streams` | ## The push model In the push model, the event source calls the function's invoke API directly the instant something happens. - **API Gateway** is the canonical example: when a client sends an HTTP request, API Gateway synchronously invokes the backing function, waits for the response, and streams it back to the client. - **S3 object-created notifications and SNS topic deliveries** work similarly but asynchronously — the source fires the event at the function and moves on without waiting for a reply, relying on the platform's internal retry queue if the invocation fails. In every push case, the source itself decides exactly when and how often the function is called; the serverless platform has no say in pacing beyond enforcing concurrency limits and throttling at the door. ## The poll model In the poll model, the source is a store the function cannot call out of — a queue like SQS, a log like Kinesis or Kafka, or a change stream like DynamoDB Streams. These sources are passive: they hold records until something reads them. So the serverless platform runs its own long-lived poller process (invisible to the developer) that: 1. continuously calls Receive/GetRecords-style APIs against the source, 2. accumulates records into a batch up to a configured size or time window, 3. and only then synchronously invokes the function with that batch as the payload. The poller is also responsible for checkpointing progress (deleting consumed SQS messages, advancing a Kinesis shard iterator) once the function returns successfully. ## Why the split exists This split exists because event sources have fundamentally different delivery semantics, and the platform has to bridge each one to a uniform 'invoke a function with a payload' contract. A queue or stream is designed to buffer and be read at the consumer's pace — there's no mechanism for it to reach out and call code. An API Gateway request, by contrast, has a live client on the other end waiting synchronously; buffering it centrally would defeat the purpose. Rather than force every function author to write bespoke polling loops for queue-based work, the platform absorbs that complexity into the poller component and exposes the same function-invocation interface everywhere. ## The trade-off The trade-off is between control and simplicity. - **Push** gives the lowest possible latency — no polling interval to wait out — and a simple mental model of 'one event in, one invocation out,' but it hands invocation-rate control to something outside the platform's process boundary. If ten thousand requests land in a second, the platform must throttle concurrently-running invocations and either queue or reject the excess; a burst that outruns configured concurrency shows up to the caller as 429s or 5xxs. - **Poll**, conversely, gives the platform itself the throttle: it will never pull records faster than the function can be invoked, so there's natural backpressure and the batch size can be tuned to trade latency for throughput. The cost is added latency (records sit until a batch fills or a timeout elapses) and a layer of polling infrastructure whose own health — poller concurrency, iterator age, empty-receive backoff — becomes something to monitor. ## Failure modes Failure modes differ correspondingly. - **On the push side**, an unconfigured or undersized concurrency limit causes silent event loss for asynchronous push sources unless a dead-letter queue is attached, because the source only retries a bounded number of times before giving up. - **On the poll side**, a single malformed or perpetually-failing record can stall an entire batch — and, for ordered sources like Kinesis, an entire shard — because the poller keeps re-delivering the same unconsumed batch instead of advancing past it (the so-called poison-pill problem). Iterator age or approximate-number-of-messages-visible metrics climbing steadily is the classic sign that a poller is falling behind or stuck. ## A concrete illustration A concrete illustration: a typical AWS Lambda-based order-processing pipeline uses API Gateway (push) to accept the initial 'place order' HTTP call synchronously, then drops a message onto SQS for asynchronous fulfillment work picked up by a second function via poll. The first hop needs sub-second, synchronous, per-request latency, so push is correct; the second hop benefits from batching, buffering against downstream slowness, and backpressure, so poll is correct. Choosing the wrong model — say, trying to synchronously push high-volume fulfillment events — would either overwhelm concurrency limits or lose events with no buffer to absorb bursts.

  • If a push-based asynchronous trigger like an S3 event fails every retry, what happens to the event, and how do you prevent silent loss?
    Asynchronous push invocations retry a limited number of times (twice by default in AWS Lambda) with backoff, and if all retries fail the event is dropped unless a dead-letter queue or a failure-destination is configured on the function. Without one, the event and its payload are gone with only a CloudWatch metric or log line as evidence. Attaching a DLQ/on-failure destination is the standard mitigation so failed events can be inspected and reprocessed.
  • Why does a poll-based trigger need a visibility timeout or its equivalent, and what happens if it's set too short?
    The poller has to hide a record from other consumers while a function is processing it, otherwise a second poller cycle would fetch and process the same record concurrently, causing duplicate side effects. If the visibility/lease timeout is shorter than the function's actual processing time, the record becomes visible again before processing finishes and gets redelivered, producing duplicate processing even though the first invocation eventually succeeds.
  • How does concurrency scale differently between a push source like API Gateway and a poll source like SQS?
    Push scales concurrency roughly 1:1 with inbound request rate, up to the function's concurrency limit, so a traffic spike directly spikes concurrent invocations. Poll scales concurrency based on how many pollers the platform assigns (bounded by queue depth and a configurable maximum), so it ramps more gradually and is capped independently of how fast messages arrive, giving the operator a knob to limit downstream load.

Push is like a doorbell — whoever's outside decides when it rings and you have to answer right then. Poll is like a mailbox with a mail carrier (the platform) who checks it on a schedule, bundles up whatever's arrived, and hands you the stack at once.

saying these in an interview costs you the question

  • Says push and poll are just synonyms for sync vs async invocation
  • Claims every event source can be configured to use either model interchangeably
  • Doesn't know that poll-based sources need explicit checkpointing/deletion to avoid redelivery
  • Assumes push sources automatically retry forever with no event loss possible
  • Thinks the developer's function code has to implement the polling loop itself

context

open as a page

When an API Gateway is configured as an HTTP trigger for a serverless function, the invocation is synchronous end-to-end: the gateway waits for the function's response before replying to the client. What practical constraints does this synchronous push model impose on the function, and what happens if the function takes too long or errors out?

level: middleimportance: must knowfreq 60%

basics

~20 s

The client is waiting live, so the function has to answer fast — there's a hard time limit shorter than the function's own max runtime. If it's too slow or crashes, the gateway just returns an error straight to the client; nothing retries automatically.

open as a page

A function is triggered by an SQS queue with a batch size of 10. If 7 of the 10 messages in a batch process successfully but 3 throw errors, what happens to each message by default, and how does 'partial batch response' (reporting individual item failures) change that behavior?

level: middleimportance: must knowfreq 65%

basics

~20 s

Normally, if any message in the batch fails, the whole batch of 10 gets retried later — including the 7 that already succeeded, so they run twice. Partial batch response lets the function tell the queue exactly which 3 failed, so only those get retried and the other 7 aren't repeated.

open as a page

A serverless function is triggered by a Kinesis (or Kafka) stream with several shards/partitions. One shard has a record that always throws an exception when processed. What happens to that shard's throughput while the bad record isn't skipped or fixed, and why doesn't it affect the stream's other shards?

level: seniorimportance: must knowfreq 55%

basics

~20 s

That one shard gets stuck: every batch keeps re-trying starting from the same bad record, so new records behind it never get processed — like a traffic jam on one lane. Other shards each have their own separate reader, so they keep moving fine; the jam doesn't spread.

open as a page

Object-storage triggers, like S3 event notifications invoking a function on every object upload, are push-based and deliver events with at-least-once semantics. What does 'at-least-once' mean operationally for a handler, and what other delivery quirks (ordering, timing) does this trigger type have that a handler needs to account for?

level: seniorimportance: should knowfreq 45%

basics

~20 s

At-least-once means the same upload event might trigger your function more than once, so it can't assume it only runs exactly one time per file. Events also might not arrive in the exact order files were uploaded, and there can be a short delay before the event fires.

open as a page

A scheduled (cron-style) trigger fires a serverless function every 5 minutes to reconcile a data feed, and each run typically takes about 4 minutes. What can go wrong if the schedule doesn't account for overlapping executions or missed invocations, and how would you design around it?

level: principalimportance: should knowfreq 35%

basics

~20 s

If a run takes almost as long as the gap between runs, sometimes the next run starts before the last one finishes, and now two copies are working on the same data at once, which can cause conflicts or double work. You need a way to stop overlapping runs and a plan for what happens if a run is skipped entirely.

open as a page