When an API Gateway is configured as an HTTP trigger for a serverless function, the invocation is synchronous end-to-end: the gateway waits for the function's response before replying to the client. What practical constraints does this synchronous push model impose on the function, and what happens if the function takes too long or errors out?
answer
- synchronous push, gateway blocks
- gateway timeout < function max timeout
- no retry/DLQ on sync errors
- cold start = visible latency
- split slow work into async job + poll
basics
~20 sThe client is waiting live, so the function has to answer fast — there's a hard time limit shorter than the function's own max runtime. If it's too slow or crashes, the gateway just returns an error straight to the client; nothing retries automatically.
solid answer
~50 sAPI Gateway-to-function integration is a synchronous, push-based RPC: the gateway opens a connection, invokes the function, and blocks until it gets a response to relay back to the client. This imposes an end-to-end timeout well below the platform's max function timeout — for example AWS's REST API Gateway caps integration latency at 29 seconds regardless of a function's configured timeout of up to 15 minutes — so any function meant to run behind an HTTP trigger must either finish within that window or return quickly with a 'processing started' response and do the real work asynchronously. If the function errors, times out, or the concurrency limit is exhausted, the gateway gets no successful response and returns an error (5xx/429) directly to the caller; there is no automatic retry, no DLQ, and no buffering, so error handling and retries become the client's or the caller's responsibility, not the platform's.
go deeper
Should know the function has to respond within some time limit because a client is waiting, and that a slow or crashed function results in an error being returned to the client.
Should name the specific mismatch between the gateway's integration timeout and the function's own max execution time, and describe the basic async-job workaround (return quickly, do work in the background).
Should articulate the full trade-off — no retry/DLQ on the synchronous path, throttling surfacing as client-visible errors, cold starts inflating tail latency — and design the split between synchronous and asynchronous work appropriately.
Should reason about this at a system-architecture level: setting concurrency/throttling policy to protect downstream systems without degrading client experience, choosing between sync API + async job patterns vs. websockets/long-polling for genuinely long operations, and the SLO implications of tying a client connection to a serverless invocation lifecycle.
## How the integration works An HTTP API Gateway trigger wires a serverless function into the request path of an HTTP endpoint using a **synchronous, push-based integration**. Mechanically, when a client sends a request, the gateway: 1. terminates the HTTP connection, 2. maps the request into an event payload (headers, path, query string, body), 3. and calls the function's synchronous invoke API. The gateway then blocks — holding the client's connection open — until the function returns a response, which it maps back into an HTTP response and streams to the client. Nothing is queued or buffered in between: the function is invoked exactly once per inbound request, and the gateway itself becomes a live, in-line hop in the request path rather than a decoupled event emitter. ## Why HTTP demands push This design exists because HTTP is inherently a request/response protocol with a client waiting for an answer. Unlike a queue or object-storage event, there is no acceptable place to buffer the work — the whole point is to give the caller a timely, synchronous answer. Push invocation is the natural fit: the gateway must call the function the instant the request arrives, and it must get a response back before it can reply, so there's no room for a poller accumulating a batch over a time window the way SQS or Kinesis triggers do. ## The latency ceiling The synchronous contract imposes a hard latency ceiling that is often tighter than the function platform's own maximum execution time. On AWS, for instance, REST API Gateway enforces a **29-second integration timeout** on every backend call, even though a Lambda function itself can be configured to run up to 15 minutes; if a function is invoked via API Gateway but takes longer than 29 seconds, the gateway gives up waiting and returns a 504 to the client regardless of whether the function eventually succeeds. This forces an architectural split: - **Work that can genuinely complete inside the gateway's timeout window** can live directly behind the HTTP trigger. - **Anything longer-running** — report generation, batch processing, complex ML inference — has to be restructured so the HTTP-triggered function does the minimal work of validating the request and kicking off the real job asynchronously (for example, dropping a message on a queue or starting a state-machine execution), then immediately returning a 202-style acknowledgment with a job ID the client can poll. ## The trade-offs The trade-offs of this model cut both ways. - **On the upside**, synchronous push gives the lowest possible latency for the class of work it fits — there's no polling interval, no batch-accumulation delay, and the client gets an immediate, direct answer, which is exactly what a typical REST API needs. - **On the downside**, there is zero built-in resilience: if the function errors, throws an unhandled exception, or is throttled because the account or function concurrency limit is exhausted, the gateway has nothing to retry against and simply surfaces an error straight to the client (a 500 for an unhandled function error, a 429 or 503 for throttling). Unlike SQS or Kinesis triggers, there is no dead-letter queue capturing the failed event for later inspection, because there was never a persisted event to capture — the request/response happened live and, once failed, is gone unless the client itself retries. ## Failure modes in production This produces distinctive failure modes in production. - **A slow downstream dependency** (a database under load, a third-party API with high latency) shows up not as a queued backlog but as client-visible timeouts and 504s, because there's no buffer absorbing the slowness. - **A traffic spike that exceeds reserved concurrency** shows up immediately as a wave of 429/503 responses to real users, rather than a growing but invisible queue depth metric. - **Cold starts** — the extra latency of spinning up a fresh execution environment — are directly visible to end users as elevated tail latency (p99 spikes), because each invocation is tied to a live client connection. That is a much bigger user-facing concern for HTTP triggers than for queue or stream triggers where a few hundred milliseconds of extra poller latency is invisible. ## A concrete scenario A concrete scenario: a checkout API backed by a Lambda function behind API Gateway validates payment details and creates an order record, both fast operations that comfortably finish in under a second, so the synchronous push model works well. If that same team later tries to have the checkout function also synchronously call a slow third-party fraud-scoring service that occasionally takes 40 seconds, they'll start seeing 504s at the gateway even though the Lambda function itself hasn't hit its own timeout — the correct fix is to move the fraud check off the synchronous path entirely, either into an asynchronous post-checkout step or a queue-backed worker, rather than trying to lengthen the gateway's timeout, which for many platforms isn't adjustable past a fixed ceiling.
- How would you redesign an HTTP-triggered function that needs to do 3 minutes of work, given the gateway's much shorter integration timeout?Split it into two hops: the HTTP-triggered function validates the request, writes a job record, and enqueues the real work (to SQS, a state machine, or a background worker), returning a 202 with a job ID in well under a second. A separate poll-based or asynchronous function does the actual 3 minutes of work, and the client polls a status endpoint or receives a webhook/notification when it's done.
- Why do cold starts matter more for an HTTP-triggered function than for a queue-triggered one?An HTTP trigger's latency is directly visible to a waiting end user as part of the response time, so a cold start directly inflates p99/p999 client-facing latency. A queue-triggered function's extra startup latency just delays when a message finishes processing in a backlog that's already asynchronous and invisible to any live client, so the same delay has far less user impact.
- If the function behind an API Gateway trigger throws an unhandled exception, what does the client actually receive, and what's missing compared to a queue-triggered failure?The client typically receives a 500-class error response directly and immediately, with no automatic retry from the platform. Compared to a queue trigger, there's no redelivery, no DLQ, and no batch-item-failure mechanism — the event was never durably stored, so once the synchronous call fails, recovering it depends entirely on the client retrying or the caller having captured the request elsewhere.
It's like calling a customer-service line and staying on hold: the agent (function) has to pick up and answer while you're still on the line, and if they take too long the call just drops — nobody automatically calls you back.
saying these in an interview costs you the question
- Believes the function can run for its full configured max timeout regardless of the gateway
- Assumes API Gateway automatically retries failed synchronous invocations
- Doesn't know a long-running task needs to be split into an async job pattern
- Thinks cold starts don't matter for HTTP-triggered endpoints
- Confuses HTTP triggers with SQS triggers when discussing DLQs or batch retries