skip to content

AWS Lambda

Lambda is the function-as-a-service tier: event triggers, managed runtimes, IAM execution roles, cold starts, and concurrency limits. Interviewers use it to see whether you can name the constraints — duration, package size, statelessness — instead of treating serverless as free scaling.

part ofAWSoverview, primer and where to startread it →
on this pageshow

explore

questions

29

What is a cold start in AWS Lambda, and why does the same function usually respond faster on the next invocation a moment later?

level: juniorimportance: must knowfreq 85%

answer

  1. new sandbox before your code runs
  2. two phases, only one repeats
  3. frozen, not destroyed, between calls
  4. Init Duration in the REPORT line
  5. scale-out pays it again

basics

~20 s

A cold start is the extra latency Lambda spends creating a new execution environment for an invocation: downloading the code, starting the runtime, and running your initialization code before the handler runs. A later invocation reuses that environment and skips all of it.

solid answer

~50 s

Lambda runs your code inside an execution environment — a small isolated sandbox. When a request arrives and no idle environment exists, Lambda creates one: it fetches the deployment package, starts the language runtime, and runs everything in your module's global scope, which AWS calls the Init phase. Only then does it call your handler. That whole preamble is the cold start, and you see it in the CloudWatch Logs `REPORT` line as `Init Duration`. After the handler returns, Lambda does not destroy the environment; it freezes it and keeps it around for a while, so the next invocation is routed to the warm environment and only pays the Invoke phase. Cold starts therefore show up on the first request after a deploy, whenever traffic scales out to more concurrent environments, and after an idle period — they hit a small share of requests but land in your tail latency.

go deeper

for a junior

Be able to say plainly that a cold start is the setup time for a new execution environment and that the next call reuses it. Knowing that init code lives outside the handler is enough at this level.

for a middle

Explain the Init/Invoke split and the freeze-thaw behaviour between invocations, and name the concrete triggers: deploys, scale-out, and reclaimed idle environments.

for a senior

Show that you treat cold starts as a tail-latency measurement problem — pull Init Duration out of the logs, quantify the p99 impact, and only then decide whether provisioned concurrency or SnapStart is worth the money.

for a principal

Own the framing question: is this workload latency-sensitive enough that serverless cold starts are a real product risk, or is the right answer a different compute model? Be ready to price warm capacity against re-platforming.

## The unit Lambda actually manages Lambda does not run "a function"; it runs your code inside an **execution environment** — an isolated sandbox (a Firecracker microVM) with your deployment package, the language runtime, a slice of memory, and its own writable `/tmp`. One environment handles **exactly one invocation at a time**. That single sentence explains nearly everything about Lambda scaling and cold starts: to serve two simultaneous requests, Lambda needs two environments; if it doesn't have a spare one, it has to build it, and building it takes time. ## The lifecycle: Init, Invoke, Shutdown An environment moves through three phases. **Init** happens once per environment. Lambda downloads and unpacks your code, starts any extensions, starts the language runtime, and then executes your function's *initialization code* — everything at module/static scope, outside the handler. In Node.js that is the top level of your module; in Python the module body; in Java static initializers and the constructor of your handler class. **Invoke** is the part you normally think of as "the function running": Lambda passes the event to your handler and waits for the response. This repeats for every request the environment serves. **Shutdown** happens when Lambda decides to reclaim the environment; the runtime and extensions get a chance to stop. A **cold start** is an invocation that had to pay for an Init phase first. A **warm start** is one routed to an environment that already exists. ## Freeze and thaw Between invocations Lambda **freezes** the environment rather than destroying it. Execution is suspended: background threads, timers, and pending I/O stop making progress, and they resume (thaw) only when the next event arrives. Two consequences follow. First, work you kicked off but did not await before returning may simply not finish — and may resume, confusingly, during a *later* invocation. Second, everything you built during Init is still in memory when the environment thaws, which is exactly why the second call is fast. ```javascript // runs once per environment (Init) — reused by every later invocation const client = new SomeClient(); export const handler = async (event) => { // runs once per request (Invoke) return client.doWork(event); }; ``` ## When cold starts actually happen - **First request after a deploy.** Publishing new code or changing configuration invalidates existing environments; the next requests all initialize fresh. - **Scale-out.** Traffic rising from 10 to 60 simultaneous in-flight requests means Lambda must stand up roughly 50 more environments, and each of those requests eats an Init. - **After idleness.** Lambda eventually reclaims unused environments. AWS does not publish the idle timeout, so never design as if a warm environment is guaranteed. - **Periodic recycling.** Environments are replaced over time even under steady traffic, for patching and health reasons. Steady, high-volume traffic therefore has a *low percentage* of cold starts, but a spiky or low-volume function can see them constantly. ## How you observe it Every invocation writes a `REPORT` line to CloudWatch Logs. On a cold start that line carries an extra `Init Duration` field: ``` REPORT RequestId: ... Duration: 12.34 ms Billed Duration: 13 ms Memory Size: 512 MB Max Memory Used: 90 MB Init Duration: 430.12 ms ``` Because only a subset of requests are cold, cold starts are a **tail-latency** problem: they barely move your p50 and can dominate p99. Judge them there, not on averages. ## What influences init time, and what to do Init time is dominated by how much work happens before your handler is reachable: runtime startup plus loading dependencies and constructing clients. Runtimes differ substantially — an interpreted or ahead-of-time-compiled runtime typically initializes faster than a JVM that must load and JIT a large classpath. Doing less at Init (lazy-loading rarely used dependencies, importing narrowly rather than pulling in an entire SDK) is the cheapest lever. When a workload cannot tolerate the tail, AWS sells two answers: **provisioned concurrency**, which keeps a pool of environments already initialized and waiting, and **SnapStart**, which restores an environment from a snapshot taken after Init instead of re-running it. Both remove the Init phase from the request path rather than making Init faster. What does **not** work is the folklore fix: a scheduled "ping" that invokes the function every few minutes keeps *one* environment warm and does nothing for the scale-out case, which is where cold starts actually hurt.

  • Does keeping a function warm with a scheduled ping every five minutes solve cold starts?
    It only keeps roughly one environment alive. Cold starts hurt most during scale-out, when Lambda must create many environments at once for concurrent requests, and a single ping does nothing there. It also costs invocations forever. If the latency tail genuinely matters, provisioned concurrency or SnapStart address it directly; a warming ping is a folk remedy that hides the problem at low traffic.
  • Where would you look to tell whether a latency spike was caused by cold starts?
    Compare the `Duration` distribution against the `Init Duration` values in the CloudWatch Logs `REPORT` lines — a Logs Insights query filtering for records that contain `Init Duration` isolates the cold invocations. If your p99 spikes line up with a burst of initialized environments right after a deploy or a traffic ramp, cold starts are the cause; if slow invocations have no `Init Duration`, look at the handler itself.
  • Can two requests ever be handled by the same execution environment at the same time?
    No. An execution environment processes one invocation at a time, which is why concurrency is measured in environments. Two simultaneous requests always need two environments. This is also why in-memory state in a single environment is safe from data races between invocations, but is not shared with any other environment.

saying these in an interview costs you the question

  • Says every invocation starts a brand-new container
  • Thinks cold starts only happen on the very first request ever
  • Claims a scheduled warming ping eliminates cold starts
  • Believes init code runs again on every invocation
  • Judges cold-start impact from average latency instead of p99

context

open as a page

You attach an SQS queue to a Lambda function, yet SQS never pushes anything anywhere. What is the event source mapping, and how do queue messages actually end up in your handler?

level: juniorimportance: must knowfreq 72%

basics

~20 s

An event source mapping is a separate Lambda resource that runs an AWS-managed poller. It reads the queue for you, groups messages into a batch, and invokes your function synchronously with that batch in the event's Records array.

open as a page

An AWS Lambda function needs to write objects to S3, and API Gateway needs to be able to invoke it. Which of Lambda's two permission mechanisms — the execution role or the function's resource-based policy — covers each, and how do they differ?

level: juniorimportance: must knowfreq 78%

basics

~10 s

The execution role says what the function may do, so it grants the S3 write. The function's resource-based policy says who may invoke the function, so it grants API Gateway the lambda:InvokeFunction action.

open as a page

In AWS Lambda, what is the difference between invoking a function with InvocationType RequestResponse and InvocationType Event, and who is responsible for retrying a failure in each case?

level: juniorimportance: must knowfreq 82%

basics

~20 s

RequestResponse is synchronous: the caller waits, receives the function's result or error, and must retry itself. Event is asynchronous: Lambda returns 202 immediately, queues the event internally, and retries a failed invocation on your behalf.

open as a page

How do you work out how many concurrent executions an AWS Lambda function needs, and what does Lambda do to a synchronous invocation once the account's concurrency limit is reached?

level: middleimportance: must knowfreq 68%

basics

~20 s

Concurrency equals invocation rate multiplied by average duration in seconds: 500 requests per second at 200 ms needs about 100 concurrent executions. Beyond the account's per-Region limit Lambda throttles, and a synchronous caller gets HTTP 429 with TooManyRequestsException.

open as a page

An AWS Lambda function invoked asynchronously (InvocationType Event) throws on every attempt. Describe what Lambda does with that event from the moment it accepts it until the event is gone for good.

level: middleimportance: must knowfreq 66%

basics

~20 s

Lambda queues the event, returns 202, then invokes the function. On failure it retries twice by default with delays of minutes, subject to a maximum event age of six hours. After that the event is discarded unless a failure destination or dead-letter queue captures it.

open as a page

In AWS Lambda, raising a function's memory setting changes more than the amount of RAM available. What else does it change, and why can allocating more memory sometimes make a function cheaper rather than more expensive?

level: middleimportance: must knowfreq 60%

basics

~20 s

Memory is Lambda's single performance dial: CPU, and network and disk throughput, are allocated in proportion to it. Because you are billed for gigabyte-seconds, doubling memory on CPU-bound work that then runs in less than half the time lowers the total bill.

open as a page

AWS Lambda lets you deploy a function either as a .zip archive or as a container image. What are the practical differences between the two packaging formats, and how would you choose one for a given function?

level: middleimportance: must knowfreq 62%

basics

~20 s

Both run on the same Lambda execution model; only packaging differs. A .zip is small, quick to deploy and can use layers, but is capped at 250 MB unzipped. A container image allows up to 10 GB from a private ECR repository and carries OS-level dependencies, at the cost of an image build pipeline.

open as a page

In AWS Lambda, what is the difference between reserved concurrency and provisioned concurrency, and when would you configure each?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Reserved concurrency partitions the account's concurrency pool: it both guarantees a function that many slots and caps it there, at no extra cost. Provisioned concurrency pre-initialises a number of environments so they answer without a cold start, and you pay for them hourly.

open as a page

An SQS-triggered Lambda processes batches of ten. When one message fails, the other nine are delivered and processed all over again. Why does that happen, and how do you make only the failing message be retried?

level: seniorimportance: must knowfreq 54%

basics

~20 s

The event source mapping deletes messages only when the invocation succeeds, so any thrown error redelivers the whole batch. Enable ReportBatchItemFailures on the mapping and return a batchItemFailures list of the failed message IDs so only those are retried.

open as a page

An AWS Lambda function worked fine until it was attached to a VPC; now every invocation hangs and times out when it calls an external HTTPS API. Explain the cause and how you would diagnose and fix it.

level: seniorimportance: must knowfreq 62%

basics

~20 s

A VPC-attached function has private addresses only, so outbound internet traffic goes nowhere unless the subnet routes it. Fix it by placing the function in private subnets whose route table sends 0.0.0.0/0 to a NAT gateway, or by using a VPC endpoint for AWS-service calls.

open as a page

An AWS Lambda function's configuration includes a runtime (for example python3.12) and a handler string (for example app.lambda_handler). What does each of those identify in your deployment package, and what does Lambda report if the handler string does not match what you shipped?

level: juniorimportance: should knowfreq 58%

basics

~20 s

The runtime selects the language environment Lambda starts; the handler names the entry point inside the package — the file or module, then the function in it. A mismatch fails during initialization, before any of your code runs.

open as a page

In an AWS Lambda function, why do you create SDK clients and database connections outside the handler, and what kind of state must never be cached there?

level: middleimportance: should knowfreq 62%

basics

~20 s

Code outside the handler runs once per execution environment and its objects survive into every later invocation that environment serves, so expensive clients are built once. Never cache per-request state there — user identity, request context, or anything you must not leak between callers.

open as a page

On a Lambda event source mapping, what do BatchSize and MaximumBatchingWindowInSeconds control, and what actually makes the poller stop collecting and invoke your function?

level: middleimportance: should knowfreq 56%

basics

~20 s

BatchSize is the maximum number of records per invocation and MaximumBatchingWindowInSeconds is how long the poller may keep gathering them. The invoke fires on whichever comes first: the batch is full, the window expires, or the payload reaches Lambda's synchronous 6 MB limit.

open as a page

Is an AWS Lambda environment variable a safe place to store a database password? Explain how Lambda protects environment variables and who is able to read them.

level: middleimportance: should knowfreq 52%

basics

~20 s

Not for a real secret. Lambda encrypts environment variables at rest with KMS, but anyone allowed to read the function's configuration sees the decrypted value, and the value is static, so a rotated password silently goes stale.

open as a page

What actually changes for an AWS Lambda function when you attach it to a VPC, and what network resources does Lambda create to make that work?

level: middleimportance: should knowfreq 58%

basics

~20 s

Attaching a function to subnets and security groups makes Lambda place its traffic on elastic network interfaces inside your VPC. The function then has private IPs only, obeys your security groups and route tables, and loses default internet access.

open as a page

For asynchronous AWS Lambda invocations, what is the difference between an on-failure destination and the older DeadLetterConfig dead-letter queue, and which should you configure on a new function?

level: middleimportance: should knowfreq 44%

basics

~20 s

A dead-letter queue receives only the original event payload and can target SQS or SNS. An on-failure destination receives a richer record with the invocation context and the error response, and can also target EventBridge or another Lambda function. Prefer destinations on new functions.

open as a page

In AWS Lambda, what does attaching a layer actually do to the function's execution environment, and what limits and gotchas come with layers?

level: middleimportance: should knowfreq 55%

basics

~20 s

A layer is a .zip of shared content that Lambda extracts into /opt before the handler runs, where each runtime already looks for libraries. A function may attach up to five layer versions, and their contents still count against the 250 MB unzipped limit.

open as a page

A synchronous AWS Lambda function behind an API starts returning throttling errors during a short traffic spike, even though account-wide concurrency stays far below the quota. What is happening, and how do you fix it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Lambda limits how fast a function adds environments, not just how many it may have. A spike that outruns that ramp is throttled while headroom still exists. Fix it with pre-warmed provisioned concurrency, a queue to absorb the burst, or client backoff.

open as a page

A Kinesis-triggered Lambda has stopped making progress on one shard: the same batch keeps failing and the iterator age climbs for hours. Why does one bad record stall the shard, and which event source mapping settings get it moving again?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A shard is processed in order, so the mapping retries the failing batch rather than skipping ahead, and by default it retries until the records expire from the stream. Bound it with MaximumRetryAttempts and MaximumRecordAgeInSeconds, isolate the record with BisectBatchOnFunctionError, and capture it via the mapping's on-failure destination.

open as a page

An asynchronously invoked AWS Lambda function charges customers, and finance reports occasional double charges even though CloudWatch shows no function errors for those invocations. Explain how a Lambda invocation that succeeded can still run twice, and how you would stop the double charge.

level: seniorimportance: should knowfreq 52%

basics

~20 s

Asynchronous Lambda invocation is at-least-once: the internal queue can deliver the same event more than once even without a handler error, and a handler that completes its side effect but then times out is retried too. The fix is idempotency in the handler, not tighter retry settings.

open as a page

An AWS Lambda function downloads a large object from S3 into /tmp, transforms it, and uploads the result. It passes every test but in production it intermittently fails with "No space left on device", and occasionally hits the function timeout. What is going on, and how would you fix it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Lambda's /tmp defaults to 512 MB and belongs to the execution environment, which is reused across invocations — so leftover files accumulate until a later invocation runs out of space. Fix by deleting files after use, streaming instead of buffering, and raising ephemeral storage if genuinely needed.

open as a page

An SQS-triggered Lambda scales out under load and exhausts the connection limit of the RDS database behind it. How do you cap how much of that queue is processed at once, and why is putting reserved concurrency on the function the wrong lever?

level: principalimportance: should knowfreq 44%

basics

~20 s

Set ScalingConfig MaximumConcurrency on the event source mapping, which tells the poller itself not to exceed that many concurrent invocations. Reserved concurrency instead lets the poller keep receiving messages and be throttled, inflating receive counts and pushing healthy messages toward the dead-letter queue.

open as a page

You are designing a workflow backed by AWS Lambda and must decide whether the caller retries a failed synchronous invoke, or the work is handed to Lambda's asynchronous invocation path so Lambda retries it. How do you make that call, and what does each choice cost you?

level: principalimportance: should knowfreq 38%

basics

~20 s

Decide by who must know the outcome and how long they can wait. Caller-side retry keeps the error visible and the policy yours, but multiplies in-flight load during an outage. Lambda-side asynchronous retry buys durability and backpressure at the cost of a coarse, untunable policy and no response channel.

open as a page

A Lambda function is invoked for every message on a busy SQS queue but discards about 90% of them immediately. How does FilterCriteria on the event source mapping change that, and what happens to the messages that are filtered out?

level: middleimportance: nice to knowfreq 40%

basics

~20 s

FilterCriteria puts event-pattern matching in the AWS-managed poller, so non-matching records never reach the function and are never billed as invocations. For SQS, filtered-out messages are deleted from the queue; for streams, the poller simply advances past them.

open as a page

What is an AWS Lambda function URL, and when would you use one instead of putting API Gateway in front of the function?

level: middleimportance: nice to knowfreq 36%

basics

~20 s

A function URL is a built-in HTTPS endpoint on a single Lambda function, created with an auth type of NONE or AWS_IAM. It invokes the function synchronously with no extra service in the path. Choose it for simple single-function endpoints, webhooks and streamed responses.

open as a page

What does AWS Lambda SnapStart do to reduce cold-start latency, and what must your function code account for when it is enabled?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

SnapStart runs your initialization once when you publish a version, snapshots the initialized execution environment, and restores from that snapshot instead of re-running init. Code must therefore assume anything captured at init — connections, cached credentials, random seeds — may be stale or duplicated after restore.

open as a page

An AWS Lambda handler calls Secrets Manager on every invocation to fetch a database password. What problems does that create, and how would you fix it without hard-coding the secret?

level: seniorimportance: nice to knowfreq 40%

basics

~20 s

Every invocation pays an extra API round trip, an API charge, and a share of the account's Secrets Manager rate quota, so a high-concurrency function throttles itself. Fetch the secret once per execution environment and cache it with a time-to-live.

open as a page

What is an AWS Lambda extension, how does it get into a function and run relative to your handler, and what are extensions typically used for?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

A Lambda extension is companion code that runs inside the same execution environment as your handler, registered through the Lambda Extensions API. External extensions run as separate processes started before the runtime; they ship as a layer with executables under /opt/extensions, or baked into a container image.

open as a page