On AWS Lambda, configuring more memory for a function also increases the CPU and network bandwidth it gets, and every function has a maximum execution timeout (15 minutes on Lambda). How do these two limits shape the way you design a FaaS-based system, and what happens when a function exceeds either one?
answer
- Lambda: memory dial also sets CPU/network
- timeout is a hard kill, no graceful finish
- long jobs = decompose via queue/step-function, or leave FaaS
- OOM = immediate fail, no partial output
- design for idempotent retries near the timeout edge
basics
~30 sFaaS functions have two hard caps: how much memory (and therefore CPU) you give them, and how long they're allowed to run before the platform kills them. If a function tries to use more memory than allowed it crashes, and if it runs past its time limit it gets forcibly stopped - so long or heavy jobs must be split into smaller pieces or moved to a different kind of compute.
solid answer
~60 sFaaS platforms cap both memory and execution duration per invocation, and on providers like AWS Lambda, memory and CPU are coupled - you pick a memory size (128MB up to 10GB) and the platform allocates proportional vCPU and network throughput, so a CPU-bound function can get faster purely by requesting more memory, even if it doesn't need the extra RAM. Timeout limits (Lambda: 15 minutes max; Cloud Functions/Azure Functions have their own comparable ceilings) mean any workload whose natural duration can exceed that ceiling - a large batch job, a long-running data export, a slow third-party API call - cannot be modeled as a single synchronous function; it must be decomposed into a chain of shorter invocations (e.g., via a queue or a step-function/orchestrator) or moved off FaaS entirely onto a container or VM. Hitting the memory limit causes an out-of-memory error and invocation failure; hitting the timeout causes the platform to forcibly terminate the invocation mid-execution, which can leave a partial write in a downstream system if you weren't careful about idempotency.
go deeper
Should know that FaaS functions have both a memory cap and a time cap, and that going over either one causes the function to fail/be killed.
Should know that on providers like Lambda, memory and CPU are coupled, and should be able to name the basic strategy for jobs that don't fit the timeout - break them into smaller pieces.
Should be able to design the decomposition concretely - queues, fan-out/fan-in, orchestration tools like Step Functions/Durable Functions - and explain the idempotency/checkpointing concerns that arise from forced mid-execution kills.
Should reason about when FaaS is the wrong compute primitive entirely for a workload class, weigh the operational and cost trade-offs of decomposing versus migrating to containers/VMs, and set organizational guardrails (e.g., alerting when a function's p99 duration approaches its timeout ceiling) to catch this failure mode before it hits production.
## Two resource ceilings Every FaaS platform imposes two resource ceilings on each invocation: a **memory limit** and an **execution time limit**, and understanding how they actually work - not just that they exist - shapes real architectural decisions. On AWS Lambda, memory is the single dial you turn (currently 128MB up to 10,240MB), and CPU allocation is not independently configurable - it scales proportionally with the memory you request, up to a point around 1,769MB where a function gets the equivalent of one full vCPU, and beyond that you get access to multiple vCPUs. This means a CPU-bound function - say, one doing image resizing or JSON parsing over a large payload - can often be made faster simply by increasing its configured memory, even though it doesn't need the extra RAM for data; you're really buying CPU time by way of the memory dial. Network bandwidth is similarly scaled with memory. Google Cloud Functions and Azure Functions have analogous but separately-configured memory and CPU settings depending on the generation/plan (Cloud Functions gen2 lets you set CPU somewhat more independently, and Azure Functions' behavior depends on the hosting plan - Consumption vs. Premium vs. Dedicated). ## The timeout limit The timeout limit is a hard wall-clock cap on how long a single invocation is allowed to run before the platform forcibly kills it: - **AWS Lambda's** ceiling is 15 minutes. - **Google Cloud Functions** (2nd gen, HTTP) allows up to 60 minutes. - **Azure Functions' Consumption plan** defaults to 5 minutes (configurable up to 10) while the Premium/Dedicated plans allow much longer duration. These limits exist because FaaS platforms are built around fast, elastic scaling of huge numbers of short-lived invocations - the billing model, the scheduler, and the whole 'spin up an environment per burst of concurrency' architecture assumes invocations complete quickly, and a platform that let functions run indefinitely would undermine both its cost model and its ability to reclaim resources predictably. ## The architectural trade-off The trade-off these limits impose is architectural, not just a config nuisance: FaaS is fundamentally not the right primitive for anything whose natural runtime is unbounded or can spike past the ceiling - a nightly batch job processing millions of records, a long-lived WebSocket connection, a video transcoding job, or a slow synchronous call to a flaky third-party API that occasionally takes 20 minutes. Two paths exist. 1. One is **decomposition**: break the long job into a chain of short invocations, each processing a bounded chunk of work and either recursively invoking itself with a continuation token, or pushing work items onto a queue (SQS, Pub/Sub, Storage Queues) that trigger the next function, or using an orchestration layer (AWS Step Functions, Azure Durable Functions) that manages state and retries across many short function calls so no single invocation has to run the whole job. 2. The other path is simply choosing a **different compute primitive** - a container on ECS/Cloud Run/AKS, or a plain VM - for workloads that are naturally long-running, and reserving FaaS for the request/response or event-handling pieces of the system. ## The failure modes The failure modes here are sharp and unforgiving in production. - **Exceeding the memory limit** causes an immediate out-of-memory termination of the invocation - no partial results, no graceful degradation, just a failed invocation that (depending on the trigger type) may or may not be retried by the platform. - **Exceeding the timeout** is worse in practice because the platform kills the invocation mid-execution with no warning inside your code (you get, at best, a chance to react to a 'time remaining' check via the context object and try to flush state before the hard cutoff) - if your function had already written half its output to a database or sent some but not all messages to a downstream queue, you can be left with partial, inconsistent side effects. This is why idempotency and checkpointing matter enormously in FaaS design: any function whose duration is close to the timeout ceiling should be designed so a forced kill-and-retry doesn't double-process or corrupt data, typically by making writes idempotent (upserts keyed by a request ID) or by checkpointing progress externally so a retried invocation can resume rather than restart. ## A concrete real-world scenario A concrete real-world scenario: a team originally implements a nightly report-generation job as a single Lambda function that queries a database, aggregates results, and emails a PDF. As data volume grows, the job creeps toward the 15-minute ceiling and starts intermittently timing out mid-aggregation, leaving no report sent and no clear record of how far it got. They refactor it into a Step Functions state machine: one Lambda queries and chunks the work, a Map state fans out many short Lambda invocations to aggregate each chunk in parallel (each well under the timeout), and a final Lambda assembles and sends the report - trading a fragile single long invocation for a bounded, retryable, observable pipeline.
- Why might a team increase a Lambda function's memory setting even though the function's data footprint is tiny and never comes close to using that memory?Because AWS Lambda ties CPU allocation to configured memory, requesting more memory is often the simplest way to get more compute power for a CPU-bound task, cutting execution time (and sometimes even total cost, since faster execution can offset the higher per-GB-second rate). It's a common tuning lever, but teams should benchmark rather than guess, since the cost/speed curve isn't linear.
- What's the risk of a function whose typical execution time is 14 minutes against a 15-minute Lambda timeout, and how would you make it safer?It's living dangerously close to a hard, no-warning termination, where any input larger than usual or any slow downstream dependency pushes it over the edge and kills it mid-execution, potentially leaving partial writes. The safer pattern is to decompose the work into smaller, bounded chunks orchestrated by a queue or Step Functions, and to make each chunk's writes idempotent so a killed-and-retried chunk doesn't corrupt data.
Like renting a parking meter spot instead of a garage: you get a fixed, ticking window of time and a fixed amount of space, and if you go over either the car (your function) gets forcibly towed (killed) mid-errand - so you plan errands (workloads) that fit the meter, or you park somewhere else (a different compute service) for the long ones.
saying these in an interview costs you the question
- Thinks FaaS timeout and memory limits can be worked around by just retrying the same monolithic function forever
- Doesn't know CPU is often tied to memory on providers like Lambda
- Assumes a killed-by-timeout invocation leaves the system in a clean state by default
- Proposes running an inherently long-lived process (e.g., WebSocket server) as a single FaaS invocation without addressing the timeout
- No mention of idempotency when discussing retries near the timeout edge