skip to content

Your team is deploying a new HTTP API on AWS and must choose between AWS Lambda, ECS on Fargate, and EC2 instances. Which decision axes drive that choice, and which way does each one point?

level: middleimportance: must knowfreq 82%

answer

  1. answer with axes, not a service name
  2. duration is the only hard disqualifier
  3. duty cycle decides the cost curve
  4. what does the process hold between requests
  5. who patches it at 3 a.m.

basics

~20 s

Drive the choice from five axes: request duration, traffic shape and duty cycle, statefulness, latency and cold-start tolerance, and how much operations the team can carry. Lambda wins on spiky, short, stateless work; containers win on steady, long-running or stateful work.

solid answer

~40 s

I would not name a service first — I would ask five questions. **How long is a request?** Anything past a few minutes rules Lambda out against its hard timeout. **What is the traffic shape?** Spiky, low-duty-cycle traffic favours Lambda, which costs nothing when idle; steady round-the-clock load favours a container or instance you have already paid for. **Is it stateful?** In-memory caches, long-lived connections and work that continues after the response is returned all point away from Lambda. **How tight is the latency budget, and can it absorb a cold start?** A p99-sensitive path with strict tail latency is easier on a warm container. **What can the team operate?** EC2 buys control at the price of patching and capacity management. Then I name a service, and say what would change my mind.

go deeper

for a junior

Be ready to list the three options and say what kind of workload each suits, and to ask about request duration and traffic pattern before committing to an answer.

for a middle

Explain each axis and which way it points, especially why per-invocation billing favours spiky traffic and per-hour capacity favours steady traffic.

for a senior

Demonstrate judgment by naming what would change your mind and what you would measure to detect it — duty cycle, handler duration, tail latency — rather than defending one model.

for a principal

Own the axis nobody puts on the slide: each compute model the organisation supports needs its own delivery, observability and on-call path, so the choice is partly about how many models you are willing to fund.

## Why interviewers open with this "EC2, containers, or Lambda?" is the AWS equivalent of "SQL or NoSQL?". It is asked because there is no right answer, only a right *method*: a candidate who answers with a service name has guessed, and a candidate who answers with axes has designed. The five axes below are the ones that actually decide real cases. ## Axis 1 — request duration This is the only axis that produces a hard disqualification rather than a tradeoff. A Lambda function has a maximum timeout of 15 minutes (as of 2025), so any single unit of work that can exceed it cannot run in one invocation, full stop. A synchronous HTTP API sitting behind a front door usually has a much tighter effective ceiling anyway, because the client and the front door both time out long before the function does. Duration also drives cost, because Lambda bills the wall-clock time of every invocation. A handler that spends 800 ms waiting on a downstream call is billed for all 800 ms, whereas a container spends that same wait handling other requests on other threads. Long, I/O-blocked handlers are the workload where per-invocation billing is least flattering. ## Axis 2 — traffic shape and duty cycle The useful number is duty cycle: what fraction of the time is your capacity actually doing work? An internal API that gets a burst of traffic during business hours and nothing overnight has a low duty cycle, and every provisioned container hour outside the burst is waste. Lambda's cost is proportional to work done and falls to zero when nothing arrives — that is what "scale to zero" buys. Invert it and the arithmetic inverts. A public API at a steady few hundred requests per second all day has a high duty cycle; a container you keep running is busy essentially all the time, so its fixed hourly cost is spread over a very large number of requests. Burstiness matters separately from volume. Traffic that goes from near zero to a large peak in seconds is exactly what elastic, per-request compute is good at; a container fleet has to be scaled by a policy that reacts to a metric, which takes time and typically means running headroom you are not using. ## Axis 3 — statefulness Ask what the process holds between requests. A warm in-memory cache, a pool of long-lived database connections, a subscription to a broker, an in-process scheduler, background work that continues after the response is written — every one of those assumes a long-lived process, and every one of them is awkward on Lambda, where each concurrent invocation is a separate isolated environment and the environment is frozen once the handler returns. If the answer is "it holds nothing; every request is self-contained", Lambda fits the model naturally. ## Axis 4 — latency budget and cold-start tolerance Every compute model has a first-request penalty; they differ in when you pay it. A container pays it at deploy and scale-out; a function can pay it whenever a new execution environment is created, which is visible in the tail rather than the median. So the question is not "do cold starts exist" but "does this endpoint have a tail-latency commitment tight enough that an occasional multi-hundred-millisecond outlier breaks it?" An internal admin API does not care. A synchronous call on a checkout path might. ## Axis 5 — operational burden and team shape The last axis is not technical. EC2 gives you kernel access, arbitrary protocols, GPUs, local disk and licensed software, and charges you in patching, AMI hygiene, capacity management and on-call surface. Fargate removes the host but keeps the container workflow the team already knows. Lambda removes the most operations and imposes the most constraints on how you write code. A small team with no platform engineers is buying something real when it gives that work away. ## Putting it together Asked cold, a defensible answer sounds like: *"Spiky, low-volume, stateless, no tight tail-latency requirement — Lambda behind an API front door, and I'd revisit it if sustained volume grows or handlers get long. Steady traffic, an existing container build, or state in the process — ECS on Fargate. EC2 when the workload needs the host itself: a GPU, a kernel feature, a licensed agent, or a protocol the managed options do not serve."* Then name the thing that would change your mind. The axes are also the migration triggers: a Lambda API that grows into steady high-volume traffic is a candidate to move to containers, and a container fleet running at 4% utilisation overnight is a candidate to move the other way.

  • Which single axis most often flips a team from Lambda back to containers as a product grows?
    Duty cycle. A workload that started spiky and low-volume becomes steady and high-volume, and per-invocation billing that was almost free at launch becomes the dominant line item. Long handler durations accelerate it, because Lambda bills wall-clock time while a container overlaps waiting requests on the same capacity.
  • Does choosing Lambda mean you never think about capacity again?
    No. You stop choosing instance counts, but you gain concurrency as the thing to reason about: how many simultaneous executions the workload implies, and what downstream systems can absorb that fan-out. Elastic compute pushes the capacity question onto whatever the function calls.
  • Where would you put a workload that is spiky but must not cold-start?
    Usually a container kept warm at a small minimum count, or Lambda with warm capacity provisioned for the latency-sensitive path only. The decision turns on whether the tail-latency commitment is worth paying for idle capacity — which is the same duty-cycle tradeoff seen from the latency side.

saying these in an interview costs you the question

  • Names a service immediately without asking about the workload
  • Says serverless is always cheaper than containers
  • Treats cold starts as a reason never to use Lambda anywhere
  • Ignores that a stateful process cannot simply be split into functions
  • Assumes EC2 is the safe default because it is the most flexible

context