A steady internal API serves a roughly constant 200 requests per second around the clock and currently runs on AWS Lambda. How would you reason about whether moving it to ECS on Fargate is cheaper, and what must the comparison include?
answer
- compare curve shapes, not unit prices
- duty cycle decides the crossover
- billed wall-clock time versus overlapped waits
- front door and NAT belong in the total
- migration cost can exceed the saving
basics
~20 sCompare cost curves, not sticker prices. Lambda's bill grows with requests times duration and never amortizes; a container's bill is set by provisioned capacity and time, so it wins at high, steady duty cycle. Include front-door, networking and engineering costs.
solid answer
~60 sThe two models have different *shapes* of cost, so the question is where the curves cross. Lambda charges per request plus GB-seconds of wall-clock duration, so at constant traffic the bill scales linearly with volume forever and idle capacity is never paid for — which is worthless here, because there is no idle. Fargate charges for provisioned vCPU and memory per second whether or not requests arrive, so its cost is set by the peak capacity you must keep running; at a high duty cycle that fixed cost is spread across a very large number of requests. A steady 200 rps with a duty cycle near 100% is the classic case where containers win. But the comparison has to be total: the front door (per-request API Gateway pricing versus an hourly load balancer), NAT and data-transfer charges the container path may add, logging volume, headroom you must run to absorb scaling lag, and the engineering time to migrate and then operate. Compute Savings Plans can cover both, so apply them to both sides or neither.
code
python · 25 lines# Illustrative comparison. Substitute current published AWS prices
# and figures measured from your own bill before trusting any output.
RPS = 200
AVG_SECONDS = 0.100
MEMORY_GB = 0.512
HOURS_PER_MONTH = 730
PRICE_PER_REQUEST = 0.0000002 # placeholder
PRICE_PER_GB_SECOND = 0.0000166667 # placeholder
PRICE_PER_VCPU_HOUR = 0.04048 # placeholder
PRICE_PER_GB_HOUR = 0.004445 # placeholder
requests = RPS * 3600 * HOURS_PER_MONTH
gb_seconds = requests * AVG_SECONDS * MEMORY_GB
per_invocation_cost = requests * PRICE_PER_REQUEST + gb_seconds * PRICE_PER_GB_SECOND
# Concurrency in flight, then a headroom factor for burst and scaling lag.
concurrency = RPS * AVG_SECONDS
vcpus = max(1.0, concurrency / 20) * 2.0
memory_gb = vcpus * 2
provisioned_cost = (vcpus * PRICE_PER_VCPU_HOUR + memory_gb * PRICE_PER_GB_HOUR) * HOURS_PER_MONTH
print(f"per-invocation model: {per_invocation_cost:,.2f}")
print(f"provisioned model: {provisioned_cost:,.2f}")
print("front door, NAT, logging and migration effort are NOT included above")go deeper
Know that Lambda bills per request and per duration while a container bills for capacity over time, so the busier and steadier a service is, the more containers tend to favour it.
Explain the crossover in terms of duty cycle, and show how to derive required concurrency from throughput and latency before sizing the container side.
Insist on a total comparison — front door, NAT and data transfer, logging, headroom, commitments — and be ready to recommend not migrating when the saving is smaller than the effort.
Frame it as portfolio economics: a second compute model carries fixed organisational cost in tooling and on-call, so a per-service saving must be weighed against the platform surface it commits the company to.
## Two cost shapes, not two prices The mistake candidates make is comparing unit prices. The right mental model is two curves plotted against *utilised capacity over time*. **Lambda** is a variable-cost model: a per-request charge plus a duration charge in GB-seconds — configured memory multiplied by billed wall-clock time, at millisecond granularity (as of 2025). Nothing is charged when nothing runs. The curve passes through the origin and rises linearly with volume forever; there is no volume at which a request becomes free. **Fargate** is a fixed-cost model: you pay for the vCPU and memory reserved by running tasks, per second, with a one-minute minimum. The curve starts above zero the moment a task exists and is flat with respect to request volume, until you need another task. So the crossover is not about requests per second in the abstract — it is about **duty cycle**, the fraction of paid capacity actually doing work. At 5% duty cycle the container is billed 20× for the work it does. At 95% it is billed almost exactly for work done, and the per-invocation surcharge on the other model becomes the expensive part. ## Doing the arithmetic honestly To size the container side you need concurrency, which follows from throughput and latency: roughly `concurrency = requests_per_second × average_seconds_per_request`. At 200 rps with a 100 ms average, about 20 requests are in flight at any moment — a small number of tasks, especially since one container process serves many concurrent requests while they wait on I/O. That last clause is where the models diverge most sharply. Lambda bills the *whole* duration of every invocation, including the time the handler sits blocked on a database. A container serving the same traffic overlaps those waits on capacity it has already paid for. An I/O-heavy handler is therefore the worst case for per-invocation billing and the best case for containers. Have both sides measured, not modelled: pull actual invocation count and billed duration from the current bill, and size the container side from a load test rather than a guess. ## What must be in the comparison A compute-only comparison is the classic incomplete answer. Add: - **The front door.** Sitting an API behind a per-request-priced gateway costs proportionally to volume; an hourly-priced load balancer costs the same at 20 rps as at 2,000. At high steady volume the front-door line can rival the compute line, and it may change independently of the compute choice. - **Networking.** A container in a private subnet reaching the internet through a NAT Gateway pays hourly *and* per-GB processed. VPC endpoints exist partly to avoid that. Cross-AZ traffic is chargeable too. - **Logging and telemetry.** More capacity usually means more log lines and more custom metrics, and ingestion is priced per GB. - **Headroom.** Containers do not scale instantly, so a real deployment runs above the average to absorb bursts and scaling lag. Budget the headroom, not the theoretical minimum. - **Commitments.** Compute Savings Plans apply to EC2, Fargate and Lambda usage, so applying a discount to one side of the comparison and list price to the other is dishonest arithmetic. - **Engineering time.** Migration is a project, and the ongoing operational load differs. If the saving is a few hundred dollars a month and the migration is six engineer-weeks, the answer is no regardless of the curves. ## The answer that earns the point "Containers are cheaper at scale" is a slogan; the interviewer wants the condition attached. Say it as: *high, steady duty cycle plus long or I/O-blocked handlers moves the crossover toward containers; spiky, low-duty-cycle traffic with short handlers keeps it on functions.* Then say what you would actually do first — take the current bill, split it into request charges and duration charges, size the container fleet from a load test, add front door and networking to both columns, and only then decide. And be willing to conclude "stay where we are", because a migration that saves less than it costs to perform is a loss dressed as an optimisation.
- Why does handler latency change the comparison so much?Because per-invocation billing charges wall-clock duration, including time blocked on a downstream call, while a container overlaps those waits across many requests on capacity already paid for. Doubling average latency roughly doubles the duration charge on one model and changes the other hardly at all.
- The migration would save 15% of a small bill. What do you recommend?Usually staying put. Weigh the saving against migration effort and the ongoing cost of operating a second compute model, and compare it with cheaper levers on the current model — reducing handler duration, right-sizing memory, or trimming log volume — which often recover a similar percentage for far less work.
- Does a Compute Savings Plan change which model wins?Not by itself, because it can cover EC2, Fargate and Lambda usage. It shifts both curves down, so the crossover moves only if the discount rates differ between the two. The real risk is comparing discounted capacity against list-price functions and calling the result an insight.
saying these in an interview costs you the question
- Compares only compute prices and ignores the front door and networking
- Says serverless is always cheaper regardless of volume
- Applies a Savings Plan discount to one side of the comparison only
- Sizes the container fleet from average load with no headroom
- Forgets that Lambda bills time spent waiting on downstream calls