skip to content

Cost & Scaling Model

You pay per request and per GB-second, scale to zero when idle, and hit concurrency limits when busy. You will learn to reason about when that beats an always-on instance and when sustained load makes serverless the more expensive option.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

6

In serverless computing, what does 'pay-per-invocation' pricing mean, and what are the two components (requests and GB-seconds) that make up the bill for a function like AWS Lambda?

level: juniorimportance: must knowfreq 70%

answer

  1. requests + GB-seconds
  2. no traffic = $0
  3. GB-seconds = memory x duration
  4. billed in fine-grained increments (ms)
  5. more memory can mean shorter duration

basics

~20 s

You pay for two things: how many times your function ran (per request), and how much compute it used each time (memory allocated multiplied by how long it ran, in GB-seconds). No traffic means no bill.

solid answer

~40 s

Serverless compute is billed on two independent meters: a fixed price per invocation (per request), and a price per GB-second of compute consumed, where GB-seconds = memory allocated to the function (in GB) x wall-clock execution duration (in seconds), typically rounded up to a small billing increment like 1 millisecond. There's no charge for idle capacity — if the function isn't invoked, you pay only the platform's baseline (often zero). Total monthly cost is roughly (invocation count x price-per-request) + (total GB-seconds across all invocations x price-per-GB-second). Cost scales directly with two levers you control: how much memory you allocate (which also affects CPU/network throughput on most platforms) and how long the function actually runs, so reducing either lowers the bill, while reducing memory too far can increase duration and net cost.

go deeper

for a junior

Should describe the two cost components in plain language and know that idle time costs nothing.

for a middle

Should be able to compute an example cost given price-per-request, price-per-GB-second, memory, and duration.

for a senior

Should reason about tuning memory allocation to minimize the cost/latency trade-off and know the cost curve versus memory isn't monotonic.

for a principal

Should model aggregate cost across a portfolio of functions with varying memory/duration profiles and connect it to broader capacity and architecture decisions.

## The two meters Serverless **function-as-a-service** platforms bill using two independent meters instead of a flat hourly server rate. | Meter | What the platform charges for | |---|---| | **The first meter** — a per-invocation charge | Every time the function is triggered — by an HTTP request, a queue message, a scheduled timer, a storage event — the platform counts one request and charges a small fixed amount for it, regardless of how long that invocation takes. | | **The second, usually larger, meter** — compute-time | The platform tracks how long each invocation ran (wall-clock duration, from when your handler code starts executing to when it returns or errors) and multiplies that by the amount of memory you configured for the function, producing a unit called a `GB-second` (one GB-second is one gigabyte of allocated memory held for one second of execution). | Duration is typically measured and billed at fine granularity — historically rounded up to the nearest 100ms, and on modern platforms often to the nearest millisecond — so short, fast functions are billed close to their true cost rather than padded to a coarse increment. The total bill is the sum of the request charge across all invocations plus the GB-second charge across all invocations' actual compute consumption. ## Why the model is shaped this way This pricing model exists because the core promise of serverless is that you should pay for **work done**, not for **capacity reserved**. - A traditional server or container is billed by the hour or month regardless of how busy it is. - A serverless platform instead meters the exact resources a specific invocation consumed and charges only for that, down to fractions of a second. This only works because the platform can allocate and deallocate execution capacity per-invocation behind the scenes, which is also what enables **scaling to zero**: an idle function has zero active GB-seconds and therefore costs literally nothing beyond any negligible fixed per-function fee the provider might charge (many charge none). ## The trade-off The trade-off is that this fine-grained metering comes with a **per-unit price premium** compared to renting raw compute by the month. A GB-second of serverless compute costs meaningfully more than the equivalent slice of a reserved virtual machine's capacity, because the provider is pricing in the elasticity, multi-tenant isolation, and on-demand availability it has to maintain to let thousands of customers burst unpredictably. The cost side named explicitly: - **You gain** zero cost for idle time and no capacity planning. - **At the expense of** a higher rate per unit of actual compute used — a trade-off whose value depends entirely on how much idle time you'd otherwise be paying for on an always-on alternative. ## Failure modes 1. **Treating memory as a performance knob.** A very common failure mode is treating memory allocation as a knob that only affects performance, not cost, and over-provisioning it 'to be safe.' Because GB-seconds are memory times duration, doubling a function's memory doubles its cost for the same duration; the only way that's cost-neutral or beneficial is if the extra memory (which usually comes bundled with proportionally more CPU on these platforms) meaningfully shortens the duration, so the two roughly cancel out or the net cost even drops. Teams that set memory generously without measuring the actual duration impact routinely end up paying two or three times more than necessary for the same workload. 2. **Underestimating GB-seconds for CPU-bound work.** A second failure mode is underestimating GB-seconds for CPU-bound work: a function doing heavy computation on too little memory (and therefore too little allotted CPU) can run so much longer that its GB-second cost ends up higher than a better-tuned, higher-memory configuration would have produced — the cost curve versus memory is not monotonic, it typically has a minimum somewhere in the middle that has to be found empirically. ## Where it shows up A concrete, widely recognized example is **AWS Lambda**: its documented pricing model is exactly this two-meter structure — a price per million requests, plus a price per GB-second of duration, with duration billed at 1ms granularity. Teams commonly right-size Lambda functions by running the same workload at several memory settings and comparing both latency and total GB-second cost, since a function configured with more memory sometimes finishes fast enough that its GB-second bill is actually lower than a low-memory configuration that runs proportionally much longer — a nightly cost-tuning exercise that pure server-based billing never required because you paid the same flat rate no matter how efficiently the code ran.

  • Why does the platform round duration up to a billing increment like 1ms instead of billing to the nanosecond?
    Metering at infinite precision would add overhead to every single invocation for negligible pricing accuracy gain, so platforms pick a small increment (historically 100ms, now often 1ms on modern platforms) that keeps the bill close to true usage while keeping the metering system itself cheap and simple to operate.
  • If a function's memory allocation is increased, why might its GB-second cost actually increase even though wall-clock latency decreases?
    GB-seconds equal memory times duration, so if you double memory but duration only drops by, say, 30%, the net GB-second figure still rises; more memory only lowers cost when the resulting speedup is proportionally larger than the memory increase itself, which isn't guaranteed for I/O-bound code that isn't CPU- or memory-constrained in the first place.
  • How does a synchronous HTTP-triggered function's billing differ from an asynchronous queue-triggered one?
    Both are billed the same way per invocation on GB-seconds and request count, but an async invocation that fails and gets automatically retried by the platform generates additional billed invocations for each retry attempt, so a flaky downstream dependency can quietly multiply the bill for an async-triggered function in a way a synchronous caller managing its own retries would not.

Like a taxi meter versus a leased car: a taxi (serverless) only charges for the exact ride you take, based on distance and time, while a leased car (an always-on server) costs the same monthly fee whether you drive it constantly or leave it parked.

saying these in an interview costs you the question

  • Believes serverless has no cost until some usage threshold is crossed
  • Confuses GB-seconds (compute-time x memory) with GB of storage
  • Doesn't know idle time between invocations is unbilled
  • Assumes serverless pricing is a flat monthly fee
  • Ignores that increasing memory allocation raises GB-second cost even if duration barely changes

context

open as a page

What is a 'concurrency limit' in a serverless compute platform, and what happens to incoming requests when a function's executions hit that limit at the same moment?

level: middleimportance: must knowfreq 65%

basics

~10 s

It's a cap on how many copies of your function can run at once. If more requests arrive than the cap allows, the extra ones get rejected or queued/retried instead of running immediately.

open as a page

When does a serverless (pay-per-invocation) compute model come out cheaper than running an always-on server or container for the same workload, and when does it flip to being more expensive?

level: seniorimportance: must knowfreq 75%

basics

~20 s

Serverless wins when traffic is spiky or low-volume, since you pay nothing while idle. An always-on server wins when traffic is steady and high, since its flat rate beats serverless's per-request premium at high utilization.

open as a page

In a serverless platform's concurrency model, what is the difference between 'reserved concurrency' (a dedicated slice of an account's shared concurrency pool set aside for one function) and 'unreserved concurrency' (the shared pool available to all other functions), and what problem does each solve?

level: middleimportance: should knowfreq 55%

basics

~10 s

Reserved concurrency carves out a guaranteed, capped slice of the shared pool for one function only. Unreserved concurrency is the leftover shared pool every other function draws from.

open as a page

What does 'scaling to zero' mean in a serverless platform, and what latency and cost trade-off does it introduce via cold starts?

level: seniorimportance: should knowfreq 60%

basics

~20 s

When idle, the platform fully shuts down your function's instances, so you pay nothing. The catch: the next request after idle has to wait for a fresh instance to start up — that delay is a cold start.

open as a page

You're designing cost and capacity governance for an organization running dozens of serverless functions that share a single account-level concurrency limit and several downstream dependencies with their own capacity ceilings (e.g., a relational database, a rate-limited third-party API). What's your approach to preventing one function's traffic pattern from silently degrading cost or reliability for the rest of the account?

level: principalimportance: should knowfreq 35%

basics

~10 s

Give critical functions a guaranteed slice of shared capacity, cap risky/bursty ones so they can't eat everything, protect fragile downstream systems like databases with pooling or their own limits, and monitor the shared pool.

open as a page