An AWS Lambda function is billed in GB-seconds. A colleague proposes cutting cost by lowering its memory setting from 1024 MB to 512 MB. Explain when that actually lowers the bill and when it raises it.
answer
- GB-seconds: memory times duration
- CPU scales with the memory dial
- cheaper per millisecond, more milliseconds
- waiting on I/O ignores extra CPU
- sweep the settings, plot the curve
basics
~20 sLambda bills memory multiplied by duration, and CPU is allocated in proportion to memory. Halving memory halves the per-millisecond price but can more than double the runtime of CPU-bound work, raising total cost. It only saves money for functions that mostly wait on I/O.
solid answer
~50 sCost is GB-seconds: configured memory times duration, billed per millisecond, plus a small per-request charge. The catch is that Lambda allocates CPU (and network throughput) *in proportion* to the memory setting, so memory is really the size dial for the whole function. For a **CPU-bound** function — parsing, compression, image work, a cold JVM or .NET start — halving memory halves the price per millisecond but roughly doubles the milliseconds, leaving cost flat at best and often worse, while latency clearly degrades. For an **I/O-bound** function that spends its time waiting on a database or a third-party API, extra CPU buys nothing, so the smaller setting is close to a pure saving. The honest answer is to measure: run the function across a range of memory settings and plot cost and duration together. AWS Lambda Power Tuning, an open-source Step Functions state machine, automates exactly that sweep.
code
python · 14 linesPRICE_PER_GB_SECOND = 0.0000166667 # us-east-1 x86 on-demand; verify current pricing
PRICE_PER_REQUEST = 0.0000002
def monthly_cost(memory_mb, duration_ms, invocations):
gb_seconds = (memory_mb / 1024) * (duration_ms / 1000) * invocations
return gb_seconds * PRICE_PER_GB_SECOND + invocations * PRICE_PER_REQUEST
CALLS = 10_000_000
# CPU-bound: halving memory doubles duration -> same bill, twice the latency
print(round(monthly_cost(1024, 400, CALLS), 2))
print(round(monthly_cost(512, 800, CALLS), 2))
# I/O-bound: duration barely moves -> the smaller setting really is cheaper
print(round(monthly_cost(1024, 800, CALLS), 2))
print(round(monthly_cost(512, 780, CALLS), 2))go deeper
Know that Lambda charges by memory multiplied by duration plus a per-request fee, and that you are billed for the memory you configured, not the memory the code used.
Explain that memory also allocates CPU proportionally, so the cost curve is memory times duration — flat for CPU-bound work and rising for I/O-bound work.
Show how you would measure the curve for a real function at a high percentile, keep headroom above the memory ceiling, and weigh the latency effect on a synchronous path against the saving.
Be ready to argue against a fleet-wide memory policy: the right setting is per-workload, and the platform's job is to make measuring it cheap rather than to mandate a number.
## What you are actually billed for AWS Lambda's compute charge is **GB-seconds**: the memory you configured (not the memory you used) multiplied by the duration the function ran, billed at 1 ms granularity, plus a per-request charge. Two consequences follow immediately: - You pay for **configured** memory. A function set to 1024 MB that touches 90 MB pays for 1024 MB. - You pay for **wall-clock** duration, including time spent blocked on the network. ``` compute cost ≈ (memory_MB / 1024) × (duration_ms / 1000) × price_per_GB_second total cost ≈ compute cost + invocations × price_per_request ``` ## Memory is the size dial, not just a memory dial The reason this question is interesting is that memory is not an isolated setting. Lambda allocates **CPU proportionally to memory** — AWS documents roughly one full vCPU at 1,769 MB, with fractional CPU below that and more than one vCPU above it. Network throughput scales with the same dial. So the memory slider is really "how big a machine is this". That gives the cost curve two competing halves: - Price per millisecond is **linear** in memory. - Duration is **inverse** in CPU for the CPU-bound part of the work, and **flat** for the I/O-bound part. Multiply them and you get a curve, not a line. ## The two regimes **CPU-bound work.** If the function spends its time computing, doubling memory roughly halves duration, so GB-seconds stay about the same while latency halves — the same bill for a faster function, which is usually a straight win. Running the argument backwards: halving memory from 1024 to 512 MB roughly doubles duration and leaves cost flat *at best*. In practice it is often worse than flat, because below 1,769 MB the function is on a fraction of a vCPU and runtimes with parallel work (a JVM's GC threads, a Node.js runtime doing crypto in the thread pool, a compiled runtime using multiple cores) lose more than proportionally. The proposal saves nothing and makes every request slower. **I/O-bound work.** If the function issues a query and waits 400 ms for the answer, extra CPU cannot shorten the wait. Duration is nearly constant across memory settings, so cost is close to linear in memory and the lower setting is a genuine saving. This is where the colleague's instinct is right. **The mixed reality.** Most real functions are both: some deserialisation and business logic around a couple of network calls. That is exactly why the curve has to be measured rather than reasoned about. ## How to measure it Run the function with representative input at a spread of memory settings — 256, 512, 1024, 1536, 2048 MB — and record duration at a high percentile, not the mean, because the tail is what your users feel. Compute GB-seconds for each and plot cost against latency. You will usually see one of three shapes: a flat cost line (CPU-bound; take the fastest setting for free), a rising line (I/O-bound; take the smallest setting that is stable), or a U (a genuine optimum in the middle). **AWS Lambda Power Tuning** is an open-source Step Functions state machine published by AWS Labs that runs this sweep for you and returns the cost/latency curve. It is the standard answer to "how did you pick 1536?". ## Guardrails - **Do not tune below stability.** A setting that occasionally hits the memory ceiling and kills the invocation is not cheap; retries and failed work cost far more than the memory saved. Watch `Max Memory Used` in the invocation report and keep headroom. - **Latency has a price too.** If the function sits in a synchronous API path, doubling its duration to save nothing is a bad trade even at flat cost. - **Cost lives elsewhere as well.** Invocation count, provisioned concurrency left enabled on a function that no longer needs it, and time spent in initialisation all move the bill; memory is one dial among several. - **Tune per function, not per account.** A blanket "everything is 512 MB" policy is exactly the reasoning error this question probes.
- How would you know whether a given function is CPU-bound or I/O-bound before running a sweep?Compare the function's duration against the latency of its downstream calls. If traced spans for the database or HTTP calls account for nearly all the wall clock, it is I/O-bound. If duration far exceeds the sum of its outbound calls, the time is being spent computing, and CPU — which means memory — will move it.
- Besides the memory setting, what else moves a Lambda function's bill?Invocation count, wall-clock duration including initialisation, and provisioned concurrency, which bills for the hours it is enabled whether or not requests arrive. Since billing is per millisecond, trimming heavy imports and client construction out of the request path pays off directly on high-volume functions.
- Why is Max Memory Used a poor guide to the right setting on its own?Because it only tells you about the memory ceiling, not the CPU. A function reporting 90 MB used at 1024 MB may still need that setting for its share of a vCPU. Dropping it to 128 MB to match observed memory can multiply duration and cost while adding latency.
Memory in Lambda works like hiring a bigger crew for a fixed job: paying twice as much per hour for a crew that finishes in half the time costs the same, unless the job is mostly waiting for a delivery — then the bigger crew just stands around at double the rate.
saying these in an interview costs you the question
- Assumes less memory always means a smaller bill
- Thinks Lambda bills the memory actually used, not configured
- Does not know CPU is allocated in proportion to memory
- Sets one memory value across every function as policy
- Tunes to the edge of the limit and ignores out-of-memory failures