skip to content

Packaging, Layers and Limits

How I package a function decides its size ceiling, its cold-start cost and how I share dependencies. I learn ZIP versus container-image deployment, what a layer actually mounts, and the hard limits interviewers love to ask for by number.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

6

In AWS Lambda, raising a function's memory setting changes more than the amount of RAM available. What else does it change, and why can allocating more memory sometimes make a function cheaper rather than more expensive?

level: middleimportance: must knowfreq 60%

answer

  1. one dial for everything
  2. CPU scales with the memory setting
  3. about one vCPU near 1,769 MB
  4. billed in gigabyte-seconds
  5. waiting on I/O gains nothing

basics

~20 s

Memory is Lambda's single performance dial: CPU, and network and disk throughput, are allocated in proportion to it. Because you are billed for gigabyte-seconds, doubling memory on CPU-bound work that then runs in less than half the time lowers the total bill.

solid answer

~60 s

In Lambda you do not size CPU directly — the memory setting allocates it. The configurable range is 128 MB to 10,240 MB, and CPU scales linearly across it: roughly one full vCPU at about 1,769 MB, and up to six vCPUs at the top. Network and disk throughput scale with it too. Billing is per **GB-second** plus a per-request charge, so cost is memory multiplied by duration. That means raising memory doubles the per-millisecond price but, for CPU-bound work, can more than halve the duration — and the product falls. The classic mistake is running everything at 128 MB "to save money" when it is actually both slower and dearer. The caveats are that a single-threaded workload sees no benefit from the extra vCPUs past roughly 1,769 MB, and a function that spends its time waiting on a downstream API gets no speedup at all, only a bigger bill. I measure rather than guess — sweeping a few memory settings against a representative payload and plotting cost against duration, which is exactly what the open-source Lambda Power Tuning state machine automates.

code

bash · 11 lines
bash
# Ask CloudWatch Logs what the function actually used and how long it ran
aws logs filter-log-events \
  --log-group-name /aws/lambda/my-func \
  --filter-pattern 'REPORT' \
  --query 'events[].message' --output text
# REPORT RequestId: ...  Duration: 1502.11 ms  Billed Duration: 1503 ms
#        Memory Size: 1769 MB  Max Memory Used: 412 MB

# Sweep a candidate setting, then re-measure
aws lambda update-function-configuration \
  --function-name my-func --memory-size 1769

go deeper

for a junior

Know that memory is the only performance setting on a Lambda function and that CPU is allocated in proportion to it, so a starved function is slow because it is short of CPU, not just RAM.

for a middle

Do the arithmetic out loud: cost is configured memory times billed duration, so if doubling memory more than halves duration the bill falls. Name the range, the roughly one-vCPU point, and the I/O-bound exception.

for a senior

Describe how you actually tune — sweeping settings against representative payloads, reading billed duration, weighing latency against cost for a user-facing path, and re-measuring after dependency or payload changes rather than trusting a stale number.

for a principal

Own this as a portfolio-level cost lever: which functions justify tuning by invocation volume, whether arm64 is the default, how tuning is kept current as code changes, and how you stop teams from copying one memory setting across every function by habit.

## One dial, several resources Lambda exposes exactly one performance control: memory, settable from **128 MB to 10,240 MB** in 1 MB increments. Everything else is derived from it. CPU is allocated in proportion — a function at 1,769 MB gets approximately the equivalent of one vCPU, and the ceiling of 10,240 MB corresponds to roughly six vCPUs. Network bandwidth and the throughput of the environment's storage scale along the same curve. This design is why "how much CPU does my Lambda get?" has no direct answer, and why the memory setting is a *performance* decision, not just a capacity one. ## How the bill is computed Two components, ignoring free tier: 1. A charge per request. 2. A charge per **GB-second**: configured memory in gigabytes multiplied by billed duration, metered in millisecond increments. ```text cost_per_invoke ≈ request_fee + (memory_GB × duration_seconds × price_per_GB_second) ``` Note it is *configured* memory, not memory actually used. Provisioning 3 GB for a function whose peak resident set is 200 MB costs the full 3 GB for every millisecond it runs. ## Why more memory can cost less Because the two factors move in opposite directions. Take a CPU-bound function — image resizing, JSON parsing of a large payload, compression, a cryptographic operation: | Memory | Duration | GB-seconds | |---|---|---| | 512 MB | 6,000 ms | 3.00 | | 1,024 MB | 2,900 ms | 2.97 | | 1,769 MB | 1,500 ms | 2.65 | | 3,008 MB | 1,350 ms | 4.06 | From 512 MB to 1,769 MB the duration falls faster than the memory rises, so the bill drops *and* latency improves by 4x. Past the one-vCPU point the single-threaded work cannot use the extra cores, duration flattens, and cost climbs steeply. The optimum is a genuine minimum in the middle, and it moves with the workload and the payload. The practical rule: **at 128 MB you are frequently paying more for a slower function.** Under-provisioning is the more common and more expensive mistake. ## When more memory does nothing - **I/O-bound functions.** If 90% of the duration is waiting on an HTTP call to a downstream service or a database round trip, more CPU cannot shorten the wait. You pay strictly more per millisecond for the same number of milliseconds. - **Single-threaded work past ~1,769 MB.** A second vCPU is only useful to a runtime that can use it. A Node.js function doing synchronous CPU work on one thread, or a Python function without multiprocessing, will plateau. A JVM function may still benefit somewhat, because garbage collection and JIT compilation run on their own threads. - **Functions already at their floor.** If duration is dominated by a fixed downstream latency, tune the downstream, not the memory. Conversely there are workloads that genuinely need the top of the range: parallel processing across multiple worker processes, large in-memory datasets, or anything that can actually saturate several vCPUs. ## Measuring instead of guessing Lambda reports `Memory Size` and `Max Memory Used` in each invocation's `REPORT` line in CloudWatch Logs, which tells you whether you are anywhere near the RAM ceiling — but it does *not* tell you the cost-optimal setting, because that depends on how duration responds to CPU. The method that works is empirical: run the same representative payload at several memory settings, record billed duration for each, compute GB-seconds, and pick the minimum — or the best latency within an acceptable cost, if the function is user-facing and latency matters more than cents. The open-source **AWS Lambda Power Tuning** project packages this as a Step Functions state machine that sweeps the settings and plots the cost/latency curve for you. Two further levers pair with this. First, **architecture**: functions can run on `arm64` (Graviton) as well as `x86_64`, at a lower price per GB-second, and many workloads are also somewhat faster — worth testing if your dependencies have ARM builds. Second, **re-tuning after change**: the optimum is a property of the code and payload, so a dependency upgrade or a change in average payload size invalidates last quarter's answer. Functions invoked millions of times a month deserve periodic re-measurement; a function invoked twice a day does not. ## The judgment an interviewer is listening for That you know memory and CPU are coupled, that you can explain the GB-second arithmetic that makes "more memory, less money" possible, that you can name the cases where it does *not* hold, and that your answer to "what setting should we use?" is a measurement rather than a number.

  • A function spends most of its time waiting on a third-party HTTP API. Will raising its memory reduce cost?
    No — it will raise it. The duration is dominated by network wait, which more CPU cannot shorten, so you pay a higher per-millisecond rate for the same elapsed time. The levers that help there are concurrency of the outbound calls, connection reuse across invocations, a shorter client timeout, or fixing the downstream. Tune memory for CPU-bound work only.
  • Your function is configured at 3 GB and REPORT shows Max Memory Used around 300 MB. Should you cut it to 512 MB?
    Not without measuring. The 3 GB may have been chosen for CPU, not RAM, and cutting it will slow the function roughly proportionally — possibly raising cost as well as latency. Sweep several settings against a representative payload, compare GB-seconds and duration, then decide. Peak memory used only proves you are not near the RAM ceiling.
  • What does switching a function to the arm64 architecture change?
    Graviton-based arm64 functions are priced lower per GB-second than x86_64 and are often comparable or faster in duration, so the combination usually reduces cost. The requirement is that every dependency has an arm64 build — native extensions and binaries in layers or images must be rebuilt for the architecture — and you should re-measure the memory sweep after switching.
  • Why is configured memory rather than used memory the basis for billing?
    Because Lambda reserves the full amount for the execution environment: the memory is allocated and the proportional CPU share is provisioned whether or not your code touches it. That is why over-provisioning is never free, and why 'set it high just in case' is a real cost decision rather than a harmless safety margin.

saying these in an interview costs you the question

  • Believes 128 MB is always the cheapest setting
  • Thinks CPU is configured separately from memory
  • Sizes memory only from Max Memory Used in the REPORT line
  • Assumes more memory always shortens duration
  • Says billing is based on memory actually consumed

context

open as a page

AWS Lambda lets you deploy a function either as a .zip archive or as a container image. What are the practical differences between the two packaging formats, and how would you choose one for a given function?

level: middleimportance: must knowfreq 62%

basics

~20 s

Both run on the same Lambda execution model; only packaging differs. A .zip is small, quick to deploy and can use layers, but is capped at 250 MB unzipped. A container image allows up to 10 GB from a private ECR repository and carries OS-level dependencies, at the cost of an image build pipeline.

open as a page

An AWS Lambda function's configuration includes a runtime (for example python3.12) and a handler string (for example app.lambda_handler). What does each of those identify in your deployment package, and what does Lambda report if the handler string does not match what you shipped?

level: juniorimportance: should knowfreq 58%

basics

~20 s

The runtime selects the language environment Lambda starts; the handler names the entry point inside the package — the file or module, then the function in it. A mismatch fails during initialization, before any of your code runs.

open as a page

In AWS Lambda, what does attaching a layer actually do to the function's execution environment, and what limits and gotchas come with layers?

level: middleimportance: should knowfreq 55%

basics

~20 s

A layer is a .zip of shared content that Lambda extracts into /opt before the handler runs, where each runtime already looks for libraries. A function may attach up to five layer versions, and their contents still count against the 250 MB unzipped limit.

open as a page

An AWS Lambda function downloads a large object from S3 into /tmp, transforms it, and uploads the result. It passes every test but in production it intermittently fails with "No space left on device", and occasionally hits the function timeout. What is going on, and how would you fix it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Lambda's /tmp defaults to 512 MB and belongs to the execution environment, which is reused across invocations — so leftover files accumulate until a later invocation runs out of space. Fix by deleting files after use, streaming instead of buffering, and raising ephemeral storage if genuinely needed.

open as a page

What is an AWS Lambda extension, how does it get into a function and run relative to your handler, and what are extensions typically used for?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

A Lambda extension is companion code that runs inside the same execution environment as your handler, registered through the Lambda Extensions API. External extensions run as separate processes started before the runtime; they ship as a layer with executables under /opt/extensions, or baked into a container image.

open as a page