skip to content

An AWS Lambda function downloads a large object from S3 into /tmp, transforms it, and uploads the result. It passes every test but in production it intermittently fails with "No space left on device", and occasionally hits the function timeout. What is going on, and how would you fix it?

level: seniorimportance: should knowfreq 40%

answer

  1. scratch space outlives the invocation
  2. fresh in test, warm in production
  3. 512 megabytes unless you raise it
  4. fifteen minutes is a hard ceiling
  5. clean up on the exception path

basics

~20 s

Lambda's /tmp defaults to 512 MB and belongs to the execution environment, which is reused across invocations — so leftover files accumulate until a later invocation runs out of space. Fix by deleting files after use, streaming instead of buffering, and raising ephemeral storage if genuinely needed.

solid answer

~60 s

Two separate limits are being hit. `/tmp` is the function's **ephemeral storage**: 512 MB by default, configurable up to 10,240 MB, and it lives in the execution environment rather than being recreated per request. Because a warm environment serves many invocations, files the handler leaves behind pile up — which is exactly why tests pass (every run is a fresh environment) and production fails intermittently on whichever invocation crosses the line. The second limit is the **15-minute maximum timeout**, a hard ceiling you cannot raise, so a job whose input keeps growing eventually cannot finish at all. The fixes, in order: delete temp files in a `finally` block and use unique filenames; better still, stream the object from S3 through the transform and out via a multipart upload so nothing large touches disk; raise ephemeral storage only if the workload really needs local random access, remembering it is billed above the free 512 MB. If the transform itself cannot complete within 15 minutes, Lambda is the wrong home for it and the work belongs in a container task or split into chunks.

code

python · 17 lines
python
import os, tempfile, boto3

s3 = boto3.client("s3")

def lambda_handler(event, context):
    # Unique per invocation: two invocations on one warm environment cannot collide
    fd, path = tempfile.mkstemp(dir="/tmp", suffix=".bin")
    os.close(fd)
    try:
        s3.download_file(event["bucket"], event["key"], path)
        out = transform(path)
        s3.upload_file(out, event["bucket"], event["key"] + ".out")
        return {"ok": True}
    finally:
        for p in (path, path + ".out"):
            if os.path.exists(p):
                os.remove(p)   # runs on the exception path too

go deeper

for a junior

Know that /tmp is the only writable directory in a Lambda function, that it starts at 512 MB, and that anything you write there should be deleted when you are finished with it.

for a middle

Explain that /tmp belongs to the execution environment and survives between invocations on a warm one, give the configurable range up to 10,240 MB and the 900-second timeout ceiling, and show cleanup on the exception path.

for a senior

Diagnose from the symptom: intermittent, traffic-correlated, invisible in tests. Prefer streaming over buffering, treat raised ephemeral storage as a billed decision, set a timeout near p99 rather than the maximum, and say plainly when the workload has outgrown Lambda.

for a principal

Own the boundary: define what belongs on Lambda versus duration-uncapped compute, set the guidance on streaming-by-default for object processing, and make sure growth in input size is monitored so a function does not silently drift toward a ceiling nobody is watching.

## The two limits in play This scenario is a favourite because it needs two facts at once, and the *intermittent* symptom is the clue that separates a candidate who has operated Lambda from one who has only read the quotas page. **Ephemeral storage.** Every Lambda execution environment gets a writable `/tmp` directory. It defaults to **512 MB** and can be configured up to **10,240 MB** through the function's ephemeral-storage setting. It is the only writable location in the environment; the deployment package under the task root and any layer content under `/opt` are read-only. **Timeout.** A function's timeout defaults to 3 seconds and can be raised to a maximum of **900 seconds — 15 minutes**. That ceiling is not negotiable and no support request lifts it. ## Why the failure is intermittent The crucial property is that `/tmp` is scoped to the **execution environment**, not to the invocation. When Lambda keeps an environment warm and routes a subsequent request to it, that request sees whatever the previous one left on disk. So: - A test run, or a low-traffic period, creates a fresh environment per request. 400 MB of scratch fits in 512 MB every time. Green. - In production one warm environment handles request after request. Two 400 MB files do not fit. The third request into that environment fails with `ENOSPC` while its neighbours, on other environments, succeed. That is the shape of the bug: failures that correlate with traffic and warmth rather than with input, and that cluster on particular environments. Each concurrent environment has its own independent `/tmp`, so this is never contention between concurrent requests — it is accumulation within one. ```python import os, tempfile def lambda_handler(event, context): fd, path = tempfile.mkstemp(dir="/tmp") # unique name per invocation os.close(fd) try: download(event["key"], path) return transform(path) finally: os.remove(path) # always reclaim the space ``` A fixed filename is the other half of the trap: it hides the leak on rewrite but leaves a stale file behind if a later invocation reads it before writing, which produces the far nastier failure — one request silently processing another request's data. ## Fixing it properly **1. Clean up unconditionally.** Delete in a `finally` block (or the language equivalent) so an exception path cannot leak. Use unique names so two invocations on the same environment can never collide. A defensive sweep of `/tmp` during initialization is a reasonable belt-and-braces measure. **2. Don't touch disk at all if you can avoid it.** The strongest fix is architectural: read the object as a stream, transform it in bounded chunks, and write it out through a multipart upload. Peak disk use goes to zero and peak memory becomes a buffer you chose rather than the size of the input. This also removes the coupling between input size and the ephemeral-storage setting entirely — which matters, because inputs grow. **3. Raise ephemeral storage when the work genuinely needs local files.** Some workloads do: a tool that memory-maps a file, a media encoder that seeks randomly, an extraction step that must materialise an archive. Configure it up to 10,240 MB — and note that storage above the included 512 MB is billed for the duration of the invocation, so it is a real cost, not a free safety margin. **4. Give the timeout a real value.** Set it a little above the observed p99 rather than at the maximum: a timeout that is far too generous turns a hung downstream call into 15 minutes of billed duration per request. Read the remaining budget from the context object and decide whether to start another chunk of work rather than being killed mid-write — a timeout kill runs no cleanup code, so it is also how orphaned files appear in the first place. ## When 15 minutes is the actual wall If the transform cannot finish inside 900 seconds no matter how much CPU you buy with the memory setting, the function has outgrown Lambda's execution model, and the honest answer in an interview is to say so rather than to keep tuning: - **Split the work.** Chunk the input by byte range or record range, process each chunk in its own invocation, and combine the results. This is the natural fit when the transform is parallelisable, and it converts one long job into many short ones that also scale out. - **Move it to a container task.** A long-running batch job with no time limit belongs on a container or instance-based compute service, where duration is not capped. A useful framing for the interviewer: Lambda's timeout is not an inconvenience to work around but a statement about what kind of work belongs there. Anything approaching the ceiling is a workload whose growth curve will cross it. ## What good answers include Naming both limits with their real values; explaining environment reuse as the reason the failure is intermittent and test-invisible; proposing cleanup *and* streaming rather than only raising the setting; noting that larger ephemeral storage is billed; and knowing when to stop and move the workload.

  • Is /tmp ever shared between two invocations running at the same time?
    No. Each execution environment has its own /tmp, and one environment serves one invocation at a time, so concurrent requests never see each other's files. The sharing that matters is sequential: a warm environment hands the same /tmp to the next request it serves, which is what makes leftover files a correctness and capacity problem rather than a race.
  • The team wants the timeout raised to 30 minutes. What do you say?
    It isn't possible — 900 seconds is Lambda's hard maximum and no setting or support request changes it. The options are to make the work fit, by streaming and by buying CPU through the memory setting; to split the input into chunks processed by separate invocations; or to move the job to compute without a duration cap. Treat the ceiling as a signal about workload fit.
  • Why does raising ephemeral storage not come for free?
    Storage above the included 512 MB is billed for the duration of each invocation, so a function configured at 10 GB pays for that allocation on every request whether it writes a byte or not. It also masks the real defect if the underlying cause is files never being deleted — the function then fails later, at higher cost, instead of immediately.
  • How would you detect this class of problem before customers do?
    Have the function report its own /tmp usage as a metric or log field at the end of each invocation and alarm on it approaching the configured size. Because the symptom follows environment warmth, also look for errors clustering on a subset of request IDs over time rather than correlating with input size — and load-test with sustained traffic, not single cold invocations.

saying these in an interview costs you the question

  • Thinks /tmp is wiped before every invocation
  • Believes the 15-minute timeout can be raised on request
  • Only raises ephemeral storage without fixing the leak
  • Assumes concurrent invocations share one /tmp
  • Cleans up after the transform but not on the error path

context