AWS Lambda gives each function instance a writable /tmp directory (up to a configurable size, e.g. up to 10 GB). Why is writing files to this directory not a substitute for durable storage, even though multiple invocations sometimes reuse the same warm instance and can still see files written by an earlier invocation?
answer
- disk scoped to one instance
- warm reuse only, not guaranteed
- concurrent invocations = separate /tmp each
- fine for single-invocation scratch work
- not a substitute for S3/DynamoDB
basics
~10 s/tmp is just a scratch disk attached to one function instance; it disappears when that instance is recycled, and other instances never see it, so it can't reliably hold data you need to keep.
solid answer
~40 sLambda's /tmp is local, ephemeral disk on the single execution environment handling the invocation — it's not networked or replicated. Content written by one invocation is sometimes still there on the next invocation only if the platform happens to reuse the same warm instance, and that reuse is an unguaranteed optimization, not a contract — the instance can be frozen or destroyed anytime, and other concurrent invocations run in entirely separate instances with their own empty /tmp. Anything durable or shared across requests belongs in S3, DynamoDB, EFS, or another externally-managed store; /tmp is only appropriate as scratch space for a single invocation's own transient work, like unzipping a file for local processing, or as a best-effort warm-start cache whose absence is handled gracefully.
go deeper
Should know that /tmp is temporary and shouldn't be counted on to still be there for a different request.
Should explain warm-instance reuse as the mechanism that makes /tmp sometimes appear to persist, and know it's not guaranteed.
Should articulate the correct usage pattern (single-invocation scratch space, or best-effort cache with a handled miss path) and identify the disk-space-leak failure mode from unclean reuse.
Should reason about this as one instance of the general 'don't trust instance-local resources' principle across a serverless architecture, and design retry/checkpoint/caching strategies that are correct regardless of instance placement, informing platform-wide engineering guidelines.
## What /tmp actually is AWS Lambda attaches a local, writable filesystem at `/tmp` to every execution environment, sized by configuration between **512 MB and 10,240 MB**. Mechanically, this is ordinary disk storage local to the single Firecracker micro-VM running your code for that environment — it is not a network filesystem, not replicated, and not visible to any other environment. - When your function invocation writes a file to `/tmp`, that write lands on the local disk of whichever specific environment is currently executing your code. - If the platform later routes another invocation to that exact same environment — which happens when the environment is 'warm,' i.e., still alive and idle since its last invocation — that second invocation's code will indeed find the file still there, because nothing has torn the environment down. This is precisely why `/tmp` can feel deceptively like durable, shared storage during light testing. ## Why it exists at all The reason this exists at all is that some workloads genuinely need a real filesystem for the duration of a single invocation: - decompressing a zip archive before parsing its contents; - using a native library or CLI tool that only knows how to read/write files; - memory-mapping a large lookup table for the current computation; - building up an intermediate file before uploading the final result to S3. None of that requires durability beyond the current invocation; it's **scratch space**, conceptually identical to a stack-allocated buffer, just backed by disk because the data is too large or the tooling too file-oriented for pure in-memory handling. ## The trade-off The trade-off is exactly the same shape as the in-memory-variable case, but on disk instead of RAM: warm-instance reuse of `/tmp` is an opportunistic, unguaranteed optimization the platform gives you for free, not a contract you can rely on for correctness. - **Concurrency** defeats it immediately — if your function is invoked five times simultaneously, the platform creates or reuses from a pool five separate environments, each with its own private `/tmp` that starts out however it happened to be left by whatever invocation last ran there, which for a brand-new environment means completely empty. - There is no way to target 'the instance that has my file' from a given invocation, and you don't even know which environment you're running on. - Idle environments are also torn down after inactivity, wiping `/tmp` along with them, and the platform can recycle environments proactively for internal load-balancing reasons at any time. ## The failure modes The failure modes this produces are subtle because they're intermittent and load-dependent, exactly like the in-memory case. 1. A batch job that writes a checkpoint file to `/tmp` so 'the next invocation can resume from where it left off' will work in a low-traffic dev environment where the same warm instance keeps getting reused, then silently start from scratch in production whenever concurrency or idle timeouts force a fresh environment, with no error raised — the checkpoint file just isn't there. 2. A caching layer that downloads a large reference file to `/tmp` on first use per instance, intending to reuse it later, is a legitimate and common pattern, but only if the code treats a cache miss as the normal path, re-downloading on demand, never as an error condition, since a miss is guaranteed to happen eventually for every environment and can happen on every single invocation under high concurrency. 3. Also worth noting: heavy `/tmp` usage without cleanup between invocations on a long-lived warm instance can accumulate leftover files across many invocations, since nothing automatically clears `/tmp` between reuses, until the environment runs out of disk space — a slow disk-space leak that manifests only on unusually long-lived warm instances. ## The correct pattern The correct pattern is to treat `/tmp` purely as ephemeral working space for the current invocation, or as a best-effort warm-start performance cache with a properly handled miss path, and to put anything that must survive between invocations or be visible to other concurrent instances into **S3** for large blobs, or **DynamoDB/Redis** for smaller structured state. AWS's own Lambda documentation is explicit that `/tmp` is 'ephemeral storage' that may not persist between invocations.
- Is it ever correct to rely on /tmp persisting between invocations in production?Only as a best-effort performance optimization where a miss is handled gracefully, such as re-downloading a reference dataset if it isn't found — never as the sole copy of data you need, since the platform gives no durability guarantee and a miss must be a normal, cheap code path, not an error.
- How does /tmp size interact with the increased memory Lambda allocates for higher CPU?/tmp size is configured independently, up to 10 GB, separate from the function's memory setting, though both consume the environment's overall resource footprint; teams sometimes conflate the two and are surprised that raising memory doesn't automatically grow /tmp.
- What's a concrete symptom of a /tmp disk-space leak from unclean warm-instance reuse?Invocations that write to /tmp without deleting temp files start failing with 'no space left on device' errors, but only on some invocations and unpredictably, because it only happens once a particular warm instance has accumulated enough leftover files across many prior invocations, which correlates with how long that specific instance has stayed warm rather than with any single request's behavior.
/tmp is like a whiteboard in one specific meeting room — if you get assigned that same room again later you'll still see your notes, but the building has many identical rooms and no way to guarantee you get the same one twice, and any room can be cleaned out at any time.
saying these in an interview costs you the question
- Treats /tmp persistence across invocations as guaranteed rather than opportunistic
- Uses /tmp to pass state between concurrent invocations of the same function
- Doesn't handle the 'file not found' case when reading from /tmp, assuming it will always be there from a prior write
- Confuses /tmp's local disk with a shared or networked filesystem
- Never cleans up /tmp writes, ignoring the disk-space-leak risk on long-lived warm instances