skip to content

A long-running batch job writes intermediate files under /tmp and occasionally fails days later with "No such file or directory" for a file it created. On a Linux host, how do /tmp and /var/tmp differ, and where should that scratch data live?

level: seniorimportance: should knowfreq 50%

answer

  1. two directories, two published guarantees
  2. one survives a reboot, one does not
  3. something ages files out on a timer
  4. free disk but the write still fails
  5. let TMPDIR decide, not the program

basics

~20 s

/tmp is volatile: it may be cleared at boot, is aged out by a periodic cleanup, and on many distributions is a memory-backed filesystem with its own size limit. /var/tmp is temporary storage that must survive reboots. Multi-day scratch data belongs in /var/tmp or the service's own directory.

solid answer

~50 s

The FHS gives the two directories different guarantees. `/tmp` is for temporary files that programs must not expect to survive a reboot, and on a systemd-based host a periodic cleaner ages out files that have not been touched recently — systemd's shipped defaults are roughly 10 days for `/tmp` and 30 days for `/var/tmp`. `/var/tmp` is explicitly for temporary files that must be preserved across reboots. So a job whose scratch files must live for days is relying on a guarantee `/tmp` never made, and the cleaner is deleting them out from under it. There is a second trap: on several current distributions `/tmp` is a memory-backed filesystem sized independently of the disk, so a large write can fail with ENOSPC while `df` shows the root filesystem nearly empty. Put multi-day scratch in `/var/tmp`, or better, in a directory the job owns under `/var/lib`, and honour `TMPDIR`.

code

bash · 4 lines
bash
# create scratch safely and let the operator redirect it
workdir=$(mktemp -d "${TMPDIR:-/var/tmp}/myjob.XXXXXXXX")
trap 'rm -rf "$workdir"' EXIT
echo "staging in $workdir"

go deeper

for a junior

Know that /tmp is scratch that may vanish and /var/tmp is temporary storage meant to survive a reboot, and that neither is a place to keep anything you care about.

for a middle

Explain the two contracts and the mechanisms that enforce them — boot-time clearing plus a periodic cleaner with a configured maximum age — and describe honouring TMPDIR and creating files with a unique name.

for a senior

Diagnose the symptom end to end: prove the file was deleted rather than mislaid, identify which cleanup applied, recognise a memory-backed /tmp behind a nonsensical ENOSPC, and relocate the data to storage whose lifetime the service owns.

for a principal

Set the fleet policy: which filesystem carries scratch and how it is sized, what retention the cleanup configuration should express, and how job design avoids depending on any guarantee the platform did not deliberately make.

## Two temporary directories, two contracts The FHS defines both directories in terms of what a program may *assume*: - **`/tmp`** — programs must not assume that files here are preserved between invocations of the program. Many systems clear it at boot. - **`/var/tmp`** — for temporary files that *should* be preserved between reboots. It is the place for scratch that is expensive to recreate. Both are world-writable so that any user can create files there. Neither is a durable store; the difference is purely how long you may count on the contents. ## What actually deletes the files There are three distinct mechanisms, and diagnosing a disappearance means knowing which one hit you. **1. Boot-time clearing.** Historically distributions emptied `/tmp` during startup. If `/tmp` is a memory-backed filesystem, this is automatic — the contents never existed on disk in the first place. **2. Periodic ageing.** On systemd-based distributions, `systemd-tmpfiles` runs from a timer and removes files under the temporary directories whose relevant timestamps are older than a configured age. systemd's shipped configuration ages `/tmp` at 10 days and `/var/tmp` at 30 days; distributions may change these, and administrators can override them with their own configuration under `/etc/tmpfiles.d/`. This is the mechanism that catches long-running jobs: the job is alive, the file is untouched for eleven days, and it is removed while still referenced. An important subtlety: deleting a file that a process still holds open does not free the process's access to it — the data stays reachable through the open descriptor until it is closed. So the failure often does not appear at the moment of deletion; it appears the next time the job tries to *open the path again by name*, which is where the "No such file or directory" for a file the job knows it created comes from. **3. Manual or ad-hoc cleanup.** Somebody's disk-pressure script. Less common, more surprising. ## The memory-backed /tmp trap Several current distributions mount `/tmp` as a memory-backed filesystem — Fedora has for years, and Debian made it the default in Debian 13. That has consequences a candidate should be able to state: - Its size limit is independent of the disk. A write can fail with `ENOSPC` while the root filesystem has hundreds of gigabytes free, which reads as a nonsensical error until you check where `/tmp` actually is. - Data written there consumes memory (backed by swap when present), so a job that stages a large file in `/tmp` is competing with the applications on the box for RAM. - Contents vanish at reboot by construction, not by policy. ```sh # is /tmp its own filesystem, and how big is it really? findmnt /tmp df -h /tmp /var/tmp ``` A further wrinkle: a sandboxed service may be given its own private view of `/tmp`, so files it creates are not the files you see at `/tmp` from an ordinary shell. If a service insists its temporary file exists and you cannot find it, that isolation is the usual explanation. ## Where the data should go In rough order of preference for a job like the one described: 1. **A directory the service owns**, e.g. `/var/lib/<service>/work` or a dedicated data volume. This gives you a known filesystem, a known size, ownership and permissions you control, and an explicit retention policy that is yours rather than the distribution's. 2. **`/var/tmp`**, when the data really is temporary but must outlive a reboot. Still subject to ageing, so it is only correct if the job's lifetime is comfortably inside the retention window. 3. **`/tmp`**, only for scratch that lives inside a single invocation and is small enough not to matter. Whatever you choose, honour the `TMPDIR` environment variable rather than hardcoding `/tmp`: it is the standard way an operator redirects a program's temporary files onto a filesystem with enough space, and libraries and utilities across the system already respect it. Create the file atomically with a unique name — `mktemp` at the shell, `mkstemp()` in C, or the language's equivalent — rather than deriving a predictable path. In a world-writable directory, a predictable name is a race an unprivileged local user can win, and a program that follows a name another user planted can be tricked into writing where it should not. ## Answering the symptom Walking the diagnosis: confirm the file's absence is a deletion rather than a path bug; check whether `/tmp` is its own filesystem and what cleanup configuration applies; correlate the disappearance with the cleanup timer's schedule and the configured age; then move the data to storage whose lifetime the job actually controls. The lesson to state out loud is that `/tmp` is not "a folder" — it is a directory with a published contract, and the job was depending on a guarantee that contract never gave.

  • The job's file is deleted while the job still has it open. Why does the failure surface later rather than immediately?
    Unlinking a path removes the name, not the data a process is still holding open — the running job keeps reading and writing through its existing descriptor as if nothing happened. The error appears only when the job next tries to open that path by name, or when a second stage of the pipeline goes looking for it. That gap makes the deletion look unrelated to the failure.
  • Why does a write to /tmp fail with ENOSPC on a host with hundreds of gigabytes free?
    Because `/tmp` is very often its own filesystem, memory-backed on current Fedora and Debian, sized independently of the root filesystem — commonly a fraction of RAM. `df` on `/` tells you nothing about it. Check `findmnt /tmp` and `df -h /tmp`, then either raise its size deliberately or move the job's staging elsewhere via `TMPDIR`.
  • Why should a program create temporary files with mktemp or mkstemp rather than a predictable name?
    Because the temporary directories are world-writable, so any local user can pre-create a predictable path or plant a symlink there. A program that opens that name may then write through it to a file it was never meant to touch, or read data an attacker controls. Atomic exclusive creation with an unpredictable name removes the race entirely.

saying these in an interview costs you the question

  • Treats /tmp as durable storage between runs
  • Says only a reboot ever empties /tmp
  • Explains an ENOSPC on /tmp by checking free space on /
  • Hardcodes /tmp instead of honouring TMPDIR
  • Builds temporary paths from the process ID

context