skip to content

Temp Files, Locking and Idempotency

Automation scripts have to be safe to rerun and safe to run twice at once: mktemp plus a cleanup trap, flock so cron cannot overlap two copies, and atomic rename instead of half-written output. Predictable temp filenames are a genuine security bug, not merely untidiness.

part ofBashoverview, primer and where to startread it →
on this pageshow

questions

5

A bash script does `workdir=$(mktemp -d)` and removes it with `rm -rf "$workdir"` on its last line. On which exits does that cleanup fail to run, and how do you make removal reliable?

level: middleimportance: must knowfreq 58%

answer

  1. the last line is only the happy path
  2. errors and interrupts skip it
  3. bind cleanup to shell exit
  4. register it the line after creation
  5. one handler, one directory

basics

~20 s

The last line only runs when the script reaches it. Any early exit, a strict-mode abort, or a signal skips it and leaks the directory. Register the removal once with a cleanup registered on EXIT, immediately after creating the directory.

solid answer

~50 s

Putting cleanup on the last line means it runs on exactly one path: success. Every error branch that calls `exit`, every command that aborts the script under `set -e`, every `Ctrl-C` or `SIGTERM` from a CI runner or service manager, and any `exec` that replaces the shell all bypass it — and long-lived agents accumulate gigabytes of orphaned scratch directories that way. The fix is to attach the cleanup to shell exit instead of to a line number: right after `workdir=$(mktemp -d)`, register `trap cleanup EXIT`, where `cleanup` does `rm -rf -- "${workdir:?}"`. In bash the EXIT handler also runs when the shell is terminated by `SIGINT`, `SIGTERM` or `SIGHUP`, so one registration covers both the exit and the interrupt paths; adding `INT TERM` explicitly keeps it working in other shells. Nothing covers `SIGKILL` or power loss, so the host still needs a reaper.

code

bash · 15 lines
bash
#!/usr/bin/env bash
set -euo pipefail

workdir=$(mktemp -d) || exit 1
cleanup() {
  if [[ -n ${KEEP_WORKDIR:-} ]]; then
    printf 'kept %s\n' "$workdir" >&2
    return 0
  fi
  rm -rf -- "${workdir:?}"
}
trap cleanup EXIT INT TERM

printf 'payload\n' > "$workdir/data.txt"
wc -c < "$workdir/data.txt"

go deeper

for a junior

Know that a cleanup written on the final line is skipped whenever the script exits early or is interrupted, and that the standard fix is to register the removal so it runs on shell exit instead.

for a middle

Explain which paths skip it — error branches, a set -e abort, Ctrl-C, a SIGTERM from CI — and write the correct idiom: create with mktemp -d, register a cleanup function on EXIT INT TERM the very next line, and guard the variable.

for a senior

Show the operational consequence: leaked scratch directories fill a build agent's disk and take unrelated jobs down. Talk about a keep-on-failure escape hatch for debugging and about the reaper you still need for kills your handler can never see.

for a principal

Own it as a platform property rather than a per-script habit: where scratch lives, who reaps it, what the disk alarm is, and how the shared script template makes the correct pattern the default so no individual author has to remember it.

## The last line is one exit path out of many A script has far more ways to stop than to finish. The cleanup on the final line runs only if control actually reaches it, and these all prevent that: - an error branch: `[[ -f $input ]] || { echo "missing" >&2; exit 1; }` - a strict-mode abort: under `set -e`, any unchecked failing command ends the shell where it stands - `set -u` firing on an unset variable, or a failing pipeline under `pipefail` - an interrupt: the operator presses Ctrl-C, or CI cancels the job and sends `SIGTERM` - `exec some-other-program`, which replaces the shell image so no later shell line ever runs - a `return` from a function that the caller then turns into an exit On a build agent this is not cosmetic. Each skipped cleanup leaves a directory under `/tmp` (or wherever `TMPDIR` points), and after a few hundred failed runs the filesystem fills and *every* job starts failing for an unrelated-looking reason. ## Bind the cleanup to the shell, not to a line The idiom is to register the removal once, immediately after creating the directory: ```bash workdir=$(mktemp -d) || exit 1 cleanup() { rm -rf -- "${workdir:?}"; } trap cleanup EXIT INT TERM ``` `EXIT` fires whenever the shell terminates: falling off the end, an explicit `exit N` from anywhere including inside a function, and an abort caused by `set -e`. The exit status the script reports is preserved — the handler does not silently turn a failure into a success unless it calls `exit` itself. In bash the `EXIT` handler also runs when the shell is terminated by a catchable signal such as `SIGINT`, `SIGTERM` or `SIGHUP`. Not every shell behaves that way — `dash` and other POSIX shells vary — so scripts that may run under `/bin/sh` list `INT TERM HUP` alongside `EXIT`. Nothing at all can run on `SIGKILL`, on a hard reset, or when the container's OOM killer removes the process, which is why a fleet still wants a periodic reaper for its scratch directory in addition to well-behaved scripts. ## Getting the registration right **Register immediately after creation.** Any code between `mktemp -d` and the `trap` line is a window in which the script can die and leak. Conversely, registering *before* the directory exists means the handler may run with the variable empty. **Guard the variable.** `rm -rf -- "$workdir"` with `workdir` empty expands to `rm -rf --` (harmless), but the common variants are not: `rm -rf "$workdir"/*` with an empty variable becomes `rm -rf /*`. Writing `"${workdir:?}"` makes the expansion abort loudly instead of proceeding on an empty value, and the `--` stops a path that begins with a dash from being read as options. **Quoting inside the trap matters.** `trap "rm -rf $workdir" EXIT` expands the variable *when the trap is registered*; `trap 'rm -rf "$workdir"' EXIT` defers expansion to when the handler runs. Calling a named function, as above, sidesteps the confusion entirely and lets the cleanup grow beyond one command. **There is only one EXIT handler.** A later `trap something EXIT` replaces the earlier one rather than adding to it, so a script (or a sourced library) that registers its own cleanup can silently discard yours. If several things need tearing down, put them all in one `cleanup` function. **The handler must not fail noisily.** It runs on the error path too, when the thing it wants to remove may already be gone or half-created. Use `rm -rf` (which tolerates a missing target) and avoid commands that could abort mid-cleanup, so the first tidy-up step is not prevented by the second one erroring. ## Prefer one directory to many files Tracking six individual temp files means six chances to forget one. `mktemp -d` gives a single object: everything scratch goes inside it, and one `rm -rf` of the directory is the whole cleanup. It also makes the failure mode debuggable — set an environment flag such as `KEEP_WORKDIR=1` that the cleanup function honours, so a developer investigating a failure can rerun the script and keep the evidence, while the default stays "leave nothing behind".

  • Which terminations still leave the directory behind despite an EXIT handler?
    `SIGKILL` (`kill -9`), a container being force-removed, the OOM killer, and any hard power loss — none of them let the process run code. A `SIGSTOP`ped script that is later killed is in the same position. That is why scratch space also needs an out-of-band reaper: a directory under `TMPDIR` that the host clears, or a periodic job that deletes stale entries by age.
  • Why should the cleanup be registered on the line after mktemp -d rather than earlier or later?
    Earlier, and the handler can run while the variable is still unset or empty, which turns a tidy-up into a dangerous `rm`. Later, and every command in between is a window where the script can die with the directory already created and nothing registered to remove it. Adjacent to the creation is the only placement with no gap in either direction.
  • Your script registers a cleanup, then sources a shared library that registers its own. What breaks?
    Bash keeps a single EXIT handler, so the library's registration replaces yours and your directory leaks. Shared libraries should not claim EXIT; if one must, it should append to a cleanup function or stack that callers own. In review, treat a bare `trap ... EXIT` inside a sourced file as a defect.
  • Does an EXIT handler change the exit status the script reports?
    Not by itself — the status of the terminating exit is preserved while the handler runs, and the script still reports it. It changes only if the handler calls `exit` explicitly, which is an easy way to turn a genuine failure into a reported success. Keep the handler side-effect free about status: clean up and return.

saying these in an interview costs you the question

  • Assumes the last line always runs
  • Registers the trap before the directory exists
  • Uses rm -rf "$dir"/* with an unguarded variable
  • Thinks EXIT covers kill -9
  • Adds a second trap ... EXIT expecting both to run

context

open as a page

In a bash script, what does `tmpfile=$(mktemp)` give you that `tmpfile=/tmp/myscript.$$` does not?

level: juniorimportance: should knowfreq 55%

basics

~20 s

mktemp creates the file itself, atomically, under a random name with owner-only permissions, and prints the path. A $$-based name is merely a predictable string that nothing has created yet, so another user can pre-create or symlink it first.

open as a page

A script refreshes a config file with `curl -sS https://example.com/config.json > /etc/app/config.json`, and readers occasionally load an empty or half-written file. What is happening, and what is the standard shell fix?

level: middleimportance: should knowfreq 38%

basics

~20 s

The redirection truncates the destination the moment it opens it, long before any bytes arrive, so readers see a growing file — or nothing at all if the download fails. Write to a temp file in the same directory and rename it over the destination.

open as a page

A cron script prevents overlapping runs with `if [ -f /var/tmp/job.pid ]; then exit 0; fi` and then writes its own PID into that file. After one run was killed with `kill -9`, the job never ran again. Why did that guard jam, and what mechanism would not have?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Nothing removes a PID file when the process dies abnormally, so the stale file blocks every later run forever. A kernel advisory lock taken with flock cannot go stale: the kernel drops it when the holder's file descriptors close, however the process died.

open as a page

You own a bash deployment script that colleagues rerun by hand after it fails halfway. What does making that script idempotent actually require, and how do you decide how far to take it?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Idempotence is about the end state, not the exit status: a second run must leave the system in the same shape as the first, whether or not the first finished. That means classifying every step and converting the ones that accumulate rather than converge.

open as a page