A cron script prevents overlapping runs with `if [ -f /var/tmp/job.pid ]; then exit 0; fi` and then writes its own PID into that file. After one run was killed with `kill -9`, the job never ran again. Why did that guard jam, and what mechanism would not have?
answer
- who deletes the file when nobody runs
- kill -9 runs no cleanup code
- test and write are two steps
- PIDs get reused
- the kernel releases it for you
basics
~20 sNothing removes a PID file when the process dies abnormally, so the stale file blocks every later run forever. A kernel advisory lock taken with flock cannot go stale: the kernel drops it when the holder's file descriptors close, however the process died.
solid answer
~50 sThe guard is a promise the dead process could not keep. The PID file is removed by the script's own cleanup, so a `kill -9`, an OOM kill or a power cut leaves it behind, and from then on every run sees the file and exits — silently, because it exits 0. Two more bugs lurk in the same three lines: `[ -f ]` and the subsequent write are separate steps, so two copies starting together can both pass the test, and PIDs are reused, so a later `kill -0` check can find an unrelated live process. `flock` fixes all three. It takes an advisory lock on an open file descriptor, and the kernel releases that lock when the descriptor closes — on exit, on a crash, on `SIGKILL`, on reboot. Acquisition is atomic, so there is no test-then-set race.
code
bash · 13 lines#!/usr/bin/env bash
set -euo pipefail
lockfile=/var/lock/myjob.lock
exec 9>"$lockfile"
if ! flock -n 9; then
printf 'another copy is already running; skipping\n' >&2
exit 0
fi
printf 'doing the work\n'
sleep 1go deeper
Know that a leftover lock or PID file blocks every future run once a script dies without cleaning up, and that the shell tool for single-instance jobs on Linux is flock rather than a hand-written file check.
Explain the three defects — the stale file, the gap between testing and writing, and PID reuse — and write the flock idiom correctly: open a descriptor on a lock file, flock -n on it, and never delete the file afterwards.
Show the operational judgment: choose deliberately between skipping quietly, failing loudly and waiting with a bounded timeout, log every skip so a silent no-op is not mistaken for health, and know where the lock's guarantee ends.
Own the question of where mutual exclusion belongs at all: a per-host file lock, a lease in a coordination service, or a scheduler that guarantees a single runner — and decide how overlapping runs are detected and alerted on across the fleet.
## Why the guard jammed A PID file is a *record* that a process claims to be running. Removing it is the responsibility of that process, which means the record survives precisely the failures you most want to survive: `kill -9`, the OOM killer, a container being torn down, a host losing power. After that, the file is a permanent gate. Because the script exits 0 when it finds the file, cron sends no mail and no monitor fires — the job simply stops happening, and it is usually discovered days later by the absence of its output. ## Three separate defects in the hand-rolled version **Staleness.** Just described. A cleanup on exit narrows the window but never closes it, because the point of `SIGKILL` is that no code runs. **A test-then-set race.** `[ -f file ]` and the write that follows are two operations with a gap between them. Two copies launched in the same second can both find no file, both write their PID, and both run. The guard is at its weakest exactly when it is being tested hardest. **PID reuse.** The usual patch is to validate the recorded PID with `kill -0 "$pid"`, which only asks "does a process with this PID exist and may I signal it". PIDs are recycled, so after enough churn some unrelated process now owns that number and the guard blocks again. Comparing the process's command line narrows it but is still heuristic, and `kill -0` also fails with a permission error for a live process owned by someone else, which naive scripts read as "not running". A `mkdir /var/tmp/job.lock` guard fixes only the race: directory creation is atomic and fails if the name exists. It still goes stale for exactly the same reason. ## What flock does instead `flock` (from util-linux) takes an advisory lock associated with an *open file description*. Two properties follow, and they are the whole answer: - **Acquisition is atomic.** There is no window between checking and taking; the kernel either grants the lock or does not. - **Release is automatic.** The lock lives with the open descriptor, so when the process ends for any reason — clean exit, crash, `SIGKILL`, the machine rebooting — the kernel closes the descriptor and the lock is gone. There is no stale state to reap, and no cleanup handler to get right. The two idioms worth knowing: ```bash # 1. Lock inside the script, on a dedicated descriptor exec 9>/var/lock/myjob.lock flock -n 9 || { echo "already running" >&2; exit 0; } # 2. Wrap the whole script from the crontab line flock -n /var/lock/myjob.lock /usr/local/bin/myjob.sh ``` Useful flags: `-n` fails immediately instead of blocking; `-w SECONDS` waits with a bound; `-E CODE` sets the exit code used when `-n` or `-w` gives up; `-s` takes a shared rather than exclusive lock; `-c` runs a command string. ## Decide what "already running" should mean This is the judgment part, and interviewers push on it. Exiting 0 is quiet and right for a job that runs every minute and simply skips a beat. Exiting non-zero surfaces the overlap — as cron mail, or as a failed unit the service manager records — which is right when overlapping means you are falling behind and want to know. A bounded `-w 30` suits a job that is nearly always short and where waiting is better than skipping. Whatever you choose, log the skip; a job that silently does nothing is indistinguishable from a job that is broken. ## Practical rules Put the lock file on a path that persists and is writable by the job's user — `/var/lock/` on Linux systems, or a directory the service owns. Do **not** delete the lock file at the end: removing it while another process holds a descriptor on it creates a fresh race where two runs hold locks on two different inodes with the same name. An empty file left in place forever is the correct steady state. One lock protects one logical job, not one host — if the same script guards different datasets, put the identifier in the lock file's name. Avoid relying on `flock` across NFS, where the guarantees get murky. And note the portability boundary: `flock(1)` is util-linux, present on mainstream Linux distributions and as a BusyBox applet in Alpine, but macOS does not ship it, so a script that must run on a developer's Mac needs a different approach there.
- Why should the lock file not be deleted when the script finishes?The lock belongs to the open descriptor, not to the name. Removing the file lets a second run create a fresh file with the same name and lock *that* inode, so two processes hold locks that never conflict and both proceed. Leave the empty file in place permanently — it costs nothing and keeps every run locking the same object.
- How would you let a job wait briefly for the lock instead of skipping the run outright?Use `flock -w SECONDS` instead of `-n`: it blocks up to that bound and then gives up with a failure status you can act on, and `-E CODE` lets you pick a distinctive exit code for "timed out waiting". It suits jobs that are usually fast, where a short queue is better than a skipped run. Keep the bound well under the invocation interval so waiters cannot pile up.
- Does flock protect against a second copy that does not use flock at all?No — it is advisory. The lock is only honoured by processes that ask for it, so a manual run of the same script without the wrapper, or a different tool touching the same data, is unaffected. The guarantee is a convention among cooperating callers, which is why the lock belongs inside the script rather than only in the crontab line.
- How do you make the guard cover a job that runs on several hosts against one shared dataset?A local file lock cannot, because each host has its own filesystem view. You need a lock the participants share: a row or advisory lock in the database being written, a lease in a coordination service, or a scheduler that runs the job in exactly one place. Say so plainly — stretching flock over a network filesystem is where people get burned.
saying these in an interview costs you the question
- Adds a cleanup to remove the PID file and calls it fixed
- Validates staleness with kill -0 and trusts the PID
- Treats [ -f lock ] then touch as atomic
- Deletes the lock file at the end of the run
- Assumes flock stops a copy that does not lock