Why is checking for an existing pidfile with os.path.exists a racy single-instance guard?
answer
- Two steps where one is needed
- The gap between looking and acting
- Atomic create fixes only half
- The marker outlives its writer
- Recycled pids fool the liveness probe
basics
~20 sBetween the os.path.exists check and the write, another copy can run the same check. Both see nothing, both create the file, both start. That gap is the race, and the file also outlives a killed process.
solid answer
~50 sThe check and the create are two separate steps, so two copies starting at the same moment can both see no pidfile, both write one, and both run — a classic time-of-check to time-of-use gap. Making the create atomic fixes only half of it: `os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY)` lets exactly one process win and raises `FileExistsError` in the loser, but the file is ordinary data that survives its writer, so one SIGKILL or power cut leaves a pidfile that blocks every later run. Probing the recorded pid with `os.kill(pid, 0)` narrows the stale case, and `ProcessLookupError` means the process is gone — but pids are recycled, so a stale entry can name an unrelated live process. Removing a stale file and recreating it is itself a race between two starters. An advisory lock avoids all of this because the kernel drops it when the holder dies.
code
python · 11 linesimport os
path = "/tmp/convert-worker.pid"
try:
fd = os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o644)
except FileExistsError:
print("pidfile exists: refusing to start")
else:
os.write(fd, str(os.getpid()).encode())
os.close(fd)
print("created pidfile for", os.getpid())go deeper
Be able to point at the gap between the existence check and the write and say what two processes do in that gap. Know that a file left on disk does not vanish when the process that wrote it is killed.
Explain the atomic alternative and its flag pair, what exception the loser sees, and why atomicity still leaves the stale-file failure. Describe the liveness probe and where pid reuse defeats it.
Argue for moving the state into something the kernel releases on death, and describe the operational symptom you would see in production: a periodic job that stops running entirely after one hard kill, with no error anywhere.
Frame it as a class of defect — markers whose lifetime is not tied to the thing they represent — and set the standard for the team's job templates so every scheduled job inherits a guard that cannot wedge itself.
### The two-step problem The naive guard reads: ```python import os path = "/tmp/convert-worker.pid" if os.path.exists(path): # step 1: check raise SystemExit("already running") with open(path, "w") as f: # step 2: act f.write(str(os.getpid())) ``` Between step 1 and step 2 the process can be descheduled for an arbitrarily long time. A second copy started in that window runs step 1, also sees nothing, and also proceeds. Both write the file — the second write simply overwrites the first — and both do the work. This is time-of-check to time-of-use: the state you tested is not the state you acted on, because nothing held it still in between. The window is small, which is exactly what makes the bug nasty: it survives every manual test and appears the first time two copies are started together, often by a scheduler firing while a slow previous run is still finishing. ### Making the create atomic The kernel can do check-and-create as one indivisible step. `os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o644)` creates the file only if it does not already exist, and fails otherwise — Python raises `FileExistsError`. Exactly one racing process gets a descriptor; every other one gets the exception. That closes the race in step 1, and it is the correct primitive whenever you need "create this only if nobody has". ### What atomicity does not fix The file is now created safely, but it is still ordinary data on a filesystem, and its lifetime has nothing to do with the process that wrote it. The instant the job dies without running its cleanup — SIGKILL, an OOM kill, a crashed interpreter, a lost power supply — the file remains. Every subsequent start sees it, concludes another copy is running, and refuses. The guard has failed closed and the job silently stops happening until a human deletes the file. That is the fundamental defect of the whole family: **the marker outlives the thing it is supposed to represent.** ### Liveness probing, and why it only narrows the gap The usual repair is to read the pid out of the file and ask whether that process still exists: ```python import os def pid_alive(pid): try: os.kill(pid, 0) # signal 0: existence check, no signal sent except ProcessLookupError: return False # nothing with that pid except PermissionError: return True # exists, owned by another user return True ``` Signal `0` performs the permission and existence checks without delivering anything, so `ProcessLookupError` is a reliable "no such process". Three problems remain. *Pid reuse.* Pids are recycled. A stale file naming pid 4711 can match a completely unrelated process that started later, and the probe then reports the job as running forever. The usual mitigation — also compare a start time or a command line — pushes the guard further into platform-specific territory for a property the kernel could have given you for free. *The removal race.* Once two starters both decide the pidfile is stale, both unlink it and both create a new one. You are back to the original race, one layer down. *Meaningless pids across boundaries.* `os.getpid()` returns the pid in the caller's own pid namespace. A pid written inside a container and read outside it, or read after a restart that reset the namespace, refers to nothing comparable. A number in a file is not an identity. ### What to do instead Let the kernel hold the state. An advisory lock taken with `fcntl.flock` on a fixed file is released automatically when the holding process ends for any reason, so there is no stale case to detect, no liveness probe, and no removal race. Keep writing the pid into the file if you like — after taking the lock — purely so an operator can see who holds it. The distinction to carry into the interview is this: a pidfile records a **claim**, and claims go stale; a lock is **held**, and holding ends with the holder. The pattern also generalises. `os.O_CREAT | os.O_EXCL` is the right tool for one-shot creation — a marker written once, a uniquely named temporary file — precisely because it is atomic; it is the wrong tool for mutual exclusion over the lifetime of a process, because file existence has no lifetime tied to a process.
- Which open flags turn the create step into a single atomic operation?os.O_CREAT together with os.O_EXCL, passed to os.open. The kernel creates the file only if the path does not already exist and fails otherwise, so exactly one of several racing processes gets a descriptor and the rest see FileExistsError. It is the standard way to express create-only-if-absent, and it is why the loser learns it lost without a separate check.
- Why does os.kill(pid, 0) not fully solve the stale-pidfile problem?It answers whether some process with that pid exists right now, not whether it is your job. Pids are recycled, so a stale file can name an unrelated live process and the guard blocks forever. It also leaves the removal race intact: two starters can both decide the file is stale and both recreate it.
It is like claiming a library desk by leaving a coat on it: two people can drop coats in the same second, and a coat left behind when someone is carried out of the building keeps the desk reserved forever.
saying these in an interview costs you the question
- Treats check-then-create as if it were one operation
- Believes a pidfile disappears when the process is killed
- Assumes a pid in a file uniquely identifies the job
- Deletes a stale pidfile without noticing the second race
- Thinks a shorter window between check and write removes the race