skip to content

A document-conversion worker holding an fcntl.flock lock is SIGKILLed — what does the next run find?

level: seniorimportance: should knowfreq 24%

answer

  1. Death closes everything the process held
  2. Crash and clean exit look identical
  3. The file stays, the lock does not
  4. One leaked child keeps it alive
  5. Shared storage moves the guarantee elsewhere

basics

~20 s

It finds the lock free. The kernel closes every descriptor of a dying process, which releases the advisory lock whether the exit was clean or a SIGKILL. The lockfile itself stays on disk; that is normal and needs no cleanup.

solid answer

~50 s

Nothing is stale. Process death closes the whole descriptor table, and an `fcntl.flock` lock is released when the last descriptor holding it closes, so a crash, an OOM kill and a clean `sys.exit` end identically from the lock's point of view. The next run acquires immediately and there is no recovery code to write — that asymmetry against a pidfile is the whole reason to use a lock. The lockfile stays on disk with whatever pid was last written into it; do not delete it and never trust its contents. The one real way a lock outlives its owner is descriptor inheritance: `os.fork` copies the descriptor table, so a forked child keeps the same open file description — and the lock — alive, and one leaked child can wedge the job indefinitely. On network storage the guarantee moves to the server's lock manager and its recovery window.

code

python · 17 lines
python
import fcntl, os, signal, sys, time

path = "/tmp/single-instance.lock"
pid = os.fork()
if pid == 0:
    fd = os.open(path, os.O_CREAT | os.O_RDWR, 0o644)
    fcntl.flock(fd, fcntl.LOCK_EX)
    time.sleep(30)
    sys.exit(0)

time.sleep(0.5)
os.kill(pid, signal.SIGKILL)
os.waitpid(pid, 0)

fd = os.open(path, os.O_CREAT | os.O_RDWR, 0o644)
fcntl.flock(fd, fcntl.LOCK_EX | fcntl.LOCK_NB)
print("lock re-acquired: the kernel dropped it when the process died")

go deeper

for a junior

Remember the headline: the kernel releases the lock when the process ends, so a killed job leaves nothing to clean up. The file staying on disk is expected and is not a stale lock.

for a middle

Explain the mechanism rather than the result: process exit closes every descriptor, and the lock goes when the last descriptor on that open file description closes. Note that a forked child inherits both.

for a senior

Diagnose from symptoms. Separate 'the job never runs' — inherited descriptor, a blocking acquire, storage-level lease recovery — from 'two copies ran', which points at an early close, a per-run path or two hosts on one mount.

for a principal

Decide where this guarantee is allowed to carry weight. Host-local locking is cheap and correct on one kernel; work that must be exactly-once across a fleet needs a lease or a scheduler that owns failure detection, and that boundary should be explicit in the design.

### Why the answer is "nothing to clean up" When any process terminates, the kernel closes every descriptor it had open. An `fcntl.flock` lock is attached to an open file description, and it is released when the last descriptor referring to that description closes. Put those two facts together and abnormal termination releases the lock for free: `SIGKILL`, an out-of-memory kill, a segfault in a C extension, a hard `os._exit`, and a tidy return from `main` are indistinguishable to the lock. The next scheduled run calls `fcntl.flock(fd, fcntl.LOCK_EX | fcntl.LOCK_NB)` and it simply succeeds. This is worth saying out loud in an interview, because it inverts the instinct that every acquired resource needs a matching release path. You do not need `atexit`, you do not need a signal handler, and you do not need a `finally` that calls `fcntl.LOCK_UN`. Those are harmless but they are not what makes the scheme correct — they cannot run under `SIGKILL`, and the scheme has to be correct exactly then. ### What stays behind, and what to do about it The lockfile itself. It is an ordinary file, usually holding the pid of whichever run wrote it last, and it survives forever. That is fine and intended: the file is only a name for the lock. Resist two temptations — unlinking it on shutdown, which races a newcomer that has already opened the same path onto a fresh inode and lets two runs each hold a valid lock, and reading the pid out of it to decide anything. Its contents are for a human running `cat`, never for the guard. ### The case where the lock genuinely outlives the process Descriptor inheritance. `os.fork` copies the descriptor table, and the copies refer to the *same* open file description, so parent and child share one lock. The lock is released only when every one of those descriptors is closed. If the worker forks a helper that outlives it — a runaway child, a process left behind by an incomplete shutdown — the parent can die and the lock stays held. The symptom is a job that never runs again while nothing obviously holds the file, and it is the most common way this pattern goes wrong in production. Two details sharpen the diagnosis. Since Python 3.4 (PEP 446) descriptors Python creates are non-inheritable by default, so a child started through the standard subprocess machinery does **not** inherit the lock across the exec — the close-on-exec flag handles that. But that flag only takes effect at exec: a `fork` with no exec inherits everything regardless, and so do `os.dup` and `os.dup2` copies. So "we start children through subprocess" is a real defence, and "we fork workers" is not. ### The scenario, worked A document-conversion queue runs a converter every minute under an `fcntl.flock` guard. Each start performs a partial-failure rollback: it finds documents left half-converted by an earlier crash and resets them. Two symptoms can arrive. *The job stopped running.* Check for a surviving child of a dead parent holding the descriptor, and for a `fcntl.LOCK_EX` call made without `fcntl.LOCK_NB` that is patiently blocking rather than exiting. A stale lock after a plain crash is **not** a valid hypothesis — the kernel does not leave one — so an operator who reports "the lock got stuck after the crash" is describing inheritance, a blocked waiter, or shared storage. *Two copies ran.* Now look for a released lock rather than a stuck one: a descriptor closed too early, a file object garbage-collected right after the call, a per-run lockfile path, a directory cleaner that removed the file between runs, or two hosts pointed at one mount. That last one matters here, because concurrent rollbacks are worse than no rollback: two converters each reset the other's in-flight documents, and a rendered-page cache that was sitting at an 83% hit rate collapses as every page is regenerated. ### Where the guarantee ends At the kernel that owns the file. Across a network filesystem the lock is mediated by the server's lock manager, which must decide on its own — after a lease timeout — that a crashed client no longer holds anything. There is a window, and its length is a property of the storage, not of your program. Anything that must be exactly-once across machines belongs to a lease or a scheduler that owns that problem, not to `fcntl`. ### The one-line summary A lock ends with its holder; a file does not. Every question about crash behaviour reduces to which of those two things you built your guard on — and to whether some other descriptor is still keeping the holder's open file description alive.

  • An operator says the lock got stuck after a crash. What are you actually looking for?
    Not a stale lock, because the kernel cannot leave one. Look for a surviving child or grandchild that inherited the descriptor across a fork and keeps the open file description alive, for a waiter that omitted the non-blocking flag and is blocked rather than exiting, or for a network mount where the server's lock manager has not yet expired the dead client's lease.
  • Does releasing the lock explicitly in a finally block add anything?
    Very little. It shortens the window in a long-lived process that continues doing unrelated work after the guarded section, which can be worth it. It contributes nothing to crash safety, since the paths you care about — SIGKILL, an OOM kill, a hard interpreter abort — never run finally blocks at all. The kernel is the release mechanism that always fires.
  • Why is deleting the lockfile during shutdown worse than leaving it?
    Unlinking races a process that has already opened the same path: it can hold a valid lock on an inode with no name, while the next start creates a fresh file at that path and locks that one. Two processes then both hold a real lock and both run. Leaving a permanent zero-byte file costs nothing and removes the race entirely.

saying these in an interview costs you the question

  • Claims a killed process leaves a stale advisory lock
  • Writes signal handlers to release the lock on crash
  • Deletes the lockfile during shutdown to tidy up
  • Forgets that a forked child inherits the lock
  • Trusts the pid written in the lockfile when diagnosing
  • Assumes the same guarantee holds across a shared mount

context