skip to content

How do you use fcntl.flock to guarantee only one copy of a Python job runs?

level: middleimportance: must knowfreq 40%

answer

  1. One copy only, enforced by the kernel
  2. A file you open but never read
  3. Ask for it without waiting
  4. Exclusive plus non-blocking flags
  5. Hold the descriptor for the whole run

basics

~20 s

Open a fixed lockfile with os.open, then call fcntl.flock(fd, fcntl.LOCK_EX | fcntl.LOCK_NB). A BlockingIOError means another copy holds it, so exit. Keep the descriptor open for the whole run; the kernel releases the lock when the process dies.

solid answer

~40 s

Open a fixed path with `os.open(path, os.O_CREAT | os.O_RDWR)` — never a truncating mode — and take an exclusive, non-blocking advisory lock with `fcntl.flock(fd, fcntl.LOCK_EX | fcntl.LOCK_NB)`. If another copy already holds it the call raises `BlockingIOError`, and the right response for a periodic job is usually to exit quietly rather than queue up behind it. The descriptor must stay open for the entire run: closing it, or letting a file object be garbage-collected, releases the lock immediately. Nothing else is needed on shutdown, because the kernel drops the lock when the process ends for any reason, including a crash or SIGKILL. Do not delete the lockfile on exit — unlinking it while another process is opening the same path lets two runs each hold a lock on a different inode.

code

python · 11 lines
python
import fcntl, os, sys

fd = os.open("/tmp/convert-worker.lock", os.O_CREAT | os.O_RDWR, 0o644)
try:
    fcntl.flock(fd, fcntl.LOCK_EX | fcntl.LOCK_NB)
except BlockingIOError:
    print("another copy is already running; nothing to do")
    sys.exit(0)
os.ftruncate(fd, 0)
os.write(fd, str(os.getpid()).encode())
print("holding the lock as pid", os.getpid())

go deeper

for a junior

Recall the shape: one fixed lockfile, an exclusive non-blocking lock, exit when it fails. Know that the file's existence is not the lock — the lock is something the kernel holds while your process lives.

for a middle

Explain each flag: why exclusive rather than shared, why non-blocking rather than waiting, and why the descriptor must stay open for the whole run. Be able to name the exception raised when another copy holds it.

for a senior

Show the operational judgement: exit quietly or exit non-zero, where the lockfile lives so it survives reboots and tmp cleaners, why you never unlink it, and where the guarantee stops when the path sits on shared storage.

for a principal

Own the policy across a fleet: which jobs genuinely need exactly-once-per-host versus idempotent work that can overlap safely, and when host-local locking should be replaced by a scheduler or lease that spans machines.

### The guarantee you actually want A job that must not overlap itself — a minute-by-minute cron entry, a queue drainer, a nightly sweep — needs mutual exclusion whose lifetime is tied to a **process**, not to a file's existence. On a single Unix host `fcntl.flock` is that primitive: the lock is kernel state attached to an open file description, and the kernel tears it down when the holding process goes away, however it goes away. ### The idiom ```python import fcntl, os, sys fd = os.open("/tmp/convert-worker.lock", os.O_CREAT | os.O_RDWR, 0o644) try: fcntl.flock(fd, fcntl.LOCK_EX | fcntl.LOCK_NB) except BlockingIOError: sys.exit(0) # another copy is running; this tick is a no-op ``` Four decisions are packed into those lines. **`os.O_CREAT | os.O_RDWR`, never a truncating mode.** The file must exist before it can be locked, and `os.O_CREAT` creates it when absent without disturbing it when present. Opening with mode `"w"` (`os.O_TRUNC`) blanks a file the running instance may have written its pid into — and the truncation happens at open time, *before* you have the lock, so it happens even in the copy that is about to lose the race. **Hold a descriptor, not a value you drop.** `fcntl.flock(open(path, "a"), fcntl.LOCK_EX)` looks fine and is useless: no reference survives the call, CPython closes the file object on the spot, and closing releases the lock. Bind the descriptor (or the file object) to a name that lives as long as the work does, at module scope or on a long-lived object. **`fcntl.LOCK_EX`, not `fcntl.LOCK_SH`.** A shared lock is compatible with other shared locks, so two copies asking for `fcntl.LOCK_SH` both succeed and both run. Only the exclusive mode gives single-instance. **`fcntl.LOCK_NB`.** Without it the second copy *blocks* until the first finishes. For a job fired every minute that turns a slow run into a growing queue of waiting processes, all of which eventually run back-to-back. With it, the call fails immediately with `BlockingIOError` (an `OSError` subclass, `EWOULDBLOCK`) and you choose the policy: skip the tick, log, or exit non-zero so a supervisor notices. ### Lifetime, precisely A `flock` lock belongs to the open file description, not to the file and not to the process. It is released when the last descriptor referring to that description is closed, when you call `fcntl.flock(fd, fcntl.LOCK_UN)`, or when the process terminates — cleanly, on an unhandled exception, or under SIGKILL. There is no cleanup handler to write and no stale state to detect on the next start. That single property is why this beats a hand-rolled pidfile, whose contents survive the process that wrote them. ### Advisory, not mandatory "Advisory" means the kernel only holds back processes that *also* lock the same file. Anything that just opens the path and writes it succeeds regardless. For single-instance that is fine, because every participant is your own program; but it means the lockfile is a token, not protection for data. Keep the lockfile dedicated to locking — do not lock the data file the job is writing. ### Things that quietly break it *Deleting the lockfile on exit.* Process A unlinks the file as it finishes while B has already opened that path and is holding the lock on the now-unlinked inode; C then creates a fresh file at the same path and locks that. B and C both believe they are the only instance. Leave the file in place — an empty file per job costs nothing. *A per-run path.* Locking a file whose name contains a timestamp or a pid excludes nothing. The path must be fixed and shared by every copy, and it must be on a filesystem that survives — a directory cleaned between runs reintroduces the same problem. *Assuming it reaches other hosts.* The guarantee is one kernel. Advisory locks on a network filesystem are mediated by the server's lock manager with its own recovery windows; single-instance across machines is a lease problem, not an `fcntl` one. *Portability.* There is no `fcntl` module on Windows; the locking call there is a different, platform-specific API, so a cross-platform tool needs a branch. ### Writing the pid in After locking, `os.ftruncate(fd, 0)` and `os.write(fd, str(os.getpid()).encode())` make the file self-describing for an operator running `ls` and `cat`. That is diagnostics only: correctness never reads the contents back, which is exactly what makes the scheme immune to the stale-content problems of a pidfile-driven check.

  • Why must the file descriptor stay open instead of being closed once fcntl.flock returns?
    The lock lives on the open file description, so closing the last descriptor that refers to it releases the lock immediately. Closing right after acquiring leaves the job running with no protection at all, and the failure is invisible until two copies overlap. Bind the descriptor to something that outlives the work, and let process exit do the releasing.
  • Should the job unlink its lockfile when it finishes?
    No. Unlinking creates a race: one process can be holding the lock on an inode that has just been unlinked while a newcomer creates a fresh file at the same path and locks that, so two processes each hold a valid lock. Leave the file on disk permanently; only the lock state matters, and the kernel manages that.
  • What happens if a second copy asks for fcntl.LOCK_SH instead of fcntl.LOCK_EX?
    A shared lock conflicts with an exclusive one but not with another shared one. So the second copy is still blocked by a running holder of fcntl.LOCK_EX, but if both copies were written to use fcntl.LOCK_SH they would both acquire and both run. Single-instance requires the exclusive mode on every participant.
  • What is the right exit code when the lock is already held?
    It depends on who is watching. For a periodic job where overlapping is expected and harmless to skip, exit 0 with a log line, so a supervisor or cron does not send alert mail every minute. If two starts should never happen — a service under a supervisor, say — exit non-zero so the failure is visible rather than silently swallowed.

It is a hotel key that the front desk takes back automatically the moment you leave the building — as opposed to a note on the door saying "occupied", which stays up after you are gone.

saying these in an interview costs you the question

  • Checks whether the lockfile exists instead of locking it
  • Closes the descriptor right after locking, dropping the lock
  • Passes open(path) inline, so the file object is collected immediately
  • Deletes the lockfile on exit, letting two runs both win
  • Omits the non-blocking flag, queueing runs instead of skipping
  • Thinks an advisory lock stops a process that never locks

context