skip to content

How do you confirm a live Python process is deadlocked on two threading locks?

level: seniorimportance: nice to knowfreq 22%

answer

  1. A hang with no traceback
  2. CPU near zero, queue growing
  3. You need every thread's stack
  4. Arm a timer that dumps tracebacks
  5. Two acquires, opposite order

basics

~20 s

Dump every thread's stack from inside the process — faulthandler.dump_traceback_later(30, repeat=True), or faulthandler.register on a signal — and look for two threads parked in acquire() at different call sites. Zero CPU with no progress and no exception is the giveaway.

solid answer

~40 s

A lock-order deadlock looks like a hang, not a crash: CPU near zero, no log lines, no traceback, and the process still answering nothing forever. Confirm it by getting stacks for **all** threads. Arm `faulthandler.dump_traceback_later(timeout, repeat=True)` at startup so a stuck process prints every thread's stack automatically, or `faulthandler.register(signal.SIGUSR1, all_threads=True)` so you can dump on demand; `sys._current_frames()` with `traceback.format_stack` does the same from a watchdog thread. Since 3.14, PEP 768 lets you attach to a running process with `sys.remote_exec` even if nothing was armed in advance. The signature is two threads blocked in `Lock.acquire`, each holding the lock the other wants. The fix is a fixed global acquisition order everywhere, or fewer locks — and `acquire(timeout=...)` so an inversion fails loudly instead of hanging.

code

python · 5 lines
python
import faulthandler
import signal

faulthandler.register(signal.SIGUSR1, all_threads=True)
print("send SIGUSR1 to this process to dump every thread's stack")

go deeper

for a junior

Know the symptom rather than the tooling: a deadlocked Python process hangs with no error and near-zero CPU, and there is no way to kill just the stuck thread. Naming the ordering rule as the fix is enough here.

for a middle

Explain how you would obtain stacks for every thread, not just the main one, and what an inversion looks like in that dump — two acquisitions at different call sites, each inside the other's critical section.

for a senior

Show that you would have armed the diagnostics before the incident, and that you would rather have a timeout fail loudly than a lock hang silently. Argue for collapsing lock pairs or handing work to one owner thread instead.

for a principal

Own the convention: a documented lock ordering, a ban on calling foreign code under a lock, and mandatory hang diagnostics in the service template, so no single incident depends on one engineer remembering the right incantation.

### The symptom A lock-order deadlock in Python is a silent hang. There is no exception, no log line, and no `RuntimeError`: the threads involved are blocked inside a C-level lock acquisition, consuming no CPU. Process CPU drops to whatever the unaffected threads are doing, memory stays flat, and any queue in front of the service grows without bound. The signature that distinguishes it from an infinite loop is exactly that CPU floor — a spinning bug burns a core, a deadlock burns nothing. ### Getting stacks out of a live process The problem is that the usual debugging reflex — read the traceback — has nothing to read. You need the stacks of *every* thread, on demand, from a process that will not respond. **Armed in advance, the cheapest option.** `faulthandler` is in the standard library and costs nothing while idle: ```python import faulthandler faulthandler.dump_traceback_later(30, repeat=True, exit=False) ``` If the process fails to cancel that timer within 30 seconds, it prints every thread's stack to standard error and keeps doing so. Placed in a watchdog loop that cancels and re-arms after each successful cycle, it turns a hang into a stack dump you can read in the logs. **On demand, via a signal.** `faulthandler.register(signal.SIGUSR1, all_threads=True)` makes the process dump all thread stacks whenever you send it that signal. That is the one line worth adding to every long-lived Python service before it ever hangs. **From inside, without faulthandler.** `sys._current_frames()` returns a mapping of thread id to the frame each thread is currently executing; pair it with `traceback.format_stack` and `threading.enumerate()` to print named stacks from a watchdog thread that is itself not blocked. **Nothing armed at all.** CPython 3.14 ships PEP 768 remote debugging: `sys.remote_exec` injects a script into an already-running interpreter, so you can attach to a process that was started with no instrumentation and have it print its own thread stacks. An out-of-process stack sampler is the other classic route when you cannot modify the target. ### Reading the dump You are looking for two or more threads whose innermost frame is a lock acquisition — a `with some_lock:` line, or an explicit `.acquire()` — at *different* call sites. Then trace outward: each blocked thread's stack shows which locks it already holds, because those are the `with` blocks it is currently inside. When thread A holds lock 1 and is waiting for lock 2, while thread B holds lock 2 and is waiting for lock 1, you have a lock-order inversion and the diagnosis is complete. If instead one thread is blocked on `Queue.get` and everything else is idle, the bug is a lost producer, not a deadlock. ### Why it happened The most common causes in Python code are: - **Two locks taken in opposite orders** by two code paths that were written months apart, often because one of them acquires the second lock inside a helper the author did not read. - **Calling unknown code while holding a lock** — a callback, a `__del__`, a logging handler, a property that lazily initialises something. That foreign code takes its own lock and the ordering is now out of your hands. - **Re-entering the same non-reentrant lock** on one thread, which is a self-deadlock and shows as a single thread parked in `acquire` with the same lock further up its own stack. Note what does *not* fix an ordering inversion: switching both locks to a reentrant lock. Reentrancy only lets one thread re-acquire a lock it already owns; two threads waiting on each other are unaffected. Candidates offer this fix often and it is a reliable signal. ### Preventing it 1. **One lock.** The cheapest fix is to notice that two locks were protecting one invariant and collapse them. Fewer locks, fewer orderings. 2. **A fixed global order.** Write down the order in which locks may be acquired and follow it everywhere. Where the objects are dynamic, sort them by a stable key — `for lock in sorted(locks, key=id)` — so any two code paths acquire the same pair in the same sequence. 3. **Never hold a lock across foreign code.** Copy the data you need, release, then call out. This also keeps critical sections short, which is worth doing anyway. 4. **Acquire with a timeout.** `lock.acquire(timeout=5)` returning `False` lets you log a clear "suspected lock-order inversion" with the current stack rather than hanging. A loud failure at 3 a.m. beats a silent one. 5. **Prefer designs with no lock pair at all.** Handing work to one owner thread through a `queue.Queue` removes the second lock entirely; confinement removes both. ### Why this is a differentiator rather than a gate Most Python services never deadlock, because most of them use one lock or none. Knowing the `faulthandler` incantations is not something a strong candidate should be marked down for missing — but a candidate who has actually run into a hung process, and can describe how they got stacks out of it, is showing production experience that is hard to fake.

  • Why does switching both locks to a reentrant lock not fix a two-lock inversion?
    Reentrancy only permits the owning thread to acquire a lock it already holds, which fixes self-deadlock from re-entering the same lock. In an inversion, two different threads each hold what the other wants, and neither owns the lock it is blocked on, so reentrancy changes nothing. The fix has to be an ordering rule, a single lock, or a timeout that turns the hang into a reported failure.
  • The hang only reproduces in production. What do you add now so the next occurrence is diagnosable?
    Register an all-thread stack dump on a signal at startup, and run a watchdog that arms `faulthandler.dump_traceback_later` and cancels it after each healthy cycle, so a stall prints stacks by itself. Add timeouts to every acquisition that can nest, logging the current stack when one expires. On 3.14 you can also attach to the live process with the remote-execution facility if nothing was armed.
  • How do you tell a lock deadlock apart from a thread stuck waiting on a queue?
    Read the innermost frames. Two threads parked in lock acquisition at different call sites, each inside a `with` block for the other's lock, is an inversion. A single thread parked in a queue's `get` with every other thread idle is a starved consumer — the producer died or never sent anything — and the fix is upstream, not in the locking.

Two people each holding one half of a pair of scissors, each waiting for the other half. Nothing is broken and nobody is busy — you only find out by walking in and looking at both of their hands at once.

saying these in an interview costs you the question

  • Expects a traceback or exception from a deadlocked process
  • Reads only the main thread's stack
  • Suggests a reentrant lock as the fix for an inversion
  • Blames the GIL for the hang
  • Proposes killing the stuck thread, which Python cannot do
  • Acquires locks in whatever order each call site finds convenient

context