How does faulthandler.dump_traceback_later work as a watchdog for a wedged call?
answer
- The process is stuck, not dead
- A timer that survives a wedged interpreter
- Arm before the work, disarm after
- repeat and exit change what it does
- dump_traceback_later plus cancel_dump_traceback_later
basics
~10 sIt arms a watchdog thread that, after the given number of seconds, dumps every thread's Python stack and optionally kills the process. Call faulthandler.cancel_dump_traceback_later() when the guarded work finishes in time.
solid answer
~40 s`faulthandler.dump_traceback_later(timeout)` starts a small C watchdog thread that sleeps for `timeout` seconds and then writes the Python stack of every thread to `sys.stderr`. Because that thread is native and does not need the GIL, it still fires when the interpreter is wedged in a long C call or blocked on a lock, which is exactly the situation where nothing else in the process can report anything. `repeat=True` makes it dump on every interval instead of once, and `exit=True` makes it terminate the process immediately after the dump, turning it into a hard deadline rather than a report. The normal shape is to arm it before the risky work and call `faulthandler.cancel_dump_traceback_later()` in a `finally`, so the dump only ever lands when the work overran.
code
python · 10 linesimport faulthandler
import time
faulthandler.dump_traceback_later(5.0, exit=True)
try:
time.sleep(0.1) # the call being guarded
finally:
faulthandler.cancel_dump_traceback_later()
print("finished inside the budget")go deeper
Know that this arms a timer that prints every thread's Python stack if the code has not finished in time, and that you disarm it with faulthandler.cancel_dump_traceback_later() once the work returns.
Explain the arm/cancel pair around the guarded call, what repeat and exit change, and why a native watchdog thread can dump while a pure-Python timer callback would be starved of the GIL.
Show judgment on the timeout value and on exit=True: a hard kill skips cleanup, so it belongs only where a supervisor restarts the worker and a wedged process is worse than a dead one.
Decide where hang evidence belongs in the operational contract - which stream captures it, how long it is retained, and whether hard-killing stuck workers is the fleet default or an opt-in for a specific workload.
## The problem it solves A process that has crashed leaves a signal behind; a process that is simply stuck leaves nothing. Something is holding a lock, or a compiled extension has gone into a long computation, or a socket read has no timeout, and the process sits there consuming its own deadline. You cannot ask the interpreter what it is doing, because whatever would ask has to run Python, and Python is exactly what is not running. ## The mechanism `faulthandler.dump_traceback_later(timeout, repeat=False, file=sys.stderr, exit=False)` arms a **watchdog implemented as a native thread**. It sleeps for `timeout` seconds and, if it has not been cancelled, writes a header naming the timeout and then a stack for every thread in the process, most recent frame first, in the same austere format faulthandler uses for a fatal signal: thread identifier, then file, line and function per frame. The important property is that the watchdog thread is written in C and **does not take the GIL** to produce that dump - it reads frame state the interpreter has already recorded. That is why it works in the wedged case. A watchdog written in pure Python with `threading.Timer` cannot fire while another thread is inside a long C call that never releases the GIL, and that is precisely the hang you most want to see. ## The three knobs - **`repeat=True`** re-arms after each dump, so a process stuck for a long time produces a series of dumps and you can compare them: identical stacks mean genuinely blocked, moving stacks mean slow rather than stuck. - **`file=`** sends the output somewhere other than `sys.stderr`, and as with `enable()` the object must stay open for as long as the watchdog is armed, because the module holds the underlying file descriptor. - **`exit=True`** changes the character of the call completely: after dumping, the watchdog terminates the process immediately, without unwinding the stack, without running `finally` blocks and without `atexit` handlers. That is a deliberate hard kill, appropriate when a stuck worker is worse than a dead one and a supervisor will restart it, and inappropriate when the process holds state that has to be flushed. ## Cancelling, and why the finally matters `faulthandler.cancel_dump_traceback_later()` disarms the pending watchdog. The idiom is to arm before the guarded call and cancel in a `finally`, so the dump is only ever produced by an overrun. - Forgetting the cancel is **the classic bug**: the watchdog stays armed after the work finished, and with `exit=True` it will eventually kill a perfectly healthy process. - Calling `dump_traceback_later()` again while one is already armed replaces the pending timer rather than stacking a second one, which makes it usable as a per-iteration deadline in a loop - re-arm at the top of each iteration and you get a rolling watchdog. - Cancelling when nothing is armed is harmless. ## Choosing the timeout The number should sit **well above the worst latency you consider healthy**, not near the median, or the dumps become noise you learn to ignore. If a unit of work normally finishes inside a fraction of its budget and only the tail approaches it, arm the watchdog somewhere above that tail: you want a dump to mean "this is pathological", not "this was a slow day". A timeout is a float and may be fractional, which is useful in tests where you want a stuck call surfaced in a fraction of a second. ## Its limits - The watchdog **reports; it does not interrupt**. Unless you pass `exit=True` the stuck call keeps running after the dump, so this is not a timeout mechanism and cannot replace one - a socket timeout, a query timeout or a cancellation token is still the right way to bound work you control. - The dump also shows **Python frames only**: if the process is wedged inside a compiled extension, you see the Python call site that entered it, not what the native code is doing. - And like every faulthandler output it goes to a **raw file descriptor**, so it does not pass through the logging system, is not formatted as JSON, and will interleave with anything else writing to the same stream. Treat it as evidence written to the crash channel, not as an application log record.
- Why can faulthandler's watchdog produce a dump when a threading.Timer callback would never run?The watchdog is a native thread that does not acquire the GIL to write the dump; it reads frame state the interpreter already holds. A `threading.Timer` callback is Python code, so it needs the GIL, and if another thread is inside a long C call that never releases it the callback simply never runs. The hang you most want to see is exactly the one that starves a Python timer.
- What is the practical difference between dump_traceback_later(30) and dump_traceback_later(30, exit=True)?Without `exit`, the watchdog prints the stacks and the stuck call carries on running - you get evidence and nothing else. With `exit=True` the process is terminated immediately after the dump, skipping `finally` blocks and `atexit` handlers. Use it when a wedged worker is worse than a restarted one and a supervisor will bring it back; avoid it where the process holds unflushed state.
- What happens if you call dump_traceback_later again while a watchdog is already armed?The pending timer is replaced rather than duplicated, so you never accumulate watchdogs. That makes re-arming at the top of each loop iteration a clean rolling deadline. Cancelling when nothing is armed is also harmless, so the arm/cancel pair is safe to put in a helper or a context manager.
saying these in an interview costs you the question
- Calling it a timeout that aborts the slow call
- Forgetting to cancel, so a healthy process is killed later
- Believing repeated calls stack multiple watchdogs
- Assuming exit=True runs finally blocks and atexit handlers
- Expecting the dump to appear in the application log stream
- Setting the timeout near the median latency