skip to content

How does faulthandler.register arm a Python process to dump every thread's stack on demand?

level: middleimportance: should knowfreq 42%

answer

  1. Arm it once at startup
  2. A spare signal nothing else claims
  3. The handler is C, not Python
  4. It does not wait for the eval loop
  5. The destination descriptor is captured at registration

basics

~20 s

faulthandler.register(signal.SIGUSR1) installs a C-level handler for a spare signal, so sending that signal makes the process print every thread's Python stack to a file and keep running. It is Unix-only, and it works even while Python code is blocked.

solid answer

~50 s

You call `faulthandler.register(signal.SIGUSR1, all_threads=True)` once at startup, on a signal your application does not otherwise use. From then on, sending that signal from outside — the process keeps running afterwards — makes CPython print the Python stack of every thread to the stream you registered, `sys.stderr` by default. The handler is installed in C and walks each thread's frames directly, so it does not need to execute Python bytecode or wait for a thread to reach a safe point; that is why it still produces a dump when the workers are wedged. Two details bite in production: the destination's file descriptor is captured at registration time, so closing or rotating that stream silently loses future dumps, and `chain` defaults to `False`, meaning any handler you had previously installed for that signal is no longer called. `faulthandler.register` is Unix-only.

code

python · 17 lines
python
import faulthandler
import os
import signal
import threading
import time


def render_page():
    time.sleep(30)


faulthandler.register(signal.SIGUSR1, all_threads=True)
threading.Thread(target=render_page, name="render-0", daemon=True).start()
time.sleep(0.2)
os.kill(os.getpid(), signal.SIGUSR1)
time.sleep(0.2)
print("still running")

go deeper

for a junior

Know that a running Python process can be made to print every thread's stack by sending it a signal, that it keeps running afterwards, and that the call to arm it belongs at startup rather than at the moment of trouble.

for a middle

Explain the mechanics: the handler is installed in C, all_threads defaults to true, the output stream's descriptor is captured at registration, and chain decides whether a previously installed handler still runs.

for a senior

Demonstrate the judgement — why a Python-level handler is useless on a wedged interpreter, why the signal must be one nothing else claims, and why arming it unconditionally beats attaching a debugger to a process you did not predict would misbehave.

for a principal

Own it as a platform decision: diagnostics armed by default in every service image, an agreed signal and destination across teams, and a documented procedure so an on-call engineer takes two dumps before restarting rather than destroying the evidence.

## The shape of the mechanism ```python import faulthandler, signal faulthandler.register(signal.SIGUSR1, all_threads=True) ``` One call at startup, and from then on the process has an on-demand introspection channel. Somebody sends the process `SIGUSR1` from a shell or an orchestrator, CPython prints a block per thread — thread id, the thread's `Thread.name` in brackets on 3.14, then that thread's Python frames innermost-first — and **the process carries on running**. Nothing is killed, nothing is unwound, no exception is raised into your code. The full signature is `faulthandler.register(signum, file=sys.stderr, all_threads=True, chain=False)`, and `faulthandler.unregister(signum)` reverses it. ## Why it works when the interpreter is stuck This is the part that matters and the part candidates usually get wrong. A Python-level signal handler installed with `signal.signal` is not really run by the operating system: the C-level handler CPython installs just sets a flag, and your Python callback runs later, in the **main thread**, when the eval loop next checks for pending signals. If the main thread is blocked in a lock acquisition, or another thread is inside a long C call holding the global interpreter lock, that check never happens and your handler never runs. A diagnostic that requires a healthy interpreter is useless on a wedged one. `faulthandler`'s handler is different: it is written in C and does its work inside the signal handler itself, walking the interpreter's per-thread state and printing frames with low-level writes rather than the Python I/O stack. It does not need to acquire the interpreter lock or run bytecode. That is the entire reason to prefer it for hang diagnosis. A related consequence: `faulthandler.register` can be called from any thread, whereas `signal.signal` raises `ValueError: signal only works in main thread of the main interpreter` off the main thread. `faulthandler` installs its handler through the C API, so a library that arms it from a worker still works. ## Choosing the signal Use a signal that means nothing else to your program. `signal.SIGUSR1` and `signal.SIGUSR2` exist for exactly this. Do not reuse: - `signal.SIGINT`, which is your keyboard interrupt and normally raises `KeyboardInterrupt`; - `signal.SIGTERM`, which your orchestrator uses for graceful shutdown; - `signal.SIGQUIT`, whose default action terminates the process and writes a core dump. If something else in the process legitimately handles your chosen signal too, pass `chain=True` so `faulthandler` calls the previous handler after dumping. The default `chain=False` silently replaces it — a good way to break shutdown if you register on `SIGTERM`. ## Where the output goes, and the descriptor trap `file` defaults to `sys.stderr`, and **the file descriptor is captured when you register, not when the signal arrives**. Consequences worth stating in an interview: - If you register with an open log file and then close it, later dumps go nowhere. The process survives, you see no error, and you conclude the mechanism does not work. - Worse, if that descriptor number is later reused by a new file, the dump is written into an unrelated file. - If a supervisor replaces `sys.stderr` after startup, dumps still go to the original stream. So register **after** you have finalised where standard error points, keep a reference to any dedicated file object you pass, and re-register if you rotate it. A dump landing in the container's standard error is usually the right answer precisely because nothing rotates it out from under you. ## What the dump does and does not show It shows **Python frames only**. A thread blocked inside a compiled extension's C function is shown at the Python frame that called into the extension — accurate, but it will not name the C function you are actually stuck in. `all_threads=True` (the default) covers every thread in the interpreter, not just the one the operating system happened to deliver the signal to, so it does not matter which thread receives it. ## Operational shape Arm it unconditionally at startup in any long-running service: the cost is one installed signal handler and nothing at all until the signal arrives. That is what makes it different from attaching a debugger — there is no runtime overhead to justify, no privileges to arrange in the moment, and no need to have predicted which process would misbehave. When a worker wedges, you take a dump, wait, take a second one, and compare. The alternative, restarting the process, destroys the only evidence you had. `faulthandler.register` is documented as Unix-only; on Windows you fall back to a Python-level dumper built on `sys._current_frames()`, which carries its own limitation.

  • Why not just install a Python-level handler with signal.signal that prints the stacks?
    Because CPython defers Python-level signal handlers: the C handler sets a flag and your callback runs later in the main thread, when the eval loop next checks. If the main thread is blocked, or another thread is inside a long C call holding the interpreter lock, that check never comes and the handler never runs — precisely the situation you armed it for. `faulthandler` prints from inside the C signal handler and needs neither the lock nor the eval loop.
  • You registered the dump to a dedicated log file and later see no output at all. What is the likely cause?
    The file descriptor is captured at registration time. If the file object was closed or garbage-collected, or a supervisor rotated the file, the dump is written to a descriptor that is closed or has been reused for something else. There is no error — the process survives and the output vanishes. Keep a reference to the file, re-register after rotation, or simply dump to standard error.
  • Can faulthandler.register be called from a worker thread?
    Yes. It installs its handler through the C API, so unlike `signal.signal` — which raises `ValueError: signal only works in main thread of the main interpreter` off the main thread — it can be armed from anywhere. That matters for libraries and plugins that want to arm diagnostics without owning process startup, though arming it once in the main path is still the clearer design.
  • What does chain=True change?
    By default `chain` is `False` and the previously installed handler for that signal is replaced, so it stops being called. With `chain=True`, `faulthandler` dumps and then invokes the previous handler as well. You need it whenever the signal already has a meaning in the process — which is also the argument for picking a signal that does not, such as `signal.SIGUSR1`.

It is a fire alarm wired straight to the siren rather than through the receptionist: pulling it still works when everyone inside the building is stuck in a meeting.

saying these in an interview costs you the question

  • Thinks the dump handler is Python code a blocked interpreter can still run
  • Claims faulthandler.register works on Windows
  • Expects the process to exit or unwind after dumping
  • Registers on SIGINT or SIGTERM and breaks shutdown
  • Assumes C-extension frames appear alongside Python frames
  • Believes only the signalled thread is dumped

context