skip to content

How would you wire faulthandler into a service whose workers die and hang silently?

level: seniorimportance: should knowfreq 24%

answer

  1. Workers vanish and nothing is logged
  2. Turn it on without touching the code
  3. Prove at boot that it is armed
  4. The dump needs a stream someone keeps
  5. PYTHONFAULTHANDLER plus a timeout watchdog

basics

~20 s

Set PYTHONFAULTHANDLER in the process environment so crashes during startup are covered too, make sure the destination stream is captured and retained, assert faulthandler.is_enabled() at boot, and arm a dump_traceback_later watchdog around the work that hangs.

solid answer

~50 s

Turn it on by environment, not by code: `PYTHONFAULTHANDLER=1` in the unit or container spec enables the handler during interpreter startup, so a crash inside an import or inside a compiled extension's initialisation is covered too, and nothing can refactor the call away. Assert `faulthandler.is_enabled()` once at boot so a dropped environment variable fails loudly rather than silently costing you the next crash. Decide where the dump lands: the default is `sys.stderr`, which is right when the supervisor captures and retains stderr; if you pass `file=` instead, keep that object open for the life of the process, because the module holds the file descriptor and log rotation will leave it writing into a deleted file. For the hang half of the problem, arm `faulthandler.dump_traceback_later()` around the unit of work and cancel it in a `finally`. The runtime cost until something goes wrong is effectively zero.

code

python · 8 lines
python
import faulthandler

crash_log = open("crash.log", "a", buffering=1)
faulthandler.enable(file=crash_log, all_threads=True)

# crash_log must stay open for the life of the process:
# faulthandler holds its file descriptor, not the object.
print("enabled:", faulthandler.is_enabled())

go deeper

for a junior

Recall the two switches: PYTHONFAULTHANDLER=1 in the environment for crashes, and a faulthandler.dump_traceback_later() watchdog armed and cancelled around work that can hang.

for a middle

Explain why the environment variable covers more than an enable() call, and why a file= destination has to stay open for the process's life because the module holds its file descriptor.

for a senior

Demonstrate the operational reasoning: verify with is_enabled() at boot, confirm the stream is captured and retained, pick a watchdog timeout above the healthy tail, and know that a missing dump points at a platform kill.

for a principal

Own it as a fleet policy - crash and hang evidence on by default, a defined destination and retention for the plain-text channel, and a clear rule on whether stuck workers are hard-killed or left for a human.

## The two silences A warehouse pick-list builder that occasionally disappears is showing you two different failures wearing the same face. - One is a **fatal signal**: an extension dereferences a bad pointer, the process dies with status 139, and the restart looks like a blip. - The other is a **hang**: the worker takes the job, never returns it, and eventually something upstream gives up - an intermittent timeout where the 92nd-percentile build sits far inside its budget and one build in a few thousand never finishes at all. Neither produces a traceback by default, and the fix is to make each of them produce one before you go looking for a cause. ## Enable it by environment `PYTHONFAULTHANDLER=1` belongs in the environment the service starts with, alongside its other runtime variables. The reason to prefer it over a `faulthandler.enable()` call is **coverage**: the environment variable is honoured while the interpreter starts, so a crash during import - which is where extension-module initialisation happens, and a common place for the ugly ones - is still reported. A call in your own startup code cannot cover anything that runs before it. `-X faulthandler` on the command line is equivalent when you control the command line rather than the environment. ## Then verify it Environment variables are lost: - by deployment refactors, - by a base image change, - or by someone copying a run command. One line at boot - assert `faulthandler.is_enabled()`, or log its value with the rest of your startup banner - converts that into a visible fact instead of a surprise you discover only when you needed the dump and did not get it. ## Make sure the output survives The dump goes to a **raw file descriptor**, by default the one behind `sys.stderr`. That is the right default when the supervisor captures stderr and ships it somewhere with enough retention to still hold the moments before a restart; it is the wrong default if stderr goes to a terminal nobody is attached to, or if the log pipeline only ingests the structured stream. If you redirect with `enable(file=...)`, remember that faulthandler keeps the underlying descriptor, not the Python object: the file must stay open for the whole life of the process, and a rotation scheme that closes and reopens the log will leave the handler writing into a file nobody can find. A dedicated append-mode file opened once at startup and never rotated by the application is the boring, correct choice. And because the handler cannot go through the logging system - it must be safe to run in a signal context, so no locks and no allocation - the dump: - will not be JSON, - will not carry your request identifier, - and will interleave with anything else writing to that stream. Plan for a plain-text crash channel rather than trying to make it look like an application log. ## Cover the hang with a watchdog The crash half is handled once the handler is installed; the hang half needs `faulthandler.dump_traceback_later()`. Arm it around the unit of work with a timeout comfortably above the healthy tail - if the slow end of a build sits well inside the budget, set the watchdog above that tail so a dump always means something pathological - and cancel it in a `finally` so it only fires on an overrun. - **`repeat=True`** is worth having when the goal is to distinguish stuck from slow: two identical dumps a few seconds apart say the frames are not moving. - **`exit=True`** turns the watchdog into a hard deadline that kills the worker after dumping; that is a real option when a wedged worker holds a job that another worker could retry, and a bad one when the process holds unflushed state, because the kill skips `finally` blocks and `atexit` handlers. ## Know the boundary faulthandler is **evidence, not protection**. - It does not keep a crashing process alive, - it does not interrupt a stuck call unless you asked it to kill the process, - and it shows Python frames only - a hang inside compiled code appears as the Python call site that entered it. It also cannot report a death that delivers no signal it can handle, such as a hard kill from the platform, so a silent disappearance with no dump and no exit signal points at the environment rather than at the interpreter. Finally, the dump is a stack, not a diagnosis: it tells you which call site was live, and the work of deciding why is still yours. ## The cost Installed handlers cost nothing per bytecode; an armed watchdog costs one sleeping thread. Compared with losing every crash and every hang, that is not a tradeoff worth deliberating - the defensible default for a long-running service is on.

  • Why is passing enable(file=...) to a rotated log file a trap?
    faulthandler keeps the underlying file descriptor, not the Python object. When a rotation scheme closes and renames the file, the handler carries on writing into the old, now-unlinked inode, so the dump you need is written somewhere nobody looks. Use `sys.stderr` where the supervisor captures it, or a dedicated append-mode file the application opens once and never rotates itself.
  • The workers disappear with no dump at all, even with the handler enabled. What does that tell you?
    That no signal faulthandler handles was delivered. A hard kill from the platform - an out-of-memory kill or an orchestrator stop - gives the process no chance to run any handler, so the absence of a dump is itself evidence: look at the platform's own kill records and resource limits rather than at Python. The same holds for a power-cut style termination of the container.
  • Would you turn faulthandler on for every service by default, or only where crashes are known?
    By default. The runtime cost is a handful of installed signal handlers and nothing per bytecode, and the value only appears on the first unexpected crash, which by definition you cannot predict. The condition is that the output stream is actually captured and retained; enabling it while stderr goes nowhere buys you a false sense of coverage.
  • Your team wants the dump in the JSON log pipeline. What do you tell them?
    It cannot go there. The handler runs in a signal context where allocating memory or taking a lock is unsafe, so it writes plain text straight to a file descriptor and cannot pass through the logging framework. The workable answer is to capture that stream separately and correlate by timestamp and process identifier, treating the dump as a crash artefact rather than a log record.

saying these in an interview costs you the question

  • Enabling it in code and calling startup crashes covered
  • Assuming stderr is captured without checking the supervisor
  • Passing a rotated log file to enable() and losing the dump
  • Expecting the dump to flow through the logging framework
  • Treating a hang as covered by the crash handler alone
  • Calling faulthandler a safety net that keeps workers alive

context