Why does atexit.register cleanup not run when a Python process is killed by SIGTERM?
answer
- Hooks belong to the interpreter, not the OS
- Only a normal shutdown reaches them
- os._exit skips the same machinery
- A handler turns the kill into an exception
- SystemExit propagating is the graceful path
basics
~20 sFunctions registered with atexit.register run only during a normal interpreter shutdown. The default disposition for signal.SIGTERM destroys the process in the kernel, so no further bytecode executes and no cleanup path is reached. Install a terminate handler to get one.
solid answer
~40 s`atexit` is not an operating-system facility — it is a list of callables the interpreter walks on its way out during a **normal** shutdown: the module falling off its last line, or a `SystemExit` propagating out of the main thread. A default `signal.SIGTERM` is not a shutdown, it is a kill: the kernel destroys the process image, so nothing runs afterwards — not `atexit` hooks, not `finally`, not `__exit__`, not `__del__`, and buffered writes are lost. The same is true of `os._exit`, which exists precisely to skip that machinery. To make the hooks fire, register a handler for `signal.SIGTERM` that raises `SystemExit` (or sets a stop flag the main loop reads); the normal shutdown path then runs and `atexit` behaves as expected.
code
python · 13 linesimport subprocess
import sys
import time
child = subprocess.Popen(
[sys.executable, "-c",
"import atexit, time\n"
"atexit.register(lambda: print('cleanup ran', flush=True))\n"
"time.sleep(30)\n"],
)
time.sleep(0.5)
child.terminate()
print("child exit status:", child.wait())go deeper
Remember that atexit is an interpreter feature, not an operating-system one, and that a process killed by a signal never gets to run it. Know that sys.exit does reach it.
Explain the finalization path precisely: normal end of the main module or SystemExit propagating out of the main thread, and list what else bypasses it — os._exit, fatal signals, daemon threads.
Show the production shape: a handler that only sets a flag, a main loop that stops at a safe point, and correctness-critical cleanup that does not depend on hooks firing at all.
Take the position that shutdown cleanup should never be load bearing. Argue for retryable, idempotent work units so that an abandoned process costs a redo rather than data loss.
## What atexit actually is `atexit.register(func)` appends `func` to a list held by the interpreter. Nothing about that list is known to the kernel, the shell or a supervisor. During **interpreter finalization** CPython walks the list in last-in-first-out order and calls each entry. Interpreter finalization is reached from exactly two ordinary places: * the main module runs off its last statement, or * a `SystemExit` propagates out of the main thread — which is all `sys.exit()` does, it raises `SystemExit`. Anything that ends the process without going through finalization skips the list entirely. The default disposition of `signal.SIGTERM` is one of those things. ## Why the default disposition skips everything When a signal's disposition is `signal.SIG_DFL` and the signal's default action is "terminate", the kernel destroys the process on delivery. Your program is not asked, notified, or given a chance to run one more bytecode. So all of this is skipped at once: * `atexit` callbacks * `finally` blocks anywhere on the stack * `__exit__` on every open context manager * `__del__` finalizers * buffered writes sitting in a file object or a logging handler that have not reached the OS That last point is the one that bites hardest in practice. A process that appends to a file and is terminated at the default disposition loses whatever was still in the userspace buffer, which is why a crashed writer's output so often stops mid-record. ## The demonstration ```python import subprocess import sys import time child = subprocess.Popen( [sys.executable, "-c", "import atexit, time\n" "atexit.register(lambda: print('cleanup ran', flush=True))\n" "time.sleep(30)\n"], ) time.sleep(0.5) child.terminate() print("child exit status:", child.wait()) ``` `Popen.terminate()` sends `signal.SIGTERM`. The child prints nothing, and the parent prints `child exit status: -15` — a negative value, meaning "killed by signal 15", not an exit code. ## The other ways to skip atexit Understanding the exception list is what turns this from trivia into judgement: * **`os._exit(status)`** ends the process immediately at the C level by design. It is the correct call in a forked child that must not run the parent's cleanup twice, and the wrong call anywhere else. * **A fatal signal** — `signal.SIGKILL`, `signal.SIGSEGV`, `signal.SIGABRT`. * **A hard crash in a native extension**, which is the same thing arriving from a different direction. * **A daemon thread** still running when the interpreter shuts down: it is not given a chance to finish, so its own `finally` blocks may never run even though the process exits normally. Conversely, `SystemExit` raised anywhere in the main thread and left uncaught *does* reach finalization, which is what makes the handler shape below work. ## Making SIGTERM reach the cleanup path The minimal correct fix is a handler that converts the signal into an exception: ```python import signal def _on_terminate(signum, frame): raise SystemExit(128 + signum) signal.signal(signal.SIGTERM, _on_terminate) ``` Now the terminate signal unwinds the stack: `finally` blocks run, context managers exit, `atexit` hooks fire, buffers flush. For a long-running worker, raising from the handler is often too blunt — the exception lands at an arbitrary bytecode boundary, possibly halfway through a network write, and the unwinding you get is correct but not *safe*. The production shape is a handler that only records intent: ```python import signal import threading stopping = threading.Event() signal.signal(signal.SIGTERM, lambda signum, frame: stopping.set()) ``` and a main loop that checks `stopping.is_set()` between work items and returns normally. Returning normally is what gets you to finalization, so the `atexit` hooks still run — you have simply chosen *where* the shutdown begins. ## What belongs in an atexit hook at all Even once the hooks run, they are the weakest form of cleanup available and should not be load bearing. They only fire on a graceful exit, they run at an unpredictable point in finalization when other modules may already be torn down, an exception raised inside one is printed and the remaining hooks still run, and none of them will save you from `signal.SIGKILL`. Use them for best-effort niceties — flushing a metrics buffer, removing a scratch directory — and put anything whose loss would be a correctness problem behind an explicit `try`/`finally`, a context manager, or a design where an abandoned unit of work is simply retried by whoever hands it out. ## Handler discipline A signal handler runs between bytecodes in the main thread while the rest of the program is mid-flight. Keep it to setting a flag or raising; do not log from it, acquire locks in it, or do I/O in it. A second terminate signal arriving while the first is still being processed is a real scenario, so make the handler idempotent — setting an already-set `threading.Event` is safe by construction.
- Which other exits skip atexit callbacks besides a default-disposition terminate signal?`os._exit()`, which ends the process at the C level by design and is the right call only in a forked child that must not repeat the parent's cleanup; any fatal signal such as `signal.SIGKILL` or `signal.SIGSEGV`; and a hard crash inside a native extension. A daemon thread is a subtler case: the process exits normally and the hooks run, but the thread is abandoned where it stands, so its own `finally` blocks may never execute.
- If the atexit hooks do run, is that enough to call the shutdown graceful?No. They are best-effort: they fire late in finalization when other modules may already be torn down, they never fire under `signal.SIGKILL`, and an exception inside one is printed and swallowed. Anything whose loss is a correctness problem belongs in an explicit `try`/`finally`, a context manager, or a design where an abandoned work item is simply retried by whoever dispatched it.
- Why is raising SystemExit from the handler a poor fit for a long-running worker?Because the raise happens at whatever bytecode boundary the main thread has reached — possibly mid-request or between a send and the bookkeeping that records it. The unwinding is correct but the timing is arbitrary. A worker usually wants the handler to set a `threading.Event` and the main loop to notice it between work items, so shutdown starts at a point the program chose.
saying these in an interview costs you the question
- Treating atexit as an OS-level guarantee
- Believing a SIGTERM handler is registered by default
- Thinking finally blocks survive a default terminate signal
- Using os._exit in normal code to exit faster
- Putting correctness-critical cleanup only in atexit
- Doing logging or I/O inside the signal handler