Why do a document-conversion worker's atexit handlers never run when its container is sent SIGTERM?
answer
- Local runs clean, production runs do not
- How the process ends decides everything
- Only one signal has a Python default
- Turn the signal into normal unwinding
- Then assume it still will not run
basics
~20 sA default SIGTERM terminates the process at the OS level, so the interpreter never reaches finalization and no registered handler runs. Install a Python SIGTERM handler that raises SystemExit, and treat the shutdown flush as best effort regardless.
solid answer
~40 sBy default Python installs a handler only for SIGINT, which raises `KeyboardInterrupt` and unwinds normally; SIGTERM keeps its OS default of immediate termination, so nothing registered with `atexit.register` is ever consulted — the worker's page-count cache is simply never flushed, and the stale value it wrote weeks ago survives another release. The fix is to call `signal.signal(signal.SIGTERM, handler)` early, with a handler that calls `sys.exit` so the interpreter unwinds and finalization runs. Two limits stay: the handler only runs between bytecode instructions in the main thread, so a long blocking C call delays it, and the runtime usually follows SIGTERM with SIGKILL after a grace period, which nothing can intercept. So make the shutdown work fast and idempotent, and write anything that must survive as it is produced rather than at exit.
code
python · 13 linesimport atexit
import signal
import sys
import time
atexit.register(print, "cache flushed")
def on_term(signum, frame):
sys.exit(0)
signal.signal(signal.SIGTERM, on_term)
print("worker up", flush=True)
time.sleep(30)go deeper
Recall that how a process ends decides whether any Python cleanup runs at all, and that a process killed by a signal usually skips it. Reproduce the difference by ending a script two ways.
Explain the mechanics: only SIGINT has a Python default handler, a SIGTERM handler calling sys.exit restores normal unwinding, and Python-level handlers run between bytecodes in the main thread.
Show the diagnosis and the limits: verify what the platform actually sends, keep the main thread responsive, fit cleanup inside the grace period, and stop letting correctness ride on an exit-time write.
Own the shutdown contract for the fleet: what the grace period is, what each service promises to finish within it, what state must be recoverable after an uncatchable kill, and how that is verified rather than assumed.
### The observed symptom A document-conversion queue keeps a per-document page-count cache in memory and flushes it with a handler registered through `atexit.register`. Restarts during deploys leave the cache unflushed, so a stale value written before a three-week release train is still being served afterwards. Locally the flush always works; in the container it never does. That asymmetry is the whole clue: locally the program ends by running out of work, and in production it ends by being signalled. ### Why the handler never runs `atexit` handlers are invoked during *normal* interpreter finalization. A signal whose disposition is the OS default does not produce normal finalization — the kernel terminates the process, and no Python code runs at all. Python installs a handler only for SIGINT, which raises `KeyboardInterrupt`; that unwinds the stack the ordinary way and therefore does reach finalization. Nothing installs one for SIGTERM, so a container stop, a supervisor restart or a plain `kill` skips every handler, every finalizer and every buffered write. Reproducing it takes one terminal: start the worker, send it a TERM, and watch the flush message never appear. ### Confirming the diagnosis * Send SIGTERM by hand and compare the output with a run that ends by itself. If the difference is exactly the cleanup output, you are done. * Check whether anything in the process installs a SIGTERM handler — a framework, a supervisor shim, or an imported library may already have. The handler that is actually in force is the one that matters, not the one you meant to install. * Confirm what the platform sends. Most container runtimes send SIGTERM, wait a grace period, then send SIGKILL. Any diagnosis has to know the length of that window, because it bounds every fix. ### The fix Install a handler that turns the signal into normal unwinding: ```python import atexit, signal, sys, time atexit.register(flush_cache) def on_term(signum, frame): sys.exit(0) signal.signal(signal.SIGTERM, on_term) ``` `sys.exit` raises `SystemExit`, which propagates out of the main thread and lets the interpreter finalize normally, so the handlers run. Two constraints govern how well that works. **Signal handlers are deferred.** The OS-level handler CPython installs only sets a flag; the Python function runs later, between bytecode instructions, and only in the main thread. If the main thread is blocked inside a long C call, the handler waits for it to return. A worker whose main thread sits in a multi-minute conversion call may be killed before the handler is ever reached. The structural fix is to keep the main thread in a short-interval wait and do the long work elsewhere, so a signal is noticed promptly. **Unwinding is not instant.** `SystemExit` propagates through whatever the main thread is doing, and cleanup must fit inside the grace period. A flush that acquires a lock, opens a connection or retries with backoff is exactly the kind of shutdown work that gets cut off midway. ### The deeper answer an interviewer wants Even a perfect handler does not make an exit-time flush trustworthy. SIGKILL cannot be intercepted; a machine can lose power; a fatal error in an extension module can end the process without a flush; and `os._exit` in some library's error path does the same. Every one of those produces the same stale cache. So the durable fix is not only the signal handler. Write the value when it changes, or write it periodically and treat the exit-time flush as an optimisation; make the write idempotent so a repeat after a partial shutdown is harmless; and give the cache a validity marker so a reader can tell a stale value from a fresh one instead of trusting it forever. A shutdown path is allowed to *improve* the outcome; a correctness argument that depends on it is a bug waiting for the next hard kill. ### Checklist worth stating aloud 1. Install SIGTERM (and, if relevant, SIGHUP) handlers that request an orderly exit. 2. Keep the main thread responsive so those handlers are actually reached. 3. Keep shutdown work short, idempotent, and well inside the grace period. 4. Never let correctness depend on it — assume the process can vanish between any two instructions. 5. Log at the start and end of shutdown, so the next incident tells you which of those steps did not happen.
- The SIGTERM handler is installed but still does not run before the process dies. What now?Python signal handlers run in the main thread between bytecode instructions, so a main thread blocked in a long C call cannot reach one until that call returns. Keep the main thread in a short-interval wait and move long conversions off it, so the flag set by the OS-level handler is acted on within milliseconds. If the work itself must run to completion, the grace period has to be long enough for it, or the work has to be resumable after a kill.
- Why is SIGINT different from SIGTERM in a plain Python program?Python installs a default handler for SIGINT that raises `KeyboardInterrupt` in the main thread. That is an ordinary exception, so it unwinds the stack, runs `finally` blocks, and reaches interpreter finalization — which is why Ctrl+C does run registered handlers. SIGTERM keeps the operating system's default disposition of immediate termination, so nothing Python-level happens unless you install a handler yourself.
- How would you make the cache correct even if the process is killed with SIGKILL?Stop depending on the exit path. Persist the value when it changes, or on a short interval, so the worst case is losing a bounded window rather than the whole run. Make the write idempotent so a repeat after a partial shutdown is harmless, and stamp the record so a reader can distinguish a fresh value from a stale one instead of trusting whatever it finds. The exit-time flush then becomes an optimisation, not a correctness requirement.
saying these in an interview costs you the question
- Assumes every process termination runs cleanup
- Claims Python handles SIGTERM by default
- Adds a slow, blocking flush to the shutdown path
- Believes a handler can intercept SIGKILL
- Ignores that handlers run only in the main thread
- Keeps correctness depending on exit-time writes