skip to content

SIGTERM and Graceful Shutdown

The shutdown contract a supervisor expects: catch the terminate signal, stop taking work, drain what is in flight, and exit before the grace period runs out and an uncatchable kill arrives.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

How do signal.SIGINT and signal.SIGTERM differ in a Python process by default?

level: juniorimportance: must knowfreq 60%

answer

  1. One of the two reaches Python code
  2. The interpreter installs one handler at startup
  3. Ctrl-C and a supervisor stop differ
  4. default_int_handler raises an exception
  5. SIGTERM stays at the OS default disposition

basics

~20 s

CPython installs its own default handler for signal.SIGINT that raises KeyboardInterrupt in the main thread, so finally blocks and atexit hooks still run. signal.SIGTERM keeps the operating system default, which terminates the process immediately with no Python cleanup.

solid answer

~40 s

Both are polite termination requests, but only one of them reaches Python code. At startup CPython sets a Python-level handler for `signal.SIGINT` (`signal.default_int_handler`), which raises `KeyboardInterrupt` in the main thread between bytecodes; that exception unwinds the stack, so `finally` blocks, `with` cleanup and `atexit` hooks all run before the process ends. `signal.SIGTERM` is left at the operating system default disposition, so the kernel tears the process down at once — no bytecode executes, no buffers are flushed, and the shell reports the process as killed by signal 15. That is why a service that only ever tested Ctrl-C looks like it shuts down cleanly and then loses in-flight work under a supervisor, which sends `signal.SIGTERM`. The fix is to register a handler for `signal.SIGTERM` yourself.

code

python · 4 lines
python
import signal

print(signal.getsignal(signal.SIGINT) is signal.default_int_handler)
print(signal.getsignal(signal.SIGTERM) is signal.SIG_DFL)

go deeper

for a junior

Be ready to say plainly that Ctrl-C raises KeyboardInterrupt in Python while a plain SIGTERM kills the process outright, and that only the first runs your finally blocks.

for a middle

Explain the mechanism: the interpreter installs default_int_handler for SIGINT at startup, leaves SIGTERM at the OS default, and runs Python handlers at bytecode boundaries in the main thread.

for a senior

Show that you test shutdown with the signal production actually sends. Demonstrating that a service loses in-flight work under a supervisor while looking clean under Ctrl-C is the point an interviewer is listening for.

for a principal

Own it as a fleet-wide default: every long-running service template registers a terminate handler, and exit codes 143 versus 137 are tracked so you can see how often the polite request is being ignored.

## Dispositions, not handlers Every Unix process carries a **disposition** for each signal number: what the kernel should do when that signal is delivered. The three possibilities are *default* (`signal.SIG_DFL` — the action baked into the kernel for that signal, which for the terminate family is "destroy the process"), *ignore* (`signal.SIG_IGN`), or *handle* (run a function). A Python process starts life with the dispositions it inherited from its parent, with one interesting exception that CPython itself adds. ## What CPython adds at startup CPython installs exactly one Python-level signal handler of its own during interpreter startup: a handler for `signal.SIGINT`, exposed as `signal.default_int_handler`. All that function does is `raise KeyboardInterrupt`. You can see the asymmetry directly: ```python import signal signal.getsignal(signal.SIGINT) is signal.default_int_handler # True signal.getsignal(signal.SIGTERM) is signal.SIG_DFL # True ``` `signal.SIGTERM` was never touched, so it still carries the kernel default: terminate the process. ## Why the difference matters so much A Python signal handler does not run inside the kernel's delivery context. The C-level handler only records that the signal arrived and sets a flag; the eval loop notices the flag at the next bytecode boundary in the **main thread** and calls your Python function there. For `signal.SIGINT`, that function raises `KeyboardInterrupt`, and from that instant on it is an ordinary exception travelling up an ordinary stack. It runs `finally` blocks. It runs `__exit__` on every context manager it passes. If nothing catches it, the interpreter shuts down normally, which flushes buffered writes and runs everything registered with `atexit.register`. `signal.SIGTERM` at its default disposition never reaches the eval loop at all. The kernel destroys the process image. Nothing in your program observes it happening, which is exactly why the failure is so easy to miss in development, where you stop things with Ctrl-C. ## Ctrl-C is not SIGTERM Pressing Ctrl-C does not send a signal from Python or from the shell — the terminal line discipline sends `signal.SIGINT` to the entire foreground process group. A supervisor stopping a service, by contrast, sends `signal.SIGTERM` to the process (often followed later by `signal.SIGKILL`). They mean roughly the same thing socially — "please stop" — and behave completely differently technically. Testing shutdown with Ctrl-C therefore proves almost nothing about how the program behaves in production. ## KeyboardInterrupt is a BaseException `KeyboardInterrupt` deliberately does **not** inherit from `Exception`; it inherits from `BaseException`, alongside `SystemExit` and `GeneratorExit`. That means a broad `except Exception:` around a worker loop will *not* swallow a Ctrl-C, which is the intended design: an interrupt should not be mistaken for an application error and retried. It also means that code which writes `except BaseException:` or a bare `except:` will swallow interrupts and make the program feel unkillable. Because the exception is raised at an arbitrary bytecode boundary, it can land anywhere — including between acquiring a resource and entering the `try` that would release it. This is the reason to keep the window narrow by acquiring resources with `with` rather than by hand. ## Exit status A process killed by a signal does not have a normal exit code. On POSIX, `subprocess.Popen.wait()` returns the negated signal number (`-15` for `signal.SIGTERM`, `-9` for `signal.SIGKILL`), and shells report `128 + signum` (143 and 137). By convention a program that handles a terminate signal and exits deliberately can mirror that with `SystemExit(128 + signum)`, which keeps dashboards that count 143s honest. ## Making SIGTERM behave The whole fix is one registration, done once, in the main thread: ```python import signal def _on_terminate(signum, frame): raise SystemExit(128 + signum) signal.signal(signal.SIGTERM, _on_terminate) ``` Now `signal.SIGTERM` unwinds the stack exactly like Ctrl-C does: `finally` blocks run, `atexit` hooks run, buffers flush. Raising `SystemExit` from the handler is the simplest correct shape for a short-lived script. A long-running worker usually prefers a handler that sets a flag — an interrupt that lands in the middle of a network write is not a clean place to unwind from — and lets the main loop notice the flag and shut down at a safe point. ## The one signal you cannot do this for `signal.SIGKILL` (and `signal.SIGSTOP`) cannot be caught, handled or ignored. Calling `signal.signal` on them raises `OSError`. That is the whole point of the supervisor's two-stage stop: a catchable request first, an uncatchable one when the grace period expires.

  • Why does KeyboardInterrupt inherit from BaseException rather than Exception?
    So that broad application error handling does not swallow an interrupt. A worker loop that wraps its body in `except Exception:` to log and retry would otherwise ignore Ctrl-C and appear unkillable. `SystemExit` and `GeneratorExit` sit beside it under `BaseException` for the same reason: they are control-flow events, not application failures, and code that genuinely wants them must ask for them explicitly.
  • What exit status does a shell report for a process terminated by SIGTERM?
    143, which is `128 + 15`. A process killed by `signal.SIGKILL` reports 137. From Python, `subprocess.Popen.wait()` returns the negated signal number instead — `-15` and `-9` — while a normal exit returns the real exit code. Services that exit deliberately on a terminate signal often raise `SystemExit(143)` so the two paths look the same to whatever counts exit codes.
  • Does registering a SIGTERM handler help if the signal arrives while a worker thread is busy?
    The handler still runs, but only in the main thread, and only at a bytecode boundary there. If the main thread is blocked in a long C call the handler is delayed until that call returns. Worker threads never run signal handlers at all, so the usual shape is a handler that sets a `threading.Event` and worker threads that check it between items.

SIGINT is a knock on the door that you can answer and tidy up before opening; SIGTERM, unanswered, is the building being demolished around you mid-sentence.

saying these in an interview costs you the question

  • Claiming SIGTERM raises KeyboardInterrupt like Ctrl-C
  • Believing atexit hooks run on a default SIGTERM
  • Thinking Ctrl-C sends SIGTERM to the process
  • Saying signal handlers run inside the kernel context
  • Catching interrupts with a bare except and calling it robust
  • Assuming handlers can be registered from any thread

context

open as a page

Why does atexit.register cleanup not run when a Python process is killed by SIGTERM?

level: middleimportance: must knowfreq 50%

basics

~20 s

Functions registered with atexit.register run only during a normal interpreter shutdown. The default disposition for signal.SIGTERM destroys the process in the kernel, so no further bytecode executes and no cleanup path is reached. Install a terminate handler to get one.

open as a page

An email-digest sender gets SIGTERM with a 30-second window before SIGKILL — how do you drain without truncating a digest?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Catch signal.SIGTERM with a handler that only sets a flag, stop pulling new digests, finish the ones in flight against a deadline shorter than the 30-second window, and exit. signal.SIGKILL cannot be caught, so anything unfinished must be safely retryable.

open as a page

How would you set the SIGTERM drain budget for a fleet of Python workers, and decide what to abandon?

level: principalimportance: should knowfreq 30%

basics

~20 s

Derive the budget from measured work-item durations — a high percentile plus margin, not the mean or the maximum — keep it inside the supervisor's window, and make abandoning work cheap by keeping every unit idempotent.

open as a page