skip to content

Why is doing real work inside a signal.signal() handler unsafe, and what belongs there instead?

level: seniorimportance: should knowfreq 40%

answer

  1. Not a thread, a re-entrant call
  2. Cuts between two bytecodes
  3. It can hold nothing it might need
  4. A held lock deadlocks against itself
  5. Record intent, act at a safe point

basics

~20 s

The handler runs on the main thread between two bytecode instructions, so it re-enters code that was mid-update and can tear a multi-step change or deadlock on a lock that code already holds. Set a flag and return.

solid answer

~50 s

A Python signal handler is not a separate thread and it is not atomic. It runs **on the main thread, between two bytecode instructions**, on top of whatever frame was executing — so it can observe half-finished state and mutate the very structures the interrupted code was updating. In a long billing run that means a total already incremented while the matching item has not yet been marked done, and a handler that touches those totals produces a race with the code it interrupted even though only one thread exists. Worse, if the handler acquires a `threading.Lock` the main thread already holds, the process deadlocks against itself; and an exception raised in the handler propagates out of whatever line was executing, aborting a sequence anywhere. The rule: the handler sets a module-level flag or writes one byte, and returns. The main loop checks it at a point where invariants hold and does the real work there.

code

python · 15 lines
python
import signal

def on_usr1(signum, frame):
    raise RuntimeError("interrupted mid-update")

signal.signal(signal.SIGUSR1, on_usr1)

charged = 0
pending = 10
try:
    charged += pending
    signal.raise_signal(signal.SIGUSR1)
    pending = 0            # never runs: the two-step update is torn
except RuntimeError as exc:
    print(exc, "-> charged:", charged, "pending still:", pending)

go deeper

for a junior

Remember the rule before the reasoning: a signal handler should set a flag and return, and the real work happens in the main loop when it notices that flag.

for a middle

Explain why the rule exists: the handler runs on the main thread between two bytecode instructions, so it can see half-finished state and can block on a lock the interrupted code already holds.

for a senior

Show you have debugged this: name torn multi-step updates and self-deadlock as the two concrete failures, and describe driving the interruption point deliberately in tests with signal.raise_signal rather than hoping to hit it.

for a principal

Own the pattern across services: one agreed shutdown protocol where handlers only record intent, invariants are checked at defined safe points, and long runs are structured into resumable units so an interruption never leaves partial financial state.

## The handler is a re-entrant call, not a thread CPython runs your Python-level handler on the main thread of the main interpreter, at a bytecode boundary, on top of the frame that was executing. There is no separate stack of execution and no isolation. Everything difficult about signal handlers follows from that one fact, and interviewers at senior level are checking whether you have felt it in production rather than read it. ## Failure one: a race with yourself Concurrency bugs are normally explained as two threads interleaving. A signal handler gives you the same class of bug with a single thread, because "between bytecodes" cuts finer than a line of Python. Consider a subscription-billing run that, per account, adds an amount to a running total and then marks the item settled. Those are separate bytecodes. A signal can land between them, and a handler that reads or rewrites the totals sees an amount counted with nothing settled — and if the handler *writes* the shared structure, the interrupted code resumes and overwrites it. A regression pack of 340 accounts that raises the signal at a different point in each case is exactly how a team finds this: the failure is not reproducible at a fixed line, because the interruption point is not fixed. Even a single augmented assignment is unsafe: `total += amount` loads, adds and stores as separate instructions, so a handler that also touches `total` can lose an update. ## Failure two: deadlocking against yourself `threading.Lock` is not re-entrant and, more importantly, has no notion of "the same thread already has it". If the main thread holds a lock and the handler tries to acquire it, the acquire blocks forever waiting for a thread that is stopped inside the handler. This is the single most common way a shutdown handler hangs a process. Any indirect lock counts too: writing to a logging handler, printing, or calling into a library that guards its own state can all take a lock that the interrupted code was already holding. Using `threading.RLock` removes the hang but not the bug — the handler now sails into a critical section whose invariants are half-established, which is a quieter version of the same corruption. ## Failure three: an exception from anywhere If the handler raises, the exception propagates from whatever line the main thread was executing. That is how `KeyboardInterrupt` works, and it means an unremarkable statement can suddenly raise, skipping the rest of a sequence. Code that is not exception-safe at every line becomes wrong in the presence of a raising handler. That is a legitimate design — it is how you stop a runaway loop — but it must be a decision, not an accident, and cleanup must live in `finally` blocks or context managers rather than in the lines that were about to run. ## What the handler should do The standard shape is: record intent, return immediately. ```python stop_requested = False def on_signal(signum, frame): global stop_requested stop_requested = True ``` A single assignment to a module-level name is one bytecode store and is safe. The main loop then tests the flag between units of work — at a point where the totals and the settled markers agree — and performs the shutdown, the flush, or the report there, with the whole language available to it and no re-entrancy risk. The variant for a loop that blocks on file descriptors is to have the handler write one byte to a pipe (or let CPython do it through the wakeup file descriptor) so the blocked wait returns, and then handle the request on the normal path. ## Testing it honestly `signal.raise_signal()` triggers the handler synchronously from the main thread, which lets a test drive the interruption at a chosen point instead of hoping to hit it. Drive it at many different points across your regression pack, assert the invariant after each, and the tearing bugs surface deterministically. Also assert the boring properties: the handler returns quickly, it never blocks, and it holds no lock when it returns. ## What to say in the interview One sentence of mechanism (main thread, between bytecodes, on top of the interrupted frame), two named failures (torn multi-step state, self-deadlock on a held lock), and the rule (set a flag, act at a safe point). That is a complete senior answer, and everything else is elaboration.

  • Is setting a module-level boolean from a handler actually safe, given the handler runs between bytecodes?
    Yes. A plain assignment to a module-level name compiles to a single store instruction, so it cannot be interrupted part-way and cannot be torn. What is not safe is a read-modify-write such as counter += 1, which loads, adds and stores separately, or any mutation of a container the interrupted code was midway through updating. Keep the handler to one store and return.
  • How would you reproduce a tearing bug like this deterministically in tests?
    Use signal.raise_signal() to fire the signal from the main thread at a chosen point rather than waiting for a real one. Parameterise a regression pack over interruption points — after the total is updated, before the item is marked settled, and so on — and assert the invariant after each. Because raise_signal runs the pending handler before it returns, the timing is deterministic instead of racy.
  • Would using threading.RLock instead of threading.Lock fix a handler that deadlocks?
    It removes the hang but not the defect. An RLock already owned by the main thread lets the handler straight into a critical section whose invariants are only half-established, so instead of a visible deadlock you get silent corruption. The fix is not to take locks in the handler at all: record intent and let the main loop enter the critical section normally.

It is a phone call that interrupts you mid-sentence at your own desk: whatever you were half-writing is still on the page, and if you start editing it during the call you finish with two overlapping versions.

saying these in an interview costs you the question

  • Assumes the handler runs in its own thread, so no races
  • Puts blocking cleanup or network I/O inside the handler
  • Acquires a lock the interrupted code may already hold
  • Thinks a line of Python is atomic against a signal
  • Believes an exception raised in the handler stays inside it
  • Swaps Lock for RLock and calls the corruption fixed

context