skip to content

Why can KeyboardInterrupt interrupt a shared-cache update midway and leave it inconsistent?

level: seniorimportance: should knowfreq 32%

answer

  1. The signal handler does almost nothing
  2. A flag is checked between instructions
  3. Delivery is asynchronous to your source lines
  4. The rollback handler named the wrong base class
  5. Only the main thread runs signal handlers

basics

~20 s

Ctrl-C sets a flag that the interpreter checks between bytecodes, so KeyboardInterrupt is raised at an arbitrary point inside a multi-step update. Because it derives from BaseException, error handlers written for Exception never run the rollback.

solid answer

~50 s

The signal itself does not raise anything. The C-level handler sets a flag, and the interpreter's evaluation loop checks that flag between bytecode instructions, raising `KeyboardInterrupt` in the main thread at whichever instruction boundary comes next. A multi-step mutation — write the entry, then bump the counter — can therefore be torn in half, leaving shared state that no single line of your code could produce. Two things make it worse: `KeyboardInterrupt` derives from `BaseException`, so a rollback block written as `except Exception` is skipped entirely, and only the main thread ever receives it, so worker threads keep running against the half-written state. The fixes are to shrink the critical section to one atomic rebinding, to use `try`/`finally` or an `except BaseException` block that rolls back and re-raises, or to install a handler with `signal.signal` that just sets an event and let the loop shut down at a safe point.

code

python · 17 lines
python
import threading

lock = threading.Lock()
memory = {"entries": 0}

def commit(key, text):
    with lock:
        try:
            memory["entries"] += 1
            memory[key] = text
        except BaseException:
            memory.pop(key, None)
            memory["entries"] -= 1
            raise

commit("hello", "hola")
print(memory)

go deeper

for a junior

Know that Ctrl-C raises KeyboardInterrupt, that it can land almost anywhere in your code, and that a handler naming Exception will not catch it. That is enough at this level.

for a middle

Explain the mechanism: the signal handler sets a flag, the interpreter raises at the next bytecode boundary, and only in the main thread. Show how try/finally still runs during the unwind while an Exception handler does not.

for a senior

Demonstrate the diagnosis and the repair order — make the mutation a single publish, then compensate and re-raise, then consider a flag-setting handler. Be ready to discuss the responsiveness tradeoff of swallowing the first interrupt.

for a principal

Own the shutdown contract for a fleet of workers: how an interrupt propagates to threads that never see the signal, what invariants must survive a hard stop, and whether recovery lives in cleanup code or in a replayable design that tolerates a torn write.

## The scenario A translation-memory updater keeps its segment cache in memory — an in-process dictionary running an 83% hit rate against incoming segments — and each commit touches two things: the entry itself and a counter used to decide when to flush. An operator presses Ctrl-C during a batch, and afterwards the counter disagrees with the number of entries. Nothing in the code can produce that state on any single line, and the same batch replayed by hand is fine. ## How Ctrl-C actually becomes an exception The interrupt is a POSIX signal, and a signal handler runs in a context where almost nothing is safe to do. Python therefore does the minimum in the C handler: it records that the signal arrived and sets a flag on the interpreter, then returns. The evaluation loop checks that flag between bytecode instructions, and *there* — back on solid ground, with the interpreter in a consistent state — it raises `KeyboardInterrupt`. The consequence is that delivery is **asynchronous with respect to your source code**. The exception appears at whichever instruction boundary happens to be next, which can be between two statements of a compound update, or between the subscript store and the counter increment. Your two-line critical section is many bytecodes, and every gap between them is a place the exception can land. This is a genuine race on shared state: the interrupt is an independent event, and the window is every instruction boundary in the mutation. Two related details matter. First, Python only runs signal handlers **in the main thread**, so a worker thread never receives `KeyboardInterrupt` — it keeps running against whatever half-updated state the main thread left behind, and it cannot even observe that a shutdown was requested unless you tell it. Second, since PEP 475 in Python 3.5, a system call interrupted by a signal is retried rather than surfacing `InterruptedError`, *unless* the Python-level handler raises — which the default interrupt handler does. So a blocking read does still get cut short by Ctrl-C; it just no longer produces a spurious `InterruptedError` for signals you handle silently. ## Why the rollback did not run The second half of the bug is hierarchy. `KeyboardInterrupt` derives directly from `BaseException`, deliberately kept off the `Exception` branch so ordinary error handling cannot swallow a shutdown request. That design is correct, and it also means a compensating block written as `except Exception:` is not entered when the interrupt lands. The state is left torn and nothing repairs it. ```python import threading lock = threading.Lock() memory = {"entries": 0} def commit(key, text): with lock: try: memory["entries"] += 1 memory[key] = text except BaseException: memory.pop(key, None) memory["entries"] -= 1 raise ``` Naming `BaseException` here is legitimate precisely because the block repairs state and immediately re-raises; it never decides to continue. A `finally` block achieves the same thing when the cleanup is unconditional, and both run during the unwind that `KeyboardInterrupt` triggers. ## The three real fixes, in order of preference **Shrink the window.** The most robust answer is to make the mutation a single rebinding rather than a sequence: build the new entry, then publish it in one operation whose intermediate state is never visible. Derive the counter from the container instead of maintaining it separately, and there is no invariant left to tear. **Repair on the way out.** Where a multi-step mutation is unavoidable, wrap it in `try`/`finally` or in a broad handler that undoes the partial work and re-raises. Keep the compensating code short — a second Ctrl-C during cleanup raises again inside your handler, so cleanup that itself takes a long time is a new hazard. **Do not let the signal raise at all.** For a batch worker, the cleanest design is cooperative: install your own handler with `signal.signal` for the interrupt signal that merely sets a `threading.Event`, and have the main loop check that event between work items and exit at a boundary you chose. The exception then never fires mid-update, and worker threads can observe the same event and wind down — which the default behaviour cannot give them, since they never see the signal. ```python import signal import threading stop = threading.Event() signal.signal(signal.SIGINT, lambda signum, frame: stop.set()) ``` The tradeoff is that you have taken responsibility for responsiveness: if the loop checks the event only rarely, Ctrl-C appears to do nothing, and operators reach for a harder kill. The usual compromise is to swallow the first interrupt into the event and restore the default handler so a second one raises immediately. ## What an interviewer is listening for That you can explain *why* delivery is asynchronous rather than just asserting it; that you connect the missing rollback to `KeyboardInterrupt` sitting outside `Exception`; that you know signal handlers run only in the main thread; and that your first instinct is to shrink the critical section rather than to add ever-broader handlers around it.

  • Why does calling sys.exit() inside a worker thread not stop the process?
    `sys.exit()` only raises `SystemExit` in the thread that calls it. Propagating out of a worker thread's run method ends that thread; the interpreter's exit logic is tied to the main thread, so the process continues. To shut down from a worker you signal the main thread — an event it polls, or a queued sentinel — and let it exit. `os._exit` terminates immediately from anywhere, but skips all cleanup.
  • Can a worker thread receive KeyboardInterrupt at all?
    Not from the signal. Python runs signal handlers only in the main thread, so a Ctrl-C raises `KeyboardInterrupt` there and nowhere else. Workers keep running until they notice a shared flag or their queue closes. The exception can still reach a worker indirectly — if the main thread propagates it while holding a resource the worker waits on — but it is never delivered to the worker directly.
  • What happens if a second Ctrl-C arrives while your cleanup code is running?
    It is delivered the same way, so a second `KeyboardInterrupt` can be raised inside the cleanup block and abandon it half-done. That is the argument for keeping compensating code short and idempotent. If cleanup genuinely must complete, the usual approach is to install a handler that ignores or defers the signal for that window and restore the previous handler afterwards.

It is like a postal worker who can only interrupt you at the end of a sentence, not the end of a paragraph: you know a pause is coming, but not which sentence it lands after, so any thought that takes three sentences to finish can be cut in half.

saying these in an interview costs you the question

  • Thinks the signal handler itself raises the exception
  • Assumes Ctrl-C can only land between whole statements
  • Writes the rollback as except Exception and expects it to run
  • Believes every thread receives KeyboardInterrupt
  • Adds broader handlers instead of shrinking the critical section
  • Claims a lock protects state from an asynchronous interrupt

context