A reconciliation retry loop catching BaseException ignores Ctrl-C -- how do you diagnose and fix it?
answer
- The symptom names the clause
- Only two forms can catch it
- Grep before you reach for a debugger
- Narrow the clause, shrink the block
- Broad catch must end in bare raise
basics
~20 sThe loop's handler catches BaseException, so KeyboardInterrupt is treated as a retryable failure and each Ctrl-C only aborts the current attempt. Narrow the handler to Exception, or keep it broad, run cleanup, and end it with a bare raise.
solid answer
~50 sThe symptom -- the process logs a failure per Ctrl-C and keeps going -- says the interrupt is being caught rather than propagating. `KeyboardInterrupt` is a direct subclass of `BaseException`, so only a bare `except:` or an explicit `except BaseException:` can swallow it; `except Exception:` cannot. So the diagnosis is a grep, not a debugger session: find the broad handler wrapping the retry body. If the process is already stuck and you cannot read the source, 3.14's remote debugging (PEP 768, `sys.remote_exec`) or a `faulthandler` traceback dump will show you which handler the loop is sitting in. The fix is to narrow the clause to `Exception`, keep the guarded block to the smallest unit of work, and log with `logging.exception` so failures are still visible. If cleanup genuinely must run on interrupt, catch `BaseException`, flush, and re-raise with a bare `raise`.
code
python · 10 linesdef reconcile(case):
if case == 2:
raise KeyboardInterrupt
raise ConnectionError("stale cached rate")
for case in range(1, 4):
try:
reconcile(case)
except BaseException as exc:
print("case", case, "swallowed", type(exc).__name__)go deeper
Recall the single fact that solves this: only a bare except: or except BaseException: can catch Ctrl-C. If a loop ignores the keyboard, look at the clause before you look anywhere else.
Explain the mechanics -- KeyboardInterrupt sits outside Exception, so narrowing the clause restores the interrupt while keeping the per-item resilience the loop was written for. Show the corrected handler and where the logging goes.
Demonstrate the whole diagnostic path on a live job: distinguishing a swallowed interrupt from a delayed one, getting a stack out of a running process, keeping the guarded block small, and the cleanup-then-bare-raise shape when interrupt-safe teardown is genuinely required.
Own the prevention: the lint rule promoted to a CI error, a test that asserts interrupt propagation, and a shutdown contract every long-running job in the estate follows so operators can always stop work predictably.
## The scenario A nightly payment reconciliation job replays a 340-case regression pack against the ledger. A stale cached exchange rate makes every case fail with a connection-style error, so the job's retry wrapper logs and tries the next case. An engineer presses Ctrl-C to stop it and nothing happens: a log line appears for each keystroke and the job carries on to case 341, 342, and around again. ## Reading the symptom Ctrl-C makes the interpreter raise `KeyboardInterrupt` in the **main thread**. `KeyboardInterrupt` is a direct subclass of `BaseException`, deliberately outside `Exception`, so only two clause forms can catch it: a bare `except:` and an explicit `except BaseException:`. That single fact turns "the job ignores Ctrl-C" into a search: ```console $ grep -rn -e 'except *:' -e 'except BaseException' src/ ``` The wrapper almost always looks like this: ```python for case in pack: try: reconcile(case) except BaseException as exc: logging.warning("case %s failed: %s", case, exc) continue ``` One extra word in the clause and one `continue` turn every signal the operator has into a logged retry. Two related diagnoses are worth ruling out before you commit to that answer. If the interrupt is being raised in a **worker thread**, it will never arrive there at all: signal handlers run only in the main thread, so a main thread that is blocked joining workers looks identical from outside. And if the loop body is blocked inside a long-running C call that never returns to the interpreter, the handler sets a flag but `KeyboardInterrupt` is only raised at the next bytecode boundary, so Ctrl-C appears to be ignored until that call completes. The distinguishing evidence is the log: a *swallowed* interrupt produces a log line per keystroke, while a *delayed* one produces nothing at all. ## Getting evidence from a live process You often meet this on a job you would rather not restart. Two stdlib routes: * `faulthandler.dump_traceback_later(timeout, repeat=True)` installed at startup periodically dumps every thread's stack to a file descriptor, which shows you exactly which frame the loop is parked in. * From 3.14, PEP 768 remote debugging (`sys.remote_exec`) lets you inject a snippet into a *running* interpreter by process id, so you can dump a traceback without having planned ahead. Either way you are confirming the same thing: the interrupt reached a handler, and the handler decided to continue. ## The fix Narrow the clause. `except Exception:` still gives the loop the "one bad case must not kill the run" property, and it lets `KeyboardInterrupt`, `SystemExit` and `asyncio.CancelledError` through untouched: ```python for case in pack: try: reconcile(case) except Exception: logging.exception("case %s failed", case) ``` Two supporting habits matter as much as the clause itself. Keep the guarded block to the smallest unit of work, so the handler cannot accidentally cover setup and teardown as well as the call you meant. And log with `logging.exception` inside the handler, so a failure that is genuinely being swallowed is at least visible with its traceback -- a job that silently retries a stale-cache failure 340 times is the second bug hiding in this scenario. ## When you really do need BaseException Sometimes an interrupt must not leave a half-written batch behind. The correct shape catches broadly, cleans up, and re-raises: ```python try: run_pack(pack) except BaseException: flush_partial_results() raise ``` The bare `raise` re-raises the exception being handled with its traceback intact, so Ctrl-C still terminates the process. A `BaseException` handler that ends in `pass`, `continue` or a `return` is always a defect. Cleanup that has no reason to inspect the exception belongs in a `finally` block or a context manager instead, which removes the temptation entirely. ## Stopping the class of bug This is a lint-able defect, and that is the durable fix: every mainstream Python linter has a rule for bare and blind `except` clauses, and promoting it to an error in CI means the pattern cannot land again. Back it with a test -- a fake unit of work that raises `KeyboardInterrupt` on the third item, asserting the runner propagates it -- so the guarantee is executable rather than a convention. On a long-running job, also make the interrupt path do something visible: a handler that logs "shutting down" on the way past is how the operator learns the difference between a job that ignored them and a job that is finishing a batch first.
- The job's main thread is blocked joining worker threads. How does that change your diagnosis?Signal handlers run only in the main thread, so `KeyboardInterrupt` is raised there and never inside a worker. If the main thread is parked in a join, the interrupt does arrive, but the workers keep running until they notice a shutdown flag -- and non-daemon threads hold the process open. The tell is that no per-keystroke log line appears, unlike the swallowed case.
- How do you stop this defect from reappearing across the codebase?Make the linter rule for bare and blind `except` clauses an error in CI rather than a warning, so the pattern cannot merge. Add an executable guarantee alongside it: a test whose fake unit of work raises `KeyboardInterrupt` partway through the batch and asserts the runner propagates it instead of continuing. Convention alone decays; a failing build does not.
- Ctrl-C produces no output at all and the process only stops seconds later. What is happening?That is delay, not swallowing. The signal handler sets a flag, and `KeyboardInterrupt` is raised when the interpreter next reaches a bytecode boundary -- so a long-running native call that does not return to the interpreter defers it until it finishes. The fix is on the blocking side: bounded timeouts, or work split into chunks that return control regularly.
saying these in an interview costs you the question
- Blames the terminal or the signal instead of the handler
- Says except Exception is what swallowed the interrupt
- Reaches for kill -9 without finding the handler
- Replaces the broad catch with a broad catch plus continue
- Catches BaseException for cleanup and never re-raises
- Assumes KeyboardInterrupt is delivered to every thread