A retry loop re-passes its 30-second timeout and catches Exception on every interruption; why does the wait never expire?
answer
- Two bugs, both in the loop
- A budget, not a per-attempt value
- Compute the deadline once
- Monotonic clock, not wall clock
- The net catches the abort too
basics
~20 sRe-passing the original timeout restarts the clock on every interruption, so a wait interrupted more often than every 30 seconds never expires. Catching Exception also swallows the abort a signal handler raised, so the shutdown request disappears with it.
solid answer
~50 sThere are two independent defects and both are in the loop, not in the platform. First, the timeout is re-passed rather than recomputed: each interruption starts a fresh 30-second budget, so if a signal arrives every few seconds — a nightly search-index rebuild reaping worker subprocesses gets `SIGCHLD` constantly across a six-hour run — the bounded wait becomes unbounded. The fix is a deadline: `deadline = time.monotonic() + timeout` once, then pass `deadline - time.monotonic()` on each pass and raise when it goes non-positive. Use `time.monotonic`, not `time.time`, so a clock adjustment cannot move the deadline. Second, `except Exception: continue` catches the exception a signal handler raised to request shutdown and restarts the wait, so the operator's signal has no visible effect. Catch only the interruption you mean. Best of all, delete the loop — since 3.5 the interpreter already retries and recomputes for you.
code
python · 32 linesimport time
interrupts = iter([0.4, 0.4, 0.4, 0.4])
def blocking_wait(timeout: float) -> None:
"""Stand-in for a wait that signals keep interrupting."""
hit = next(interrupts, None)
if hit is None or hit >= timeout:
time.sleep(timeout)
raise TimeoutError("waited the whole timeout")
time.sleep(hit)
raise InterruptedError("interrupted early")
def wait_until(timeout: float) -> None:
deadline = time.monotonic() + timeout
while True:
remaining = deadline - time.monotonic()
if remaining <= 0:
raise TimeoutError("deadline expired")
try:
return blocking_wait(remaining)
except InterruptedError:
continue
start = time.monotonic()
try:
wait_until(1.0)
except TimeoutError:
print(f"total wait bounded at {time.monotonic() - start:.2f}s")go deeper
The takeaway is that a timeout is a total budget. If you retry, work out how much time is left instead of handing the same number to the next attempt, and never catch broad exception classes just to loop again.
Explain both defects and their fixes: a deadline computed once from time.monotonic and passed as remaining time, and a narrow except clause that catches only the interruption. Note that the interpreter has done this for stdlib calls since 3.5.
Demonstrate the diagnosis, not just the fix: connect the hang to signal frequency in the real workload, get a live stack rather than adding print statements, and recognise the over-broad except as the reason the error that explained everything was thrown away.
Own the pattern across services: deadlines rather than per-attempt timeouts as a codebase convention, propagated end to end; a policy on what may be caught around a wait; and how signal-driven shutdown is expressed so an operator's request cannot be swallowed by a library's loop.
## The shape of the bug The loop under discussion looks like defensive engineering and behaves like a hang: ```python while True: try: return wait_for_batch(timeout=30.0) # same 30 every pass except Exception: continue # swallow, go round again ``` It was probably written to survive `EINTR` in the pre-3.5 era, and it carries two defects that only show up under signal pressure — which is exactly what a six-hour nightly search-index rebuild produces, because every worker subprocess that exits delivers `SIGCHLD`, and every log rotation or config reload delivers another signal on top. ## Defect one: the timeout is re-passed, not recomputed A timeout is a *budget*, but the loop treats it as a *per-attempt* value. Each interruption starts a fresh 30 seconds. If signals arrive on average more often than the timeout, the wait can never expire: the loop is in a race between the deadline and the next signal, and the signals win. On a quiet developer machine no signals arrive, the loop is never exercised, and the code ships looking correct. The fix is to convert the timeout to a deadline exactly once, before the first attempt, and to pass the remaining time on every attempt: ```python deadline = time.monotonic() + timeout while True: remaining = deadline - time.monotonic() if remaining <= 0: raise TimeoutError try: return wait_for_batch(timeout=remaining) except InterruptedError: continue ``` Use `time.monotonic`, never `time.time`. The wall clock can step backwards or forwards — an NTP correction, a manual set, a container resuming — and a deadline built on it either fires immediately or never. `time.monotonic` only moves forward, at a steady rate, which is the entire reason it exists. This is also precisely what CPython does internally. PEP 475 (Python 3.5) not only reissues an interrupted call, it recomputes the remaining timeout from a monotonic deadline, so a bounded stdlib wait stays bounded however often it is interrupted. The hand-rolled loop is re-implementing a solved problem and getting it wrong. ## Defect two: the swallowed exception `except Exception: continue` is a much bigger net than the author intended. The interruption it was written for does not even arrive as an exception any more — since 3.5 the interpreter retries transparently. What *does* arrive as an exception is everything the loop should not be eating: a connection reset, a bad descriptor, a permission error, a bug in `wait_for_batch` raising `TypeError`. Each of those now becomes an infinite busy loop instead of a stack trace, which is a far worse outcome than the original error. Worse for this scenario specifically: if the process installs a signal handler that raises a custom exception to request an orderly stop, that exception is an ordinary `Exception` subclass and this loop eats it. The operator sends the signal, the handler raises, the loop swallows it and goes back to waiting, and the rebuild refuses to stop. (Ctrl-C survives only by accident: `KeyboardInterrupt` inherits from `BaseException`, so `except Exception` does not catch it — but a bare `except:` would.) Catch narrowly. `except InterruptedError` if you genuinely mean the interruption; check `exc.errno == errno.EINTR` if you are handling a raw `OSError` from native code; let everything else out. ## How to diagnose it in the wild The symptom is a job that neither finishes nor errors. Useful moves, in order: - Confirm signal pressure is the trigger — a run that hangs only when many workers are churning, and completes fine on a small input, points straight at interruption frequency. - Get a stack from the live process. The interpreter's remote-debugging support in 3.14 (PEP 768, `sys.remote_exec`) lets you attach to a running process; `faulthandler` registered on a spare signal is the older, always-available option. Either way, the stack shows the process parked in `wait_for_batch` inside the retry loop, not making progress. - Look at what the loop catches before you look at what it waits on. In practice, the over-broad `except` is what turns a five-minute diagnosis into an all-night one, because the error that would have explained everything was discarded. ## The one-line version A timeout must be a deadline, computed once on a monotonic clock; a retry must catch only the condition it retries; and on any supported Python this particular loop should not exist at all.
- Why insist on time.monotonic rather than time.time for the deadline?Because the wall clock is not monotonic. An NTP correction, a manual clock set, or a suspended machine resuming can move `time.time` backwards or forwards, and a deadline built on it then fires instantly or effectively never. `time.monotonic` is guaranteed to move forward at a steady rate and is undefined only in its absolute value, which is exactly the property a deadline needs.
- The team wants to keep a retry loop for safety. What is the narrowest correct version?Compute the deadline once, pass the remaining time each attempt, and catch only `InterruptedError` — or an `OSError` whose `errno` is `errno.EINTR` if the call comes from native code. Everything else propagates. And say plainly that around a stdlib call the loop is redundant on any supported Python; keep it only at a genuine boundary such as a `ctypes` or C-extension call the interpreter does not wrap.
- How would you prove the wait is the hang rather than the work being slow?Get a stack from the live process rather than guessing: 3.14's remote debugging (PEP 768, `sys.remote_exec`) can attach to a running interpreter, and `faulthandler` registered on a spare signal works everywhere. A hang shows the same frame parked in the wait across repeated samples, with no progress in between; slow work shows the frame moving.
It is a kitchen timer someone resets to thirty minutes every time the phone rings; with enough calls, dinner is never ready.
saying these in an interview costs you the question
- Blames the platform instead of the retry loop
- Re-passes the original timeout on each retry attempt
- Uses time.time for a deadline across retries
- Catches Exception around a wait and continues
- Assumes signal pressure is the same in development
- Keeps a hand-rolled EINTR loop around stdlib calls