A telemetry collector arms signal.alarm per sensor read; why do descriptors leak and unrelated code raise TimeoutError?
answer
- One timer, many guarded regions
- Cancelled only on the failure path
- The exception lands between two bytecodes
- A gap between acquiring and protecting
- finally disarms and restores the handler
basics
~20 sAlmost certainly the alarm is never disarmed on the fast path, so it fires later in code that never asked for a deadline. And because the exception lands at an arbitrary bytecode boundary, it can arrive after a socket is opened but before anything is responsible for closing it.
solid answer
~50 sTwo defects with one root cause. There is a single wall-clock timer per process, so arming `signal.alarm(5)` before a read that returns in 40 ms leaves nearly five seconds pending; nothing cancels it, and the `TimeoutError` surfaces in whatever the main thread is doing when it fires — a different sensor, the cleanup path, or an unrelated test in a 27-minute regression suite. Second, the handler runs between bytecodes, so the exception can land in the window after a connect or open call has returned but before the object is bound into a `with` or a `try`. Nothing closes it and descriptors accumulate until the process hits its limit. The fix is a context manager that arms with `signal.setitimer` and, in `finally`, disarms and restores the handler returned by `signal.signal` — plus keeping every acquisition inside a `with` in the guarded region.
code
python · 17 linesimport signal, time
def on_alarm(signum, frame):
raise TimeoutError("deadline")
signal.signal(signal.SIGALRM, on_alarm)
def read_batch():
signal.alarm(1)
time.sleep(0.1) # the read finished well inside the budget
return "batch" # ...and nothing disarmed the timer
print(read_batch())
try:
time.sleep(2) # unrelated work pays for it
except TimeoutError as exc:
print("stale alarm fired in unrelated code:", exc)go deeper
Take away one habit: whatever arms a timer must disarm it in a finally block. If a timeout error appears in code that has no timeout of its own, suspect an alarm someone else left running rather than the code in the traceback.
Explain why the exception can land anywhere in the guarded region and what that does to a resource acquired but not yet protected by a with statement. Know that signal.signal returns the previous handler and why that value matters.
Demonstrate the diagnosis under production conditions: assert the timer is clear at region boundaries, correlate descriptor growth with timeout events, and narrow the guarded region to the blocking call itself. Package arm, disarm and handler restoration into one context manager rather than scattering them.
Own the layering decision. Deadlines belong in the transport, with an out-of-band signal reserved as a coarse backstop, and workloads that cannot honour a deadline in-process belong behind a process boundary you can kill. Decide who in the codebase is allowed to own SIGALRM at all.
This is the failure mode that makes people distrust `signal.alarm`, and it is worth taking apart carefully, because both halves are consequences of the same two facts: there is **one** wall-clock timer for the whole process, and the exception it produces lands **wherever the main thread happens to be**. ### Symptom one: the stale alarm The collector arms a five-second deadline around each sensor read. Nearly all reads finish in tens of milliseconds. If the code disarms only in the `except TimeoutError` branch — a very common shape — then the success path leaves 4.96 seconds pending on a process-wide timer. Some later stretch of main-thread work is interrupted instead: the next sensor's setup, a flush, an unrelated code path in the same process. The traceback points at a frame that never mentions a deadline, which is why this is usually misdiagnosed as a flaky sensor. It is worst under test. A 27-minute regression suite runs thousands of guarded regions in one process, and a stale alarm from test 40 fails test 41 at a random point. The failure does not reproduce in isolation, which is the signature of this bug rather than a reason to dismiss it. The cure is mechanical: disarm in `finally`, not in `except`. ### Symptom two: the leaked descriptor CPython runs the Python handler between bytecode instructions, so `TimeoutError` can be raised at *any* boundary in the guarded region. Consider: ```python sock = socket.create_connection(addr) # returns, then the alarm fires here with sock: # never reached payload = sock.recv(4096) ``` The connection exists and its descriptor is open, but the exception unwinds before `with` takes responsibility for it. The same hole exists in the hand-rolled form — if the alarm fires between `f = open(path)` returning and the name being bound, the `finally` that calls `f.close()` raises `NameError` instead, hiding the original timeout. Repeat that a few thousand times over a long run and the process hits its descriptor limit and starts failing with an OS error that looks nothing like a timeout. The narrower you make the guarded region, the smaller the window. Guard the *blocking call*, not the whole acquire-use-release sequence, and let a `with` own every resource inside it. Where the acquisition itself is the thing that can hang, accept the window and add a defensive sweep: track the objects you created and close them in the handler-free part of the unwind. ### Symptom three, usually also present: the hijacked handler `signal.signal` returns the handler it replaced. Code that ignores that return value permanently installs its own `SIGALRM` handler for the process. Anything else in the interpreter that used `SIGALRM` — a library, a debug tool, an outer deadline — now delivers into your handler. Save the previous handler and restore it, remembering that it is often `signal.SIG_DFL`. ### The shape that fixes all three ```python import signal from contextlib import contextmanager @contextmanager def deadline(seconds): def on_alarm(signum, frame): raise TimeoutError(f"exceeded {seconds}s") previous = signal.signal(signal.SIGALRM, on_alarm) signal.setitimer(signal.ITIMER_REAL, seconds) try: yield finally: signal.setitimer(signal.ITIMER_REAL, 0) signal.signal(signal.SIGALRM, previous) ``` `signal.setitimer` rather than `signal.alarm` because a per-read budget is rarely a whole number of seconds. The `finally` disarms unconditionally and restores the previous disposition. `signal.getitimer(signal.ITIMER_REAL)` at a region boundary returns `(0.0, 0.0)` when the manager is behaving, which makes the invariant assertable in tests. ### How to confirm the diagnosis before you change anything Assert `signal.getitimer(signal.ITIMER_REAL) == (0.0, 0.0)` on entry to each guarded region; a non-zero delay means someone upstream left a timer armed. Log the descriptor count periodically and correlate its slope with timeout events. And read the tracebacks for the *frames*, not the exception type: `TimeoutError` raised inside a function that takes no timeout argument is the tell. ### The deeper judgement The out-of-band deadline is a last resort. If the sensor transport exposes a timeout of its own, use it — a call that fails cleanly at its own boundary leaks nothing, needs no signal machinery, and works off the main thread. Keep `signal.alarm` or `signal.setitimer` as a coarse outer backstop with real slack above the transport's own timeout, so it fires only when the inner mechanism has genuinely failed. And when the blocking work is compiled code that never returns to the eval loop, no signal will save you: only running it in a child process you can terminate gives a hard bound.
- How would you prove it is a stale alarm rather than a genuinely slow sensor?Assert `signal.getitimer(signal.ITIMER_REAL)` is `(0.0, 0.0)` on entry to every guarded region; a non-zero delay proves an earlier region left one armed. Then read the traceback frames rather than the exception type — a TimeoutError raised inside code that takes no deadline argument cannot be that code timing out. Correlating the descriptor count's slope with timeout events confirms the second half.
- Why is restoring the previous SIGALRM handler part of the fix and not a nicety?`signal.signal` returns the handler it replaced, and there is one disposition per signal for the whole process. Code that discards that return value permanently owns SIGALRM: an outer deadline, a library, or a debugging tool that also uses it now delivers into your handler and raises your exception. Save the previous value — often `signal.SIG_DFL` — and restore it in the same `finally` that disarms.
- When would you abandon signals here and use a child process instead?When the blocking work is compiled code that never returns to the eval loop, so the handler cannot run until it finishes, or when the read must happen off the main thread, where Python handlers never run. A child process gives a hard bound: you terminate it and the kernel reclaims its descriptors. The cost is serialization and startup, which is why it is the second choice, not the first.
- How should the alarm deadline relate to a timeout the transport already offers?Use the transport's own timeout as the primary mechanism — it fails at its own boundary, leaks nothing and works off the main thread. Keep the alarm as a coarse backstop set well above it, so it fires only when the inner timeout has itself failed. Two deadlines set to similar values just produce ambiguous failures.
It is a hotel wake-up call you booked for a nap you cut short: nobody cancelled it, so the phone rings in the middle of someone else's meeting.
saying these in an interview costs you the question
- Blames the sensor rather than the un-disarmed timer
- Disarms in the except branch instead of finally
- Thinks the exception can only land at the guarded call
- Leaves the previous SIGALRM handler permanently replaced
- Relies on garbage collection to close leaked sockets
- Arms a nested deadline without saving the remaining time