A nightly Python report job hangs in a while loop some nights — how do you diagnose it and guarantee termination?
answer
- Get evidence before restarting the job
- Which loop, and what does its condition read?
- Find the path that changes nothing
- Name a decreasing bounded quantity
- Deadline and iteration cap as backstop
basics
~20 sDump a traceback from the live process to see which loop spins, then find the body path that fails to shrink what the condition reads. Fix that path, and bound the loop with a deadline and an iteration cap.
solid answer
~50 sFirst prove where it is stuck rather than restarting: `faulthandler.dump_traceback_later(seconds, exit=True)` armed at start-up turns a silent 6-hour hang into a traceback for every thread, and `faulthandler.register` on a signal lets you dump on demand. On 3.14, `sys.remote_exec` (PEP 768) can inject an inspection script into the running interpreter without a restart. The traceback names the loop; the cause is almost always a **path through the body that makes no progress** — an intermittent timeout leaves the work item in place, so the length the condition tests never decreases. The durable fix is to state the loop variant: a bounded quantity every path must strictly decrease, asserted in the body so a stalled pass raises instead of spinning. Add a wall-clock deadline with `time.monotonic()` and an iteration cap so the failure is a loud `TimeoutError` on a bad night, not a hang.
code
python · 12 linesimport time
deadline = time.monotonic() + 5.0
pending = [3, 2, 1]
while pending:
if time.monotonic() > deadline:
raise TimeoutError(f"stalled with {len(pending)} items left")
before = len(pending)
pending.pop()
if len(pending) == before:
raise RuntimeError("loop made no progress")
print("drained")go deeper
Know that a loop hangs when the body stops changing what the condition reads, and that the first move is capturing evidence — a stack dump from the stuck run — rather than restarting and hoping.
Explain the mechanics: a watchdog armed at start-up, reading the traceback to find the spinning loop, and spotting the branch where the collection or counter is left unchanged.
Show production judgement: name the loop variant, make every path decrease it, assert progress per pass, and add a monotonic deadline so a bad night ends in a loud failure with the remaining work in the message.
Own the standard — no unbounded loop in a scheduled job without a stated variant and an outer bound — and weigh the cost of failing a run fast against completing it partially when the job feeds downstream systems.
### Step one: prove where it is stuck A job that hangs intermittently is usually restarted, succeeds, and teaches nobody anything. The first job is to make the next hang produce evidence. **Arm a watchdog before the work starts.** `faulthandler.dump_traceback_later(seconds, exit=True)` schedules a dump of every thread's stack after the given number of seconds and then kills the process. Called once at start-up with a bound comfortably above a healthy run, it converts a silent hang into a traceback that names the exact source line the interpreter is executing. Cancel it on the normal path with `faulthandler.cancel_dump_traceback_later()` so a healthy run is unaffected. **Make dumps available on demand.** `faulthandler.register(signal.SIGUSR1)` installs a handler that dumps stacks when the signal arrives, so an operator watching a run go long can sample it twice a minute apart. Two samples pointing at the same loop, with the same locals unchanged, are much stronger evidence than one. **Attach to the live process.** Python 3.14 added remote debugging support (PEP 768): `sys.remote_exec` asks a running interpreter, identified by process id, to execute a script at its next safe point. That lets you dump state from the stuck process without a restart and without having planned ahead — the modern replacement for guessing. What you are looking for is narrow: which `while` is spinning, and what its condition reads. ### Step two: find the path that makes no progress A `while` re-evaluates its condition before every pass, so it ends only if the body changes what the condition reads. Non-termination is therefore always the same defect — **some path through the body leaves that state untouched** — and the reason it is intermittent is that the bad path is a rare one. The shape that produces "fine for months, then a six-hour hang" is a drain loop whose body puts the item back: ```python while pending: item = pending.pop() try: emit(item) except TimeoutError: pending.append(item) # length is unchanged; nothing forces it down ``` On a good night nothing times out and the list empties. On a bad night one item fails every time — a slow downstream, a row that always exceeds a limit — and the length oscillates forever. The condition is doing its job perfectly; the body simply never satisfies it. Other recurring shapes worth checking on the traceback's line: - The counter is advanced on only one branch of an `if`, and the rare branch is the one that skips it. - The condition compares floats with `!=` or `==`; binary floating point steps straight past the target and equality never holds. Use `<=`/`>=`, integer counters, or `math.isclose`. - The condition reads a length captured before the loop, or a value rebound in an inner scope, so real changes are invisible to the test. - The producer refills the collection at least as fast as the loop drains it, so the loop is *live* rather than hung — a distinction the second traceback sample makes for you, since the locals will have moved. ### Step three: make termination provable, not hoped for The fix that lasts is to name the **loop variant**: a quantity bounded below that *every* path through the body strictly decreases. `while pending:` has an obvious candidate in `len(pending)`; the buggy path above simply fails to decrease it. Once the variant is named, three changes follow: 1. **Give every path a decrement.** A failing item must leave the pending set — moved to a failures list, retried against a per-item attempt count that itself decreases, or dropped with a log line. "Put it back unchanged" must not remain a path. 2. **Assert the variant in the body.** Measure it at the top of the pass and check it at the bottom; if it did not move, raise. A loop that raises on its first stalled pass reports a defect in seconds instead of consuming a night. 3. **Bound the loop from outside.** A deadline computed with `time.monotonic()` — monotonic precisely so a clock adjustment mid-run cannot extend or collapse it — plus a maximum iteration count, raising `TimeoutError` with the remaining work in the message. This is a backstop against the variant you failed to think of, not a substitute for step one. The operational payoff is the difference between two failure modes. Unbounded, the job occupies its window, blocks whatever runs after it, and is discovered by someone noticing missing output. Bounded and asserted, it fails fast with a traceback naming the stuck item, the scheduler's alerting fires on a non-zero exit, and the partial work is visible. The loop is no more correct in the second case — but the failure is now a five-minute diagnosis instead of a morning of archaeology. ### What not to conclude Resist "it must terminate, the data is finite". Finite input guarantees nothing if the body can re-add work, and it guarantees nothing about *when*. Equally, resist raising the timeout until the hang stops reproducing: that hides the stalled path rather than removing it. The question a reviewer should always be able to answer about a `while` loop is which quantity goes down, and by how much, on every single path through its body.
- What is a loop variant, and how do you use one to argue a while loop terminates?A quantity that is bounded below — usually a non-negative integer — and that every path through the body strictly decreases. If both hold, the loop cannot run forever, because the value would have to fall below its bound. In review it is a concrete question: name the variant and point at each path that decreases it.
- Why can a while loop whose condition compares floats with `!=` never terminate?Binary floating point cannot represent most decimal steps exactly, so an accumulation overshoots the target and equality is never satisfied. Adding `0.1` ten times does not give exactly `1.0`. Use an ordering comparison, an integer counter converted at the end, or `math.isclose` with an explicit tolerance.
- How do you tell a hung loop from one that is merely slow?Sample the stack twice, a minute or so apart, with `faulthandler.register` on a signal or by attaching to the process. If the line and the loop-relevant locals are identical across samples, it is stuck; if the counter or collection size has moved, it is progressing and the question becomes throughput, not termination.
saying these in an interview costs you the question
- Restarts the job and calls the hang a fluke
- Raises the timeout until it stops reproducing
- Assumes finite input guarantees the loop ends
- Cannot get a traceback out of a live process
- Compares floats with `!=` in a loop condition
- Never names what the body must decrease