When Popen.communicate(timeout=...) raises subprocess.TimeoutExpired, what is still running and what must you do?
answer
- The deadline was yours, not the child's
- Nothing was signalled when it raised
- Two calls, not one, to clean up
- Reaping is what removes the zombie
- kill then communicate again
basics
~20 sThe child is still running, unsignalled and still holding its pipes; the timeout ended only your wait. Kill it, then call Popen.communicate again to drain the pipes and reap it, or you leak a process and a zombie.
solid answer
~50 s`subprocess.TimeoutExpired` from `Popen.communicate` means only that your patience ran out — nothing was done to the child. It is still executing, still holding the write ends of its pipes, and still owned by you. The documented recovery is `Popen.kill()` followed by a second `Popen.communicate()`: the kill stops the work, and the second call drains whatever is left in the buffers and reaps the exit status so no zombie remains. Skipping the second call is the common half-fix — the signal is delivered but the child may still be blocked, and you never collect its status. `subprocess.run` does this cleanup for you before re-raising: it kills the child and waits. A production-grade version escalates, `Popen.terminate()` first for a clean shutdown, then `Popen.kill()` after a short grace period, and treats the timeout as a real failure rather than a path that silently falls back to some previous result.
code
python · 12 linesimport subprocess
import sys
slow = [sys.executable, "-c", "import time; print('start', flush=True); time.sleep(30)"]
proc = subprocess.Popen(slow, stdout=subprocess.PIPE, text=True)
try:
out, _ = proc.communicate(timeout=0.5)
except subprocess.TimeoutExpired as exc:
print("partial, undecoded:", exc.stdout)
proc.kill()
out, _ = proc.communicate()
print(repr(out), proc.returncode)go deeper
Recall that catching subprocess.TimeoutExpired does not stop the command — you still have a live child to deal with, and the safe default is subprocess.run with a timeout.
Explain the three obligations after the exception: signal the child, drain its pipes, and reap it so no zombie is left and Popen.returncode is actually set.
Show the production version — escalate from terminate to kill with a grace period, decode and log the partial output, and treat every timeout as an alertable failure rather than a quiet fallback path.
Own the policy: where the deadline comes from, what a timed-out run is allowed to serve, and how a fallback is made visibly different from a fresh result so silent degradation cannot survive in production.
## The exception ends your wait, not the child `Popen.communicate(timeout=...)` gives *your* call a deadline. When it expires, `subprocess.TimeoutExpired` propagates and everything about the child is exactly as it was: the process is running, its stdout and stderr pipes still exist, whatever it had already written is still sitting in the parent's buffers, and the parent still owns the obligation to reap it. Nothing was signalled, nothing was closed. That is why the documented recovery has two steps: ```python try: out, err = proc.communicate(timeout=30) except subprocess.TimeoutExpired: proc.kill() out, err = proc.communicate() ``` The `kill()` sends `SIGKILL` on POSIX (`TerminateProcess` on Windows). The **second** `communicate()` is not decoration: it reads the pipes to end-of-file and calls `Popen.wait`, which is what actually collects the exit status. A killed-but-unreaped child stays in the process table as a zombie, and `Popen.returncode` stays `None`, so any later `check_returncode`-style logic sees a process that neither succeeded nor failed. ## What goes wrong when you skip a step - **No kill at all** — you catch the timeout, log it, move on. The child keeps running, keeps writing, keeps consuming CPU. Repeat that under a scheduler and the box accumulates orphaned work until it falls over. - **Kill without the second communicate** — the signal is delivered but you never reap, so a zombie remains and the parent's read-end descriptors stay open. In a long-lived process this is a descriptor leak with a hang at the end of it. - **Kill and then `wait()` instead of `communicate()`** — safe only because the child is dying; if it were merely `terminate()`d and chose to keep writing, you would be back in the original deadlock, waiting on a child blocked against a full pipe. ## What the exception carries `subprocess.TimeoutExpired` exposes `cmd` and `timeout`, and on 3.14 the partial output read before the deadline is attached to `TimeoutExpired.stdout`. Two cautions: it is *partial*, so a parser that assumes complete output will misread it, and it is attached as **undecoded bytes even when the `Popen` was created with `text=True`**. Log it as bytes or decode it defensively; do not feed it to code expecting `str`. `subprocess.run` wraps this cleanup for you, but not identically on every platform: on timeout it kills the child, then on POSIX it waits, while on Windows it re-runs `communicate` to refill the exception's `stdout` and `stderr`. So `run` never leaks a child, and that is a good reason to prefer it whenever you do not need incremental control. ## Escalation, and what a timeout means to your system `Popen.kill()` is abrupt: `SIGKILL` cannot be caught, so the child gets no chance to flush a file, remove a lock or finish a database transaction. In production the usual shape is graceful-then-forceful — `Popen.terminate()` (`SIGTERM`), a short wait, then `Popen.kill()` if the process is still alive. Note that both signal only the direct child; a shell wrapper or a tool that spawned its own workers will leave descendants behind, which is a separate problem with a separate solution. The policy question matters as much as the mechanics. Consider a ticket-triage bot that shells out to a classifier over a 6,800-row batch with a 30-second budget. The batch grows, the classifier stops finishing in time, and the `except subprocess.TimeoutExpired` branch quietly falls back to the previously cached scores. Every run now serves a stale cached value; the dashboard is green because no exception escaped, and triage silently degrades for days. The lesson is that a timeout is a failure with a cause, not a routing decision: increment a counter, log the command and the partial output, and make the fallback visibly distinct from a fresh result. Set the deadline from the work's real distribution rather than a round number, and remember that a hard timeout is also your last defence against the pipe-buffer deadlock — with one, an unbounded hang becomes a bounded, alertable error. ## Interview shape Say the exception did not touch the child; give the kill-then-drain-then-reap sequence; mention that `subprocess.run` does it for you; then add the escalation and the alerting policy. The candidates who stop after `proc.kill()` are the ones who leave zombies in production.
- How does subprocess.run's timeout handling differ from doing it yourself with Popen?`subprocess.run` catches the timeout internally, calls `Popen.kill()`, and then cleans up before re-raising — on POSIX it waits for the killed child, on Windows it re-runs `communicate` so the exception carries the final output. You therefore never leak a running child or a zombie. The tradeoff is that you get no chance to try a graceful `Popen.terminate()` first, and no incremental access to the output while it runs.
- Why prefer Popen.terminate over Popen.kill as the first step?`Popen.kill()` sends `SIGKILL`, which cannot be caught: the child dies immediately with no chance to flush buffers, delete temporary files, release a lock or roll back a transaction. `Popen.terminate()` sends `SIGTERM`, which a well-behaved program handles to shut down cleanly. The production pattern is terminate, wait a short grace period, then kill only if it is still alive.
- Can you trust the partial output attached to the raised subprocess.TimeoutExpired?Only as a diagnostic. It is whatever had been read when the deadline hit, so it may end mid-line or mid-record and any parser expecting a complete document will misbehave. On 3.14 it also arrives as raw bytes even when the `Popen` was created with `text=True`, so decode it explicitly before logging. Treat it as evidence for a human, never as a result.
saying these in an interview costs you the question
- Believes the timeout kills the child for you
- Calls Popen.kill and skips the second communicate
- Leaves the child running and just logs the timeout
- Treats partial output as a complete result
- Falls back to stale data without alerting
- Assumes kill reaches the child's own descendants