skip to content

When subprocess.run's timeout= expires in a nightly index rebuilder, what happens to the child and what can survive?

level: seniorimportance: should knowfreq 42%

answer

  1. The clock is wall-clock, and it is per call
  2. Something dies, but only one thing
  3. A shell in the middle changes who dies
  4. Groups are how you kill a whole tree
  5. Two sibling exceptions, one shared base

basics

~20 s

subprocess.run kills the direct child, waits for it, then re-raises TimeoutExpired carrying the command and whatever output was captured. Only that one process dies: under shell=True the shell dies while the program it launched keeps running, so grandchildren survive and can hold the pipes open.

solid answer

~50 s

`timeout=` bounds the wall-clock wait inside `run()`, which delegates to `Popen.communicate(timeout=...)`. When it expires, `run()` kills the child, waits for it to be reaped, and re-raises `subprocess.TimeoutExpired`, whose `cmd`, `timeout`, `stdout` and `stderr` attributes carry what was collected so far. The critical limit is that it kills exactly one process - the one `Popen` started. With `shell=True` that process is the shell, so whatever the shell launched keeps running as an orphan, still holding your pipes and still doing work you believe you cancelled. The fix is to make the child the leader of its own process group with `start_new_session=True` on POSIX and signal the whole group with `os.killpg` when you need the tree gone. Note also that `TimeoutExpired` is not a `CalledProcessError`: they share `SubprocessError` as a base, so a call site with both `check=True` and `timeout=` needs both handlers.

code

python · 15 lines
python
import subprocess
import sys
import time

start = time.monotonic()
try:
    subprocess.run(
        [sys.executable, "-c", "import time; time.sleep(30)"],
        capture_output=True,
        text=True,
        timeout=1,
    )
except subprocess.TimeoutExpired as exc:
    print("gave up after", exc.timeout, "s")
    print("elapsed", round(time.monotonic() - start, 1))

go deeper

for a junior

Remember to pass timeout= at all, and to wrap the call so TimeoutExpired is handled rather than crashing the job. Know that the exception is a different type from the one check=True raises.

for a middle

Explain the sequence: communicate times out, run kills the child, waits for it, and re-raises with the captured output attached. Be able to say why env= must usually start from a copy of the current environment.

for a senior

Demonstrate the teardown judgement: only the direct child is killed, a shell in the middle orphans the real worker, and start_new_session plus a process-group signal is the fix. Talk about escalating from a polite signal to a forced one, and about what you log so on-call can tell a cancellation from a crash.

for a principal

Own the policy for scheduled work: every external call bounded, an overall budget rather than per-attempt timeouts multiplied by retries, an overrun treated as an alertable event, and one shared runner that gets the process-group teardown right so each team is not reinventing it.

A nightly search-index rebuilder owned by a four-person team calls out to a helper program under a timeout, because a rebuild that hangs must not block the next night's run. Getting that teardown right is a normal senior interview conversation. ## What run() does when the clock runs out `subprocess.run(..., timeout=N)` starts the child and calls `Popen.communicate(timeout=N)`. If the child has not exited within N seconds of wall clock, `communicate()` raises `subprocess.TimeoutExpired`. `run()` catches it, calls `kill()` on the child, waits for the process to be reaped so it does not linger as a zombie, and re-raises. The exception carries the command, the timeout value, and whatever output was captured up to that point in its `stdout` and `stderr` attributes - so if you passed `capture_output=True`, a timeout still gives you the child's last words, which is often the whole diagnosis. The timeout is wall-clock, not CPU time, and it covers only this one call. A retry loop of three attempts at 60 seconds is a three-minute worst case, not one minute; if the job has an overall budget you must track it yourself and shrink each subsequent timeout. If you use `Popen` directly the automatic kill is not there: `communicate(timeout=...)` raises and leaves the child running, and the documented recovery is `proc.kill()` followed by a second `communicate()` to reap it and collect the output. ## Only the direct child dies This is the part that bites in production. `Popen` knows one process id. `kill()` signals that process and nothing else. Any process the child itself started is untouched, and if it inherited the write end of your pipes it keeps them open - so even after you "cancelled" the work, a read on those pipes may not see EOF. `shell=True` makes this the normal case rather than an edge case, because the direct child is the shell and the program you care about is a grandchild. You kill the shell, the rebuild keeps running, the next scheduled run starts a second copy, and the two fight over the same output. The `shell=True` habit and the timeout habit interact badly in exactly this way. The POSIX remedy is a process group. Pass `start_new_session=True` so the child starts a new session and becomes the leader of its own process group, then on timeout signal the group as a whole with `os.killpg(proc.pid, ...)` - every descendant that has not deliberately left the group goes with it. Escalating politely first (a terminate, a short grace period, then a kill) gives a well-behaved child a chance to flush and clean up; a child ignoring the polite signal still dies at the end. Windows has its own equivalent involving process groups and job objects, so a cross-platform runner needs a branch. ## cwd and env, since a timed job usually sets both `cwd=` runs the child in a chosen directory. Prefer it over `os.chdir`, which changes the working directory for the entire parent process - every thread included - and is therefore an unpleasant surprise in any concurrent program. `env=` **replaces** the environment rather than adding to it. Handing over `{"INDEX_SHARD": "7"}` means the child sees only that variable: no `PATH`, no `HOME`, no locale, no proxy settings, no credentials helper. The usual consequences are a program that cannot be found, a program that cannot find its own resources, or output that decodes differently because the locale vanished. Build the mapping with `os.environ.copy()` and mutate the copy, and reserve a hand-built minimal environment for the cases where isolating the child is the actual goal - at which point put back the handful of variables it genuinely needs, `PATH` first. ## Getting the exception handling right `TimeoutExpired` and `CalledProcessError` are siblings under `subprocess.SubprocessError`, not one under the other. A call site with `check=True, timeout=60` therefore has three interesting outcomes: `TimeoutExpired` for a hang, `CalledProcessError` for a bad exit code, and `OSError` (`FileNotFoundError`) if the program could not be started at all. Only the first two are commonly handled, and code that catches `CalledProcessError` alone lets a timeout escape as an unhandled exception at 3 a.m. Also note that after a kill the exit status on POSIX is negative - `-9` for a SIGKILL - which is worth logging explicitly so the on-call reader knows the job was cancelled rather than that the program failed on its own.

  • Your timeout fires but the work carries on. What is the most likely cause?
    The process you killed was not the process doing the work - typically `shell=True`, so `Popen` owned a shell and the real program was its child, or the program itself forks a worker. Start the child with `start_new_session=True` so it leads its own process group, and on timeout signal the group with `os.killpg` rather than the single pid.
  • Why prefer cwd= over calling os.chdir before the subprocess call?
    `os.chdir` mutates the working directory of the whole parent process, so every thread and every later relative path in that process is affected, and restoring it in a `finally` is racy under concurrency. `cwd=` applies to the child alone and leaves the parent untouched.
  • What breaks when you pass env={"INDEX_SHARD": "7"} to subprocess.run?
    The child gets that variable and nothing else - no PATH, no HOME, no locale, no proxy configuration. The usual symptoms are the program not being found, resources it expects being missing, or output decoding differently because the locale is gone. Start from `os.environ.copy()`, or, when isolation is the goal, add back the specific variables the child genuinely needs.
  • You set both check=True and timeout=60. Which exceptions must the call site handle?
    `subprocess.TimeoutExpired` for a hang, `subprocess.CalledProcessError` for a non-zero exit, and `OSError` (in practice `FileNotFoundError`) if the program could not be started. The first two are siblings under `SubprocessError`, so catching one does not catch the other.

saying these in an interview costs you the question

  • Thinks a timeout terminates the whole process tree
  • Assumes CalledProcessError also covers a timeout
  • Believes env= adds variables to the inherited environment
  • Uses os.chdir instead of cwd= in a threaded program
  • Expects timeout to bound CPU time rather than wall clock
  • Thinks a timeout discards the output captured so far

context