skip to content

Popen.terminate() sometimes fails to stop a translation-memory updater's export child; how would you build a shutdown that always ends it?

level: seniorimportance: should knowfreq 46%

answer

  1. Shutdown is a sequence, not a call
  2. Ask first, insist afterwards
  3. A deadline between the two signals
  4. Signal the group, not the pid
  5. Wait again after the second signal

basics

~20 s

Signal the child's whole process group with SIGTERM, wait a bounded grace period, then send SIGKILL to the same group and wait again. Start the child with start_new_session=True so that group exists and excludes your own process.

solid answer

~40 s

Shutdown is a three-stage sequence, not a call. Spawn the child with `start_new_session=True` and cache `os.getpgid(p.pid)` immediately. To stop it, `os.killpg(pgid, signal.SIGTERM)`, then `p.wait(timeout=grace)`: the grace period is what lets the exporter finish the record it is midway through writing, because a locale-dependent output format has no marker that distinguishes a truncated final record from a complete one, and the next run silently mis-parses it. If the wait times out, `os.killpg(pgid, signal.SIGKILL)` and `p.wait()` again — SIGKILL is not instantaneous, and only the second wait proves the child is gone rather than signalled. Wrap both signals for `ProcessLookupError`, which just means the group already emptied, and log which stage ended the child so a rising share of SIGKILL exits is visible rather than silent.

code

python · 22 lines
python
import os, signal, subprocess


def shutdown(p, pgid, grace=5.0):
    if p.poll() is not None:
        return p.returncode
    try:
        os.killpg(pgid, signal.SIGTERM)
    except ProcessLookupError:
        return p.wait()
    try:
        return p.wait(timeout=grace)
    except subprocess.TimeoutExpired:
        try:
            os.killpg(pgid, signal.SIGKILL)
        except ProcessLookupError:
            pass
        return p.wait()


p = subprocess.Popen("sleep 60 & wait", shell=True, start_new_session=True)
print("returncode:", shutdown(p, os.getpgid(p.pid), grace=2.0))

go deeper

for a junior

Learn the shape before the details: ask the child to stop, give it a little time, then force it. Never start with the forceful signal if the child writes files.

for a middle

Be ready to write the loop: killpg with SIGTERM, wait with a timeout, catch TimeoutExpired, killpg with SIGKILL, wait again, and handle ProcessLookupError at both signals.

for a senior

Show that you sized the grace period against both the child's work and your own shutdown budget, made the routine idempotent and observable, and can say exactly what state a SIGKILL leaves behind.

for a principal

Argue the boundary: which of this belongs in application code at all versus in a supervisor or container runtime, and what contract you publish to teams about cleanup time and data durability on cancellation.

"Terminate then kill" is one of the few genuinely universal patterns in process management — it is what init systems, container runtimes and orchestrators all do — and an interviewer asking this wants the whole state machine, including the parts people leave out. ## The scenario, concretely A translation-memory updater runs nightly for an 11-person team. For each locale it spawns an export child that streams entries out to a file, formatting dates and numbers in that locale's own conventions. When the run is cancelled, the current code calls `Popen.terminate()` and moves on. Two failure modes follow. Some children are still alive afterwards, because SIGTERM reached a wrapper and not the exporter, or because the exporter was inside a call that delayed delivery. Others were killed midway through writing a record, and because the output format is locale-dependent there is no unambiguous end-of-record marker — a truncated line looks like a shorter but valid one, so the next run merges corrupt entries instead of failing loudly. Both failures come from treating shutdown as a single call. ## Stage 0: make the target signalable, at spawn time You cannot fix this at shutdown time. Start the child with `start_new_session=True`, which calls `os.setsid()` in the child so it leads a new session and process group and every helper it forks inherits that group. Cache the group id right away with `os.getpgid(p.pid)` — after `Popen.wait()` reaps the child the pid is gone and may be recycled, so a late `os.getpgid()` either raises `ProcessLookupError` or, worse, resolves against an unrelated process. ## Stage 1: ask, with a deadline `os.killpg(pgid, signal.SIGTERM)` reaches the exporter and anything it started. Then `p.wait(timeout=grace)`. The grace period is a real design parameter, not a magic number: it should be a little longer than the child's worst realistic time to finish the record it is on and close the file, and shorter than whatever deadline you are operating under — the job's own timeout, the deployment's shutdown budget, the orchestrator's grace period before it SIGKILLs *you*. Being killed yourself mid-escalation is the failure people forget, so your budget has to fit inside your supervisor's. For this workload the child should also make the grace period useful: writing to a temp file and renaming it into place on success means an interrupted export leaves no half-file at all, which is a stronger guarantee than any grace period alone. ## Stage 2: escalate, and wait again On `subprocess.TimeoutExpired`, `os.killpg(pgid, signal.SIGKILL)` and then `p.wait()` with no timeout. The second wait is not decoration. `os.killpg()` returning only means the signal was queued; the child is not gone until the kernel has torn it down and you have reaped it. Skipping that wait is how you get a shutdown routine that reports success while the process still holds the output file. Both `killpg` calls can raise `ProcessLookupError` — the group emptied between your check and your signal — and that is a normal, ignorable outcome, not an error to propagate. Guard the whole sequence with an early `p.poll() is not None` return so you never signal a group id whose members have all exited. ## Stage 3: make it observable and reentrant Log which stage ended each child: exited on its own, exited after SIGTERM, or required SIGKILL. That single field turns an invisible problem into a trend — if the SIGKILL share climbs after a release, some child stopped honouring SIGTERM, and you will know before the corrupt output shows up in the memory. Make the routine idempotent, because it will be called concurrently from a signal handler and from the normal cleanup path, and both may fire during a real shutdown. ## Propagating your own shutdown The parent must also handle being asked to stop. Install `signal.signal(signal.SIGTERM, handler)` in the parent and have the handler run the same escalation for every live child before exiting. Do not rely on `atexit` alone: it runs on a normal interpreter exit but not when the parent itself is SIGKILLed, and in that case nothing of yours runs at all. That is the argument for the process group being a real one — an external supervisor, a systemd unit or a container runtime can then clean up the whole group even when your parent dies without warning. One trap: `with subprocess.Popen(...) as p:` does not kill anything. `__exit__` closes the pipes and then calls `wait()`, so a child that never exits turns the context manager into an indefinite hang. If you want the block to bound the child's life, the escalation belongs in a `finally`. ## What the sequence buys you SIGTERM first gives well-behaved children their cleanup path; the deadline stops a hung child holding the pipeline; SIGKILL guarantees termination; the group makes both signals reach the whole tree; the second wait makes the outcome true rather than assumed. Drop any one of the five and you have one of the two failures the team started with.

  • How would you choose the length of the grace period between SIGTERM and SIGKILL?
    From two directions. The floor is the child's realistic worst case for finishing its current unit of work and closing its files — measure it rather than guess. The ceiling is your own shutdown budget: whatever supervisor, unit or runtime manages your parent will eventually SIGKILL *you*, and your whole escalation has to complete inside that window. Make it configurable, and log every escalation so the number can be revisited with evidence.
  • Why call Popen.wait() again after os.killpg with SIGKILL, rather than assuming the child is gone?
    Because killpg only queues the signal. The child is not gone until the kernel has torn it down and the parent has reaped it, and until then it can still hold the output file and its inherited descriptors. The second wait is what makes the routine's return value truthful; without it you report a clean shutdown while the process is still finishing.
  • How do you make sure children are stopped when the parent process itself is asked to shut down?
    Install signal.signal(signal.SIGTERM, handler) in the parent and run the same escalation for every live child from it, with the routine written to be idempotent because the normal cleanup path may run concurrently. atexit alone is not enough — it runs on a normal exit but not when the parent is SIGKILLed. For that case, rely on the children being in a process group that an external supervisor or container runtime can clean up.
  • Does using `with subprocess.Popen(...) as p:` bound the child's lifetime?
    No. Popen.__exit__ closes the pipes and then calls wait() with no timeout, so a child that refuses to exit turns the with-block into an indefinite hang rather than a terminated child. The context manager guarantees the descriptors are closed and the process is reaped, not that it stops. If you want a bound, put the terminate-grace-kill escalation in a finally block.

saying these in an interview costs you the question

  • Sends SIGKILL immediately, corrupting half-written output
  • Sends SIGTERM and never verifies the child exited
  • Waits without a timeout after terminate
  • Signals only p.pid, leaving grandchildren alive
  • Assumes killpg returning means the child is gone
  • Relies on atexit to clean up children

context