skip to content

Your CI runner kills a build step with exec.CommandContext, yet processes it spawned keep running. Why?

level: seniorimportance: nice to knowfreq 30%

answer

  1. you signalled one process, not a family
  2. the shell had children of its own
  3. PID 1 adopts whatever is left running
  4. a negative pid addresses a whole group
  5. the default cancel needs replacing too

basics

~10 s

Because os.Process.Kill signals exactly one process: the shell you launched. Its own children are never told, are re-parented to PID 1, and keep running. Killing a tree needs a process group.

solid answer

~50 s

`exec.CommandContext` kills the process it started and nothing else. A build step is usually `/bin/sh -c "..."`, so the shell dies and the compiler, test binary or daemon it forked survives, gets re-parented to PID 1, and keeps holding the workspace open. The fix is to give the child its own process group at launch — `cmd.SysProcAttr = &syscall.SysProcAttr{Setpgid: true}` on Unix — and then signal the group rather than the process, by passing a negative PID to `syscall.Kill`, since the new group's ID equals the child's PID. You must also replace `cmd.Cancel`, because the default still kills only the direct child, and remember that the fallback kill after `Cmd.WaitDelay` expires is likewise process-only. Descendants that create their own groups, or that daemonise deliberately, still escape; on a shared runner the durable answer is a container or a cgroup, not a signal.

code

go · 13 lines
go
cmd := exec.CommandContext(ctx, "/bin/sh", "-c", step)
cmd.SysProcAttr = &syscall.SysProcAttr{Setpgid: true}

// the default Cancel kills only the direct child, so replace it
cmd.Cancel = func() error {
	return syscall.Kill(-cmd.Process.Pid, syscall.SIGTERM)
}
cmd.WaitDelay = 10 * time.Second

err := cmd.Run()
// WaitDelay's own fallback is also process-only; sweep the group here
_ = syscall.Kill(-cmd.Process.Pid, syscall.SIGKILL)
return err

go deeper

for a junior

Take away the one fact: killing a process does not kill the processes it started. They are re-parented and keep running, which is why a timed-out step can leave work behind.

for a middle

Explain the mechanism — Setpgid in syscall.SysProcAttr at launch, a negative PID to syscall.Kill at cancellation — and why the default cancel function has to be replaced for any of it to take effect.

for a senior

Diagnose it from the host: survivors with PPID 1 carrying the old group ID and holding the workspace open, and know the remaining escapes, including that WaitDelay's fallback kill is process-only.

for a principal

Decide the posture. Signals are cooperative, so if the promise is that nothing from a job outlives it, the enforcement belongs in a per-job cgroup or container, and the question becomes what you allow user-supplied steps to detach in the first place.

## The symptom A CI runner executes user-supplied build steps on a shared host. A step exceeds its deadline, the runner's context fires, `os/exec` kills the command, `cmd.Wait` returns, and the runner reports a clean timeout. Some time later the host is inspected and the picture is different: a test binary from that job is still running, still writing, and still holding the job's workspace open so the cleanup that should have deleted it silently failed. Its parent PID is 1. ## Why signals stop at the child When a Go program launches a process it gets back exactly one PID, and every signalling API in the standard library — `os.Process.Kill`, `os.Process.Signal` — targets that one PID. There is no "and its descendants" variant, because Unix does not have one at the process level. A build step is almost never a single process. `/bin/sh -c "make test"` is a shell that forks a build tool that forks a compiler, and each of those may fork more. Killing the shell removes one node from the middle of that tree. The rest are not notified: a Unix child is not signalled when its parent dies. They are simply re-parented to PID 1 and continue, which is why they show up later with a PPID of 1 and the original job's working directory still open. This is why the diagnosis is a process listing on the host after the runner has exited. A tree view of the survivors — their PIDs, their parent PIDs, and their process group IDs — tells you immediately that they share the old step's group ID while their parent is now init, which is the signature of exactly this bug. ## Process groups: the mechanism that does exist Unix does have a grouping abstraction one level up: the process group. Every process belongs to one, `kill` can address a whole group, and a negative PID is how you say so — `kill(-pgid, sig)` signals every member. Go exposes this in two places. At launch, `cmd.SysProcAttr = &syscall.SysProcAttr{Setpgid: true}` (a Unix-only struct) asks the runtime to put the child into a new process group of its own. The new group's ID is the child's PID, and — importantly — the descendants it forks inherit that group unless they deliberately leave it. At kill time, `syscall.Kill(-cmd.Process.Pid, syscall.SIGKILL)` signals every member of that group. Put together with the cancellation fields, a correct runner looks like this: ```go cmd := exec.CommandContext(ctx, "/bin/sh", "-c", step) cmd.SysProcAttr = &syscall.SysProcAttr{Setpgid: true} cmd.Cancel = func() error { return syscall.Kill(-cmd.Process.Pid, syscall.SIGTERM) } cmd.WaitDelay = 10 * time.Second ``` The custom `Cancel` is not optional. The default that `exec.CommandContext` installs calls `Kill` on the process, which is exactly the process-only behaviour you are trying to escape. And note the remaining gap: when `Cmd.WaitDelay` expires, the package's own fallback also terminates only the direct child. A runner that must guarantee nothing survives sends a group-wide SIGKILL itself after `Wait` returns. ## What still escapes Process groups are a big improvement and not a guarantee. A descendant can call `setsid` or `setpgid` and leave the group deliberately — which is what daemonising means, and user-supplied build steps do it by accident all the time when they start a background service "to be helpful". A process stopped or in an uninterruptible state may not be reaped promptly. And a step that writes its own supervisor is free to re-launch anything you killed. There are two Unix details worth knowing here. On Linux `syscall.SysProcAttr` also offers `Pdeathsig`, which asks the kernel to signal the child when its parent dies — but it fires on the death of the parent *thread*, which in a multi-threaded Go program is not the same as the process exiting, and it covers only the direct child. It is a useful backstop, not a solution. And `Setpgid` has a side effect: the child is no longer in the runner's group, so it will not receive a terminal-generated interrupt aimed at that group. Whatever the runner did with those signals, it now has to forward deliberately. Windows has no equivalent field. `syscall.SysProcAttr` there is a different struct entirely, and the real answer on that platform is a job object, which the standard library does not wrap. ## The decision behind the code For a platform engineer bounding what user-supplied code may leave behind, the process-group technique is the floor, not the ceiling. It reliably cleans up the well-behaved 95% and turns "processes randomly survive" into "processes survive only when a step actively detached". If the requirement is absolute — nothing from a job outlives the job — the enforcement has to come from a boundary the child cannot step out of, meaning a per-job cgroup or a per-job container that is torn down as a unit. Signals are cooperative at the edges; a resource boundary is not. The practical middle ground most runners land on: launch every step in its own process group, terminate the group with SIGTERM on cancellation, escalate to a group-wide SIGKILL after a bounded delay, then sweep — after `Wait` returns, check whether anything in that group still exists and log it loudly, because a step that survives a group kill is a step that detached, and that is a policy question about what you allow, not a bug in the runner.

  • Why is setting SysProcAttr with Setpgid not enough on its own?
    It only creates the group. The cancellation path still has to use it: `exec.CommandContext`'s default `Cancel` calls `Kill` on the single process, and the fallback kill after `Cmd.WaitDelay` expires is process-only as well. Without replacing `Cancel` with a negative-PID `syscall.Kill`, you have built the group and then never signalled it.
  • What can still survive a group-wide SIGKILL?
    Anything that left the group deliberately — a descendant that called `setsid` or `setpgid`, which is exactly what daemonising does. Processes stuck in an uninterruptible state also linger. If the requirement is that nothing outlives the job, the boundary has to be a cgroup or a container torn down as a unit rather than a signal.
  • How would you confirm this diagnosis on the host after a job finished?
    List the surviving processes with their parent PID, process group ID and working directory. The signature is a set of processes whose PPID is 1, whose group ID still matches the killed step, and whose working directory is the job's workspace — which also explains why the workspace cleanup failed to remove the directory.
  • Does Setpgid have any downside for the parent program?
    Yes. The child leaves the parent's process group, so a terminal-generated interrupt aimed at that group no longer reaches it. Anything the runner relied on there — Ctrl-C reaching the whole tree during local use, for instance — now has to be forwarded explicitly by the runner's own signal handling.

Sending the manager home does not close the office. You have to address the whole floor, which means the floor had to be assigned a number when the team moved in.

saying these in an interview costs you the question

  • Believes killing a parent process terminates its children
  • Sets Setpgid but leaves exec.CommandContext's default Cancel in place
  • Thinks Cmd.WaitDelay's fallback kill covers the whole process tree
  • Assumes a negative PID passed to syscall.Kill is an error
  • Claims process groups guarantee nothing survives, ignoring setsid