skip to content

A Windows build agent must guarantee that when a build ends, every process it spawned is gone — including grandchildren. Why isn't terminating the top process enough, and what does the OS give you?

level: seniorimportance: should knowfreq 42%

answer

  1. parent id is a label, not a leash
  2. process ids get recycled
  3. one kernel object holding a set
  4. membership is inherited by children
  5. handle closes, members die

basics

~20 s

Terminating a process on Windows kills only that process; there is no supervisory link that carries the kill to descendants. A job object does: assign the top process to a job, set JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE, and every process in the job — children included — dies when the job handle closes.

solid answer

~40 s

`TerminateProcess` acts on one process object. Windows records a parent process id but keeps no enforced parent-child supervision the way Unix process groups do, so a build tool's grandchildren happily outlive it — and because process ids are reused, walking the tree afterwards to kill them by id is a race, not a fix. The supported answer is a **job object**: `CreateJobObjectW`, then `SetInformationJobObject` with `JobObjectExtendedLimitInformation` and the `JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE` flag, then `AssignProcessToJobObject` for the root process. Anything it spawns joins the job automatically unless the job permits breakaway. When the agent closes its job handle — or crashes, so the handle closes with it — the kernel terminates every member. Jobs also cap memory, active process count, CPU rate and priority, and can report events on a completion port.

code

c · 39 lines
c
#include <windows.h>
#include <stdio.h>

int main(void)
{
    HANDLE job;
    JOBOBJECT_EXTENDED_LIMIT_INFORMATION li;
    STARTUPINFOW si;
    PROCESS_INFORMATION pi;
    wchar_t cmd[] = L"cmd.exe /c start /wait cmd.exe /c timeout /t 60";

    job = CreateJobObjectW(NULL, NULL);
    if (job == NULL) return 1;

    ZeroMemory(&li, sizeof li);
    li.BasicLimitInformation.LimitFlags = JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE;
    SetInformationJobObject(job, JobObjectExtendedLimitInformation, &li, sizeof li);

    ZeroMemory(&si, sizeof si);
    si.cb = sizeof si;
    ZeroMemory(&pi, sizeof pi);

    /* Suspended first, so nothing runs before the process is a job member. */
    if (!CreateProcessW(NULL, cmd, NULL, NULL, FALSE, CREATE_SUSPENDED,
                        NULL, NULL, &si, &pi)) {
        printf("CreateProcessW failed: %lu\n", GetLastError());
        CloseHandle(job);
        return 1;
    }
    AssignProcessToJobObject(job, pi.hProcess);
    ResumeThread(pi.hThread);

    WaitForSingleObject(pi.hProcess, 5000);

    CloseHandle(pi.hThread);
    CloseHandle(pi.hProcess);
    CloseHandle(job);   /* kills every surviving member, grandchildren included */
    return 0;
}

go deeper

for a junior

Know that killing a process on Windows does not kill the processes it started, and that a job object is the object that groups processes so they can be managed together.

for a middle

Explain job creation, assignment and the kill-on-close limit, and state that children automatically join their parent's job. Be able to name a second limit a job can enforce, such as a memory or active-process cap.

for a senior

Show the failure modes you have hit: the assign-after-create race and the suspended-create fix, process id reuse making tree-walking kills unsafe, and cleanup surviving a supervisor crash because handle closure is what triggers it.

for a principal

Frame it as a containment policy for a shared agent: what every workload runs inside, which limits are mandatory, whether breakaway is ever allowed, and how job accounting and completion-port events feed the platform's capacity and abuse signals.

## Why the obvious approach fails Windows stores the creating process's id in a new process, and tools display it as a tree, but that link is descriptive rather than supervisory. Nothing propagates a termination down it. `TerminateProcess(hProcess, code)` marks one process object for termination; the children keep running, now with a parent id that points at something dead. There is no Unix-style process group that a single kill can address, and no session leader whose exit takes the group with it. The tempting workaround — enumerate processes, match parent ids, kill the matches — has a genuine correctness hole. Process ids are recycled aggressively on Windows. Between the moment you read a parent id and the moment you act on it, the parent may have exited and its id may have been reissued to an unrelated process. Recursive "kill the tree" utilities built this way can and do kill the wrong process on a busy machine. It is also incomplete: a child that deliberately detaches, or one started by a service on the build's behalf, never appears in the tree at all. ## The job object A job object is a kernel object that holds a set of processes and applies limits and accounting to the set as a whole. It is the Windows primitive for exactly this problem, and it is what Windows containers and server silos are built on top of. The lifecycle is small: - `CreateJobObjectW(NULL, NULL)` returns a handle to an anonymous job. Named jobs exist too, which lets a supervisor reattach. - `SetInformationJobObject(job, JobObjectExtendedLimitInformation, &info, sizeof info)` applies limits from a `JOBOBJECT_EXTENDED_LIMIT_INFORMATION`, whose `BasicLimitInformation.LimitFlags` selects which are active. - `AssignProcessToJobObject(job, hProcess)` puts a process in. - `CloseHandle(job)` releases your reference. The flag that solves the build-agent problem is `JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE`: when the last handle to the job closes, the kernel terminates every process still in it. That includes the case that matters most — the agent itself crashing or being killed, because process teardown closes its handles, which closes the job, which reaps the build. Cleanup no longer depends on the supervisor running its own shutdown code. ## Membership is inherited A process created by a member of a job joins that job automatically. That is what makes the guarantee transitive to grandchildren without cooperation from the tools in between. The exception is deliberate: if the job was created with `JOB_OBJECT_LIMIT_BREAKAWAY_OK`, a child created with `CREATE_BREAKAWAY_FROM_JOB` may leave, and `JOB_OBJECT_LIMIT_SILENT_BREAKAWAY_OK` makes children leave without asking. Absent those flags, escape is not available to the child. ## Closing the assignment race Between `CreateProcessW` returning and `AssignProcessToJobObject` executing, the new process is already running and could spawn a child that would not be in the job. The standard pattern removes the window: ```c CreateProcessW(NULL, cmd, NULL, NULL, FALSE, CREATE_SUSPENDED, NULL, NULL, &si, &pi); AssignProcessToJobObject(job, pi.hProcess); ResumeThread(pi.hThread); /* nothing ran before it was a member */ ``` The process is created suspended, assigned, and only then resumed, so its first instruction executes inside the job. ## What else a job buys Once the processes are corralled, the job becomes the natural place to bound them: - `JOB_OBJECT_LIMIT_ACTIVE_PROCESS` caps how many processes may exist at once — a fork-bomb guard for untrusted builds. - `JOB_OBJECT_LIMIT_JOB_MEMORY` and `JOB_OBJECT_LIMIT_PROCESS_MEMORY` cap committed memory for the whole job or per process. - `JOB_OBJECT_LIMIT_JOB_TIME` bounds total user-mode CPU time; `JOB_OBJECT_LIMIT_PRIORITY_CLASS` and `JOB_OBJECT_LIMIT_AFFINITY` constrain scheduling for every member. - CPU rate control through `JobObjectCpuRateControlInformation` throttles the job to a share of the machine, so one build cannot monopolise the agent. - `JobObjectBasicUIRestrictions` blocks members from touching the desktop, clipboard or other windows. - `JobObjectAssociateCompletionPortInformation` delivers notifications — process added, process exited, limit exceeded, active process count zero — to an I/O completion port, which is how a supervisor observes the job rather than polling it. - Accounting information (`JobObjectBasicAccountingInformation`) reports aggregated CPU time, page faults and process counts for the whole job, which is far more meaningful than per-process numbers for a build. Jobs may be nested since Windows 8 and Server 2012, so an agent can wrap a build that itself wraps its test processes in a job; limits compose, with the most restrictive winning. On earlier releases a process could belong to only one job, which is why older agents sometimes failed when run under another job-using supervisor. ## The interview point The answer that lands is not "use taskkill /T". It is that Windows offers a *kernel-enforced container for a set of processes* whose membership children cannot escape by default and whose destruction is tied to a handle, so cleanup survives the supervisor dying — and that this, not the process tree, is the object you should reason about when you need guarantees.

  • Why create the child suspended before assigning it to the job?
    Otherwise there is a window between `CreateProcessW` returning and `AssignProcessToJobObject` running in which the new process is already executing and could spawn a child outside the job. `CREATE_SUSPENDED` plus `ResumeThread` after assignment closes it, so the process's first instruction runs as a member and every descendant inherits membership.
  • Can a process escape the job it was created into?
    Only if the job allows it. With `JOB_OBJECT_LIMIT_BREAKAWAY_OK` a child created with `CREATE_BREAKAWAY_FROM_JOB` leaves, and `JOB_OBJECT_LIMIT_SILENT_BREAKAWAY_OK` makes children leave automatically. Without those flags the request fails, which is what makes the containment a guarantee rather than a convention.
  • How does a supervisor learn that the last process in a job has exited, without polling?
    Associate an I/O completion port with the job using `JobObjectAssociateCompletionPortInformation`. The kernel posts messages for new processes, exited processes, exceeded limits and the active-process-count-zero condition, so the supervisor waits on the port and reacts to real events instead of scanning the process list.
  • Besides cleanup, what would you bound on a shared build agent?
    `JOB_OBJECT_LIMIT_ACTIVE_PROCESS` to survive a runaway build, `JOB_OBJECT_LIMIT_JOB_MEMORY` so one job cannot push the machine into paging, and CPU rate control via `JobObjectCpuRateControlInformation` so a single build cannot monopolise the cores. Job accounting then reports CPU and fault totals for the build as a whole rather than per process.

saying these in an interview costs you the question

  • Assuming terminating a parent kills its children
  • Killing by parent process id despite id reuse
  • Believing a child can always break away from a job
  • Thinking cleanup requires the supervisor to shut down cleanly
  • Confusing a job object with a process priority setting

context