skip to content

A long-lived container's process list fills over days with finished helper entries until new ones fail — why?

level: middleimportance: should knowfreq 42%

answer

  1. finished is not the same as gone
  2. somebody has to read the status
  3. the first process inherits the duty
  4. slots leak until spawning fails
  5. a minimal init reaps by design

basics

~20 s

Nothing is collecting the exit statuses of the children that end inside the container. Each finished child keeps a process-table entry until its status is read, and with the first process not doing that job the entries accumulate until the table is exhausted.

solid answer

~50 s

When a process ends, its entry survives until whoever is its parent collects its exit status — reaping it. Inside a container, anything whose parent has already gone is handed to the **first process**, so position one inherits that duty for the entire container, not just for the processes it started itself. A report renderer or a wrapper script put at position one was never written to do it. So every finished helper leaves a slot behind, the count grows monotonically for days, and eventually nothing new can be started at all. The failure surfaces far from its cause: a render that spawns a converter simply fails to spawn it, while memory and CPU look flat. The fix is a first process that reaps by design — a minimal init — or an application that collects what ends beneath it if it is going to hold position one.

go deeper

for a junior

Hold on to the basic fact: a process that has ended still takes up a slot until someone reads its exit status. Nothing reclaims that slot on its own.

for a middle

Explain the container-specific part: anything orphaned is handed to the first process, so position one owns reaping for the whole container, and an ordinary application there does not do it.

for a senior

Show how you would separate this from its look-alikes in a real incident: flat memory and CPU, a count that never falls even when idle, and a failure that appears wherever something tried to spawn.

for a principal

The interesting question is why this reached production at all. Decide whether reaping is a property of every image by default or a thing each team gets right, and where that is verified.

## Why a process that has finished still occupies a slot When a process ends, the system cannot throw everything about it away at once: whoever started it may still want to know how it went. So an ended process keeps a **process-table entry** holding its exit status until someone collects that status. Collecting it is called **reaping**, and it belongs to the parent. A parent that never asks — because it is busy, because nobody wrote that part, or because it has itself already gone away — leaves the entry standing. It is worth being precise about what such an entry costs, because the answer shapes the whole symptom: - it holds no memory beyond the record itself, runs no code, and consumes no CPU; - it occupies one slot in a **finite table**, and that table is the resource that runs out; - it cannot be cleaned up by pressure of any kind — nothing evicts it, because nothing is wrong with it. ## What a container boundary changes Two things differ inside a container, and together they turn a minor housekeeping duty into a production failure: - **Everything orphaned lands on the first process.** When a parent ends before its children do, those children are handed to the container's first process — so position one inherits the reaping duty for the whole container, not only for what it started directly. - **Whatever is at position one was usually not built for the job.** Outside a container that position is held by a process written precisely to collect orphans forever. Inside, it is held by a report renderer, or a wrapper script, or whatever the image declared to run — none of which were written with this duty in mind. A container that starts a helper for each report, or that shells work out to short-lived converters, therefore leaks one slot per helper, forever, for as long as the container lives. A container restarted every deploy may never show it. One that runs for weeks will. ## The symptom, and the two false leads it invites The shape of the symptom is distinctive once you know to look for it: 1. the count of process entries rises monotonically and never falls, including while the workload is completely idle; 2. memory, CPU and request latency stay flat the whole time, because the entries are not doing anything; 3. eventually something fails to start a process, and the error appears wherever the spawning happens to live — a report that cannot launch its converter, a periodic task that never runs — nowhere near the cause; 4. a restart clears it instantly and the service is healthy again, which is exactly why the defect survives investigation for months. The two leads it sends people down are both wrong. It is not a memory leak: memory use is flat, and an entry holds none. And it is not the workload spawning too much: live processes come and go with load, so their count falls when the workload is idle, while uncollected entries only ever accumulate. That difference — does the number ever go down? — is the cheapest discriminator you have. ## Who collects, in each shape | First process | Who collects finished children | What happens when nobody does | |---|---|---| | A minimal init | It does, by design; that is half of why it exists | Not applicable — this is the shape that does not leak | | The application, written for the position | It collects whatever ends beneath it, including processes it did not start | Entries accumulate exactly as with a wrapper | | A wrapper script that only waits for one child | Only that one child, if even that | Every other finished process leaves a slot behind permanently | ## Fixing it, and what is not a fix Three routes, in the order most estates take them: 1. put a **minimal init** at position one: it forwards stop requests to its children and collects them as they end, which is exactly the pair of duties the position carries; 2. if the application is going to hold position one, have it collect what ends beneath it — including orphans it never started, which is the part teams miss; 3. if a wrapper script must stay, make it do both jobs deliberately, and test it, because it is now signal-handling and reaping code that you own. What is not a fix is the scheduled restart that clears the table. It works, in the sense that the count goes back to zero, and it is why this defect hides so well: the leak is camouflaged by the very routine that masks it. Nor does raising any limit help for long; it changes the date of the outage, not the slope of the line.

  • The application holds position one and starts no children of its own. Does it still need to collect anything?
    Possibly. The duty follows the whole process tree beneath it, not just the processes it starts directly. A library that forks a helper, a background task that spawns a converter, or anything started by something it started can end up orphaned and handed to it. If nothing under it ever forks, there is genuinely nothing to collect — but that is a property to verify, not to assume.
  • How do you tell a slow leak of process entries from a workload that is simply spawning too much?
    Watch whether the number ever falls. Live processes come and go with load, so their count drops when the workload is idle or the burst ends. Uncollected entries never drop; the line only goes up, including overnight with no traffic. A monotonic count through an idle period is the leak.

saying these in an interview costs you the question

  • Thinks a finished process frees its entry with nobody reading its status
  • Assumes only the original parent can ever collect a child's status
  • Blames a memory leak when it is process slots that ran out
  • Treats the nightly restart that clears it as the actual fix
  • Expects any application put at position one to reap by default