skip to content

Classic Unix daemons double-forked and wrote a PID file, while a modern service manager supervises the process directly as its own child. What concrete failure modes does the PID-file model have that direct supervision removes?

level: seniorimportance: should knowfreq 48%

answer

  1. the daemon reports on itself
  2. orphaned means the manager is not the parent
  3. a number in a file can be recycled
  4. one PID cannot describe a process tree
  5. cgroup membership is a kernel fact

basics

~20 s

PID files are self-reported, racy and can go stale: the manager may read the file before it is written, a leftover file can name a recycled PID belonging to an unrelated process, forked children are invisible, and because the manager is not the parent it never learns the exit status. Direct supervision fixes all four.

solid answer

~50 s

The PID-file model asks the daemon to report on itself, and every failure follows from that. The daemon double-forks to detach, so the manager's child exits immediately and the real process is reparented away — the manager is no longer the parent and can never call `wait()` to learn how it died. The PID file is written by the daemon at some unspecified later moment, so a start script can read it too early or read a leftover from a previous run. Worse, PIDs are recycled: a stale file can name a completely unrelated process, and a naive stop path will signal it. And only one PID is recorded, so any children the daemon forked survive a stop. A modern manager keeps the service in the foreground as its own direct child, so it receives `SIGCHLD` and the real exit status, and it places every process in a per-service cgroup so all descendants are tracked and can be signalled regardless of how the daemon forked.

go deeper

for a junior

Know what a PID file is and where it lives, and be able to say that a leftover one can make a start script wrongly believe the service is already running.

for a middle

Explain the mechanics of double-forking and why it costs you the parent relationship, so the manager can only poll for existence instead of being told the exit status.

for a senior

Bring the production judgment: enumerate the concrete failures you have seen — stale files, recycled PIDs signalled by restart scripts, workers surviving a stop and blocking the port — and explain how foreground execution plus cgroup tracking removes each one.

for a principal

Frame it as a contract between software and platform: mandate foreground-running, explicit readiness signalling and supervisor-captured output as the standard for services your organisation ships, and treat forking daemons as a compatibility exception with a cost.

## What double-forking was actually for On a classic Unix system, a program started from a shell inherits that shell's session and controlling terminal. If the terminal goes away, the process group gets `SIGHUP`. A long-running daemon must escape that, so the traditional daemonisation dance is: `fork()`, let the parent exit so the shell's prompt returns; call `setsid()` in the child to become a session leader with no controlling terminal; `fork()` a second time so the final process is not a session leader and can therefore never acquire a terminal; then `chdir("/")`, reset the umask, and close or redirect the standard file descriptors. The surviving process is orphaned and reparented to PID 1. This is a perfectly sensible answer to the terminal problem. The trouble is what it does to *management*. ## The PID-file contract Because the daemon has run away from whoever launched it, some other channel is needed to identify it. The convention became a file, typically under `/run` (historically `/var/run`), containing the decimal PID and a newline. The init script writes or reads it; `stop` means read the number and send a signal to it. Notice what this contract is: the process reports its own identity, at a time of its own choosing, into a mutable file, and everyone else trusts it. ## Failure mode 1 — the start race The daemon writes the PID file at some point after it starts; nothing defines when. A start script that returns immediately may report success before the service is listening, and a health check that reads the file may find it absent or half-written. There is also no way to distinguish "still starting" from "failed during startup", because the only signal is the presence of a file. ## Failure mode 2 — staleness and PID reuse If the daemon is killed with `SIGKILL`, or the machine loses power, the file survives with a dead PID in it. Two bad outcomes follow. The benign one: the start path sees the file, concludes the service is already running, and refuses to start — the service stays down while the tooling insists it is up. The dangerous one: PIDs are recycled, so by the time anyone reads that file the number may belong to a completely unrelated process. A stop path that blindly signals it kills the wrong thing. Careful scripts try to mitigate this by comparing `/proc/<pid>/comm` or the cmdline against an expected name, but that is a heuristic bolted onto a broken premise, and it is still racy. ## Failure mode 3 — the forgotten children Only one PID is recorded. A daemon that forks worker processes, or spawns helpers, leaves those entirely outside the model. Stopping the service signals the recorded PID and leaves the workers running — holding ports, holding file locks, still writing to the database. The next start then fails with an address-already-in-use error whose cause is invisible from the PID file alone. ## Failure mode 4 — no exit status This is the deepest one. In Unix, only a process's parent can reap it and learn its exit status via `wait()`. Because the daemon deliberately made itself an orphan, the manager is not its parent, so it cannot be told when the service dies, cannot learn whether it exited cleanly or was killed, and cannot see which signal killed it. All it can do is poll: check periodically whether the PID still exists. Polling is late by construction and blind to the reason. ## How direct supervision fixes each one A modern service manager inverts the arrangement. The service is asked *not* to daemonise — it stays in the foreground, and the manager keeps it as a direct child. - Being the parent, the manager receives `SIGCHLD` the instant the process dies and can `wait()` for the exact status, distinguishing exit code 1 from death by `SIGKILL`. Restart policies become meaningful because the trigger is exact and immediate. - Every process of the service is placed in a per-service cgroup, so membership is a kernel fact rather than a self-report. Stopping the service means signalling the whole cgroup, which catches forked children no matter how they were spawned, and no PID-reuse mistake is possible because the cgroup, not a recycled number, defines membership. - Readiness gets a real channel. Under systemd, a service can call `sd_notify(3)` to send `READY=1` when it is genuinely serving, so "started" means started rather than "the file appeared". Socket activation goes further: the manager holds the listening socket, so it is ready before the service is. - Standard output and error are captured by the manager instead of the daemon inventing its own log handling. ## Where PID files still appear Compatibility remains: systemd supports `Type=forking` together with `PIDFile=`, for software that genuinely cannot run in the foreground. It is explicitly the fallback — the manager has to guess which process is the main one, and everything above still applies to that guess. Cgroup tracking still catches the strays, which is why even a badly behaved legacy daemon is more manageable under systemd than under an rc script. But when you control the software, running in the foreground and letting the supervisor be the parent is strictly better. ## Interview framing This question is a good discriminator because both models are simple to describe and the difference only becomes vivid if you have been burned. Concrete war stories — a stale PID file that made a restart script kill an unrelated process, or worker processes surviving a stop and blocking the next start — are worth more than a taxonomy.

  • Why can a stale PID file be actively dangerous rather than merely useless?
    Because PIDs are recycled. Once the original process is gone the kernel is free to hand that number to anything else, so a leftover file may name an entirely unrelated process. A stop or restart path that reads the file and signals the number will then kill the wrong thing, and the failure looks unrelated to the service that caused it.
  • What does a per-service cgroup give a supervisor that a PID file cannot?
    Membership becomes a kernel-maintained fact rather than a self-report. Every process the service spawns, however it forks or re-parents, remains in the cgroup, so the supervisor can signal the whole set on stop and account resources for all of it. Nothing escapes by daemonising, and there is no recycled-number ambiguity to guess about.
  • If a service genuinely cannot run in the foreground, what do you lose and how do you limit the damage?
    You lose precise exit status and unambiguous main-process identification, because the supervisor has to guess which forked process is the real one. Configure the forking type together with the PID file the daemon actually writes, keep that file under a root-owned runtime directory, and rely on cgroup-wide stop to catch strays. Treat it as a compatibility path, not a design choice.
  • Why does readiness deserve a separate signal rather than being inferred from the process existing?
    A process exists long before it can serve: it may still be reading configuration, opening a database connection, or warming caches. Treating start as readiness makes dependent services race against it. An explicit readiness notification, or having the supervisor hold the listening socket, turns readiness into something the service asserts rather than something the manager guesses from a file appearing.

saying these in an interview costs you the question

  • Thinks a PID file is authoritative because the daemon wrote it
  • Says PIDs are never reused so staleness is harmless
  • Assumes stopping the main PID stops all workers
  • Believes the manager can read the exit code of an orphaned process
  • Treats double-forking as still the modern best practice

context