A macOS Launch Daemon configured with KeepAlive crashes a second after it starts, every time. What does launchd do about it, and what keeps it from respawning the process in a tight loop?
answer
- restarts, but not as fast as it can
- a floor on time between launches
- ten seconds unless you say otherwise
- the plist key with Throttle in the name
- find the exit status, not a faster knob
basics
~20 slaunchd keeps restarting it, but rate-limits the respawns: a job is not started more often than once per ThrottleInterval, ten seconds by default. So a crash loop shows up as a process reappearing roughly every ten seconds, not as a spinning CPU.
solid answer
~50 slaunchd honours `KeepAlive` and relaunches the job, but it will not respawn a job more often than once every `ThrottleInterval` seconds — the default is 10. That throttle is what stops a broken daemon from becoming a fork bomb, and the symptom is characteristic: the process reappears with a new PID on a steady ten-second cadence and the system log notes that the service ran for less than its throttle interval, so the respawn is being deferred. Diagnosis is to find out *why* it dies, not to shorten the interval: check the job's last exit status with `launchctl print` on its service target, and add `StandardOutPath` and `StandardErrorPath` to the plist so the program's own output lands somewhere readable, because a daemon that crashes during startup usually never gets far enough to log anything itself. Typical causes are a missing binary or wrong path, permissions on a data directory, a dependency that is not up yet, or the program daemonising itself.
go deeper
Know that KeepAlive means launchd restarts the job after it exits, and that launchd deliberately spaces those restarts out instead of relaunching instantly.
Explain the throttle concretely — a minimum interval between launches of the same job, ten seconds by default, set by ThrottleInterval — and why the fix belongs in the program, not the interval.
Show how you diagnose it: read the last exit status from launchd's own record, add stdout and stderr paths to capture startup output, and reproduce by running the exact arguments as the job's user.
Speak to the pattern across a fleet: crash-looping helpers are a real availability and battery cost, so define how services declare dependencies, what health signals are collected, and when a repeatedly failing helper should be disabled rather than left looping.
## Why the throttle exists `KeepAlive` makes launchd a supervisor: exit and it starts you again. Taken literally, a program that fails instantly would be restarted thousands of times a second, burning a CPU core and filling the log. So launchd applies a minimum spacing between launches of the same job. The `ThrottleInterval` key sets it, and the default is 10 seconds. If a job runs for less than that interval, the next start is deferred until the interval has elapsed since the *previous* start. This is the same idea as a supervisor's restart backoff on other platforms, with one difference worth saying out loud: it is a flat minimum spacing, not an escalating backoff, and it does not give up. A permanently broken job keeps reappearing every ten seconds forever. ## What you actually see The fingerprint of this failure is a process that exists, then does not, then exists again with a different PID, on a suspiciously round cadence. People often misread that as "the service is running" because a snapshot happens to catch a live process. The system log carries launchd's own note that the service only ran for a short time and its respawn is being pushed out, which is the clearest confirmation that you are looking at a crash loop rather than a healthy daemon. ## Getting the real cause The throttle is a symptom-limiter, not information. Three moves get you the cause: 1. **Ask launchd what happened.** `launchctl print system/com.example.job` reports the job's state and its last exit status or terminating signal. A last exit status of 127 or a spawn failure points at a bad `Program`/`ProgramArguments` path; a signal points at a genuine crash. 2. **Capture the program's own output.** A daemon has no terminal, so anything it prints is discarded unless you ask for it. Adding `StandardOutPath` and `StandardErrorPath` to the plist redirects both streams to files, and for startup failures that is usually where the answer is — a config parse error, a missing certificate, an unwritable directory. 3. **Run it by hand as the same user.** Execute exactly the `ProgramArguments` under `sudo -u` with the job's `UserName`. Many crash loops are environment, not code: a daemon does not inherit a login shell's `PATH`, has no `$HOME` you would recognise, and starts in `/` unless `WorkingDirectory` says otherwise. ## The causes that recur - **Self-daemonising.** The program forks and the parent exits. launchd watches the parent, sees a fast exit, and restarts — a crash loop made entirely of successful startups. Stay in the foreground. - **Ordering assumptions.** A daemon that needs the network, a mounted volume or another service and exits when it is missing will loop at boot until that thing appears. Either wait and retry inside the program, or express the dependency — `KeepAlive` with `PathState`, or on-demand launch through a socket — rather than assuming boot order. - **Permissions.** A daemon dropped to a service account via `UserName` cannot write the directory the developer tested as root. - **A quarantined or unsigned binary.** Gatekeeper and code-signing policy can refuse the exec, which reads as a spawn failure rather than a crash. ## What not to do Do not "fix" this by setting `ThrottleInterval` to 0 or a tiny value: that removes the only thing protecting the machine, and a fast crash loop can starve the box and flood the log. Lowering it is defensible only for a job whose short lifetime is legitimate — a helper that genuinely runs for a second and exits by design — and even then you are telling launchd its safety net is unnecessary. Equally, do not paper over it by dropping `KeepAlive`; then the job simply stays dead after the first failure and you have converted a loud failure into a silent one. ## The judgment being tested Interviewers ask this to see whether you treat the platform's supervisor as something you understand rather than fight. The strong answer names the throttle and its default, reads the log line as evidence, and then goes looking for the exit status — rather than reaching for the knob that makes the restarts faster.
- Someone sets ThrottleInterval to 1 to make a flapping daemon recover faster. What is wrong with that?It removes the protection rather than the fault. A job crashing on startup now respawns ten times as often, burning CPU and flooding the log on a machine that is already unhealthy. Lowering the interval is only reasonable for a job whose sub-second lifetime is by design; a crash loop needs the exit status investigated instead.
- How would you tell a crash loop apart from a job that simply refuses to start at all?A crash loop shows a live process with a changing PID on a regular cadence and launchd log entries about a short-lived service being deferred. A job that never starts has no process and its state in launchctl print shows a spawn failure or that it is disabled — different fix entirely.
- The daemon logs nothing anywhere. Why, and what do you change?It has no terminal, so stdout and stderr go nowhere by default and an early startup failure prints into the void. Add StandardOutPath and StandardErrorPath to the plist so both streams land in files you can read, then reproduce. Later, have the program log through the unified logging system rather than relying on redirects.
saying these in an interview costs you the question
- Says launchd gives up after a few failed restarts
- Thinks a crash loop pins a CPU because restarts are instant
- Lowers ThrottleInterval instead of finding the crash
- Assumes a live process means the service is healthy
- Expects a daemon's stdout to appear somewhere by default