skip to content

In a macOS launchd job plist, what do the RunAtLoad and KeepAlive keys each control, and how do they differ?

level: middleimportance: must knowfreq 68%

answer

  1. one is about starting, one about restarting
  2. loading a job is not running it
  3. both default to false
  4. exit status can be a restart condition
  5. never let a supervised job daemonise itself

basics

~20 s

RunAtLoad is a start trigger: launchd runs the job once, immediately, when the job is loaded. KeepAlive is a restart policy: launchd keeps the job running and starts it again whenever it exits. Both default to false.

solid answer

~50 s

They answer two different questions. `RunAtLoad` says *when to start* — with it set to true, launchd launches the program once at load time, which for a daemon means at boot. `KeepAlive` says *what to do when it exits* — with the boolean true, launchd restarts the program every time it stops, whatever the exit status, so the job is effectively a supervised long-running service. `KeepAlive` also implies an initial start, so a job with only `KeepAlive` still comes up at load. Neither key is on by default: a plist with no trigger at all loads fine and sits idle until something starts it — a socket connection, a Mach message, a timer key, or `launchctl kickstart`. `KeepAlive` can also be a dictionary rather than a boolean, so you can say "restart only if it exited non-zero" or "only while this path exists".

code

xml · 20 lines
xml
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
    <key>Label</key>
    <string>com.example.sync</string>
    <key>ProgramArguments</key>
    <array>
        <string>/usr/local/libexec/example-sync</string>
        <string>--foreground</string>
    </array>
    <key>KeepAlive</key>
    <dict>
        <key>SuccessfulExit</key>
        <false/>
    </dict>
    <key>StandardErrorPath</key>
    <string>/var/log/example-sync.err</string>
</dict>
</plist>

go deeper

for a junior

Know that RunAtLoad starts the program once at load time and KeepAlive keeps it running by restarting it after it exits, and that both are off by default.

for a middle

Explain trigger versus policy, that a load only registers the job, and that KeepAlive can be a dictionary keyed on conditions such as SuccessfulExit rather than a plain boolean.

for a senior

Demonstrate the field consequence: a supervised job must stay in the foreground because self-daemonising reads as a death, and KeepAlive on an on-demand socket job destroys the on-demand benefit.

for a principal

Own the service-lifecycle policy across a fleet — which helpers are resident versus on-demand, what restart conditions are appropriate, and how those choices show up as idle memory and battery cost on user machines.

## Two orthogonal decisions Every launchd job answers two independent questions: what causes it to start, and what happens after it stops. `RunAtLoad` belongs to the first, `KeepAlive` to the second, and confusing them is the most common launchd mistake there is. Loading a job is not running it. `launchctl bootstrap` (or the legacy `load`) registers the job with launchd in a domain; launchd then honours whatever triggers the plist declares. If the plist declares none, the job is known but idle — which surprises people who expect a load to be a start. ## RunAtLoad ```xml <key>RunAtLoad</key><true/> ``` This means "run the program once, right now, at load time". For a Launch Daemon, load time is boot; for a Launch Agent, it is login. It is a one-shot: launchd does not run it again when it exits. `RunAtLoad` is the right key for a job that does a piece of work and terminates — a boot-time migration, a cache warm-up, a one-off registration. Note it is emphatically *not* a guarantee of ordering. launchd has no runlevels and no "start me after that other job"; jobs come up concurrently and you are expected to handle dependencies by demand — by connecting to a socket or Mach service that launchd will start on your behalf — rather than by sequencing. ## KeepAlive as a boolean ```xml <key>KeepAlive</key><true/> ``` This makes launchd the supervisor: the program is started, and every time it exits — cleanly, with an error, or on a signal — launchd starts it again. This is the supervision that makes launchd an init system rather than a launcher, and it is why a macOS daemon does not need its own restart wrapper script, watchdog or double-fork. In fact a launchd job must *not* daemonise itself: forking into the background and letting the parent exit looks to launchd exactly like the job dying, so it gets restarted in a loop. Stay in the foreground and let launchd own the lifecycle. Because `KeepAlive` implies the job should be running at all times, launchd starts it at load whether or not `RunAtLoad` is present. The two keys together are harmless but redundant. ## KeepAlive as a dictionary The more interesting form takes conditions rather than a flag: - `SuccessfulExit` — true means restart only after a zero exit; false means restart only after a non-zero exit, which is the "restart on failure, stay down on success" policy people usually want. - `Crashed` — restart when the program died from a signal typically associated with a crash. - `PathState` — a dictionary of paths to booleans: keep the job alive only while a given path does (or does not) exist. Useful for jobs tied to a mounted volume or a lock file. With a dictionary, launchd keeps the job running while the stated conditions hold, so it is genuinely a policy expression and not just an on switch. ## The interaction that bites An on-demand job — one launchd starts when a connection arrives on a socket it holds — is supposed to exit when it is idle so the machine reclaims the memory. Adding `KeepAlive` to such a job defeats the entire design: launchd immediately restarts it, and it is now a resident process that merely happens to have a socket. Pick one model. Similarly, `RunAtLoad` on a job whose real trigger is a timer just gives you one extra run at boot, which may or may not be what you meant. ## What an interviewer is checking They want to hear the distinction stated cleanly — trigger versus policy — plus the two practical consequences: that loading is not starting, and that a KeepAlive job must not fork into the background. Those two facts explain most of the launchd bugs people actually ship.

  • A job with KeepAlive set to true keeps being restarted even though the program appears to start fine. What is the usual cause?
    The program daemonises itself — it forks and the parent returns. launchd only watches the process it started, so the parent's exit reads as the job dying and it is relaunched. Run in the foreground and delete any double-fork, setsid or background-and-exit logic; launchd is the supervisor.
  • How do you express "restart this job only if it failed, but leave it alone when it exits successfully"?
    Use the dictionary form of KeepAlive with SuccessfulExit set to false, which tells launchd to restart only in the inverse condition — a non-zero exit. A plain boolean true cannot express that, because it restarts on every exit regardless of status.
  • Does adding RunAtLoad to a job that already has KeepAlive change anything?
    No, in practice. KeepAlive already means the job should be running, so launchd starts it at load. The pair is redundant rather than wrong. The combination people should question is KeepAlive on an on-demand socket job, because that turns a job meant to exit when idle into a permanently resident process.

saying these in an interview costs you the question

  • Thinks RunAtLoad restarts the job whenever it exits
  • Believes loading a plist always starts the program
  • Says KeepAlive only restarts on a non-zero exit
  • Has the daemon fork into the background itself
  • Adds KeepAlive to an on-demand socket-activated job

context