In a systemd service unit, what is the difference between `Restart=on-failure` and `Restart=always`, which terminations count as a failure, and why does neither of them restart the service after an operator runs `systemctl stop`?
answer
- clean exit versus unclean exit
- default is not to restart at all
- a requested stop is not a failure
- SuccessExitStatus= moves the line
- RestartSec= defaults to 100ms
basics
~20 sRestart=on-failure restarts only after an unclean exit — a nonzero exit status, a fatal signal, a timeout or a watchdog trip. Restart=always additionally restarts after a clean exit 0. Neither applies to an operator-issued systemctl stop, which is an explicit job, not a failure.
solid answer
~40 s`Restart=` tells systemd what to do when the service's main process goes away on its own. The default is `Restart=no`. `on-failure` covers *unclean* termination: a nonzero exit status, death by an uncaught signal such as SIGSEGV, a timeout, or a watchdog trip. `always` covers all of that plus a clean `exit 0`, which is what you want for a daemon that should never legitimately end. In between sit `on-abnormal` (signal, timeout or watchdog, but not a plain nonzero exit) and `on-abort`, `on-success`, `on-watchdog`. What counts as success is tunable with `SuccessExitStatus=`. Crucially, none of these fire when *you* stop the unit — `systemctl stop` is an explicit stop job, and systemd distinguishes a requested stop from the process dying underneath it, so the unit simply goes inactive and stays there.
code
ini · 6 lines[Service]
ExecStart=/usr/local/bin/exampled --config /etc/exampled.conf
Restart=on-failure
RestartSec=5s
SuccessExitStatus=0 78
RestartPreventExitStatus=3go deeper
Be able to name Restart=no as the default and say that on-failure covers crashes and nonzero exits while always covers a clean exit too.
Explain each value in terms of what termination it reacts to, mention SuccessExitStatus= and RestartPreventExitStatus=, and state plainly that an operator stop bypasses the policy.
Argue about pacing: what RestartSec= should be for a service with external dependencies, when an automatic restart hides a defect, and when failing loudly beats recovering silently.
Own where recovery responsibility lives across the stack — host supervisor, container runtime, orchestrator — and set fleet defaults so a crash-looping service is visible in alerting rather than absorbed at every layer.
## What Restart= actually governs `Restart=` is a `[Service]` directive that decides whether systemd re-runs `ExecStart=` when the service's main process exits. It is the whole reason a systemd unit is a supervisor rather than a launcher. The default is `Restart=no`: the process dies, the unit becomes `inactive` or `failed`, and nothing else happens. ```ini [Service] ExecStart=/usr/local/bin/exampled Restart=on-failure RestartSec=5s ``` ## The value ladder systemd accepts seven values, and the interview usually turns on the middle two: - `no` — never restart. The default. - `on-success` — restart only when the process exited *cleanly*. - `on-failure` — restart on a nonzero exit status, termination by an uncaught signal, a start/stop timeout, or a watchdog timeout. This is the everyday choice. - `on-abnormal` — restart on a signal, a timeout or a watchdog, but **not** on a plain nonzero exit status. Useful when the program uses exit codes deliberately to say "stop trying". - `on-abort` — restart only on an uncaught signal. - `on-watchdog` — restart only when the watchdog timeout fires. - `always` — restart in every case, clean exit included. ## What "failure" means, and how to redefine it systemd classifies an exit as clean when the status is 0 (and, for a few well-known cases, `SIGHUP`/`SIGINT`/`SIGTERM`/`SIGPIPE` count as clean terminations). Anything else is a failure, which also puts the unit into the `failed` state visible in `systemctl status` and `systemctl --failed`. Two directives let you redraw that line without changing your program: ```ini SuccessExitStatus=1 250 SIGUSR1 RestartPreventExitStatus=78 RestartForceExitStatus=99 ``` `SuccessExitStatus=` declares extra exit codes and signals to be treated as success. `RestartPreventExitStatus=` names codes after which systemd will *not* restart even under `Restart=always` — the standard way for a program to say "my configuration is broken, do not loop on me". `RestartForceExitStatus=` is its mirror. ## Why an operator stop is not a failure This is the part candidates get wrong. When you run `systemctl stop foo.service`, systemd enqueues a *stop job*. It knows the termination it is about to observe is the one it asked for, so the restart logic is skipped entirely and the unit settles at `inactive (dead)`. The same is true for `systemctl restart`, for a stop triggered by a conflicting unit, and for a shutdown. `Restart=always` is a policy about the process dying *unexpectedly*, not a promise that the unit can never be stopped. The practical consequence is reassuring: you can stop a `Restart=always` service by hand during an incident and it will stay stopped. If a service really does come back after you stopped it, the cause is elsewhere — something else activated it, not the restart policy. ## Pacing the restarts `RestartSec=` is the delay before the restart attempt; the default is a very short `100ms`. That default is aggressive: a service that crashes instantly will be respawned ten times a second until systemd's start rate limiter intervenes. For anything that depends on an external resource — a database that is still coming up, a mount that is not ready — a few seconds is far kinder to the machine and to your journal. On systemd 254 and newer you can get real backoff instead of a fixed delay by adding `RestartSteps=` and `RestartMaxDelaySec=`, which make the interval grow geometrically from `RestartSec=` up to the maximum. On older versions the delay is a single fixed value. ## Choosing a value in practice For a long-running network daemon, `Restart=always` is usually right: there is no clean exit that you want to honour, and a process that returned 0 unexpectedly is as broken as one that segfaulted. For a `Type=oneshot` job, or a program whose nonzero exits carry meaning ("nothing to do", "configuration rejected"), `on-failure` combined with `RestartPreventExitStatus=` expresses the intent much better. For a unit whose failure should page a human rather than be papered over, `Restart=no` plus alerting on the `failed` state is a legitimate and underused choice — an automatic restart that always succeeds is also an automatic way to never find out why it keeps crashing.
- Your program exits 3 to mean "configuration is invalid". How do you stop Restart=always from looping on it forever?Add `RestartPreventExitStatus=3`. systemd will then skip the restart for that specific exit status even under `Restart=always`, leaving the unit in the failed state where an operator or an alert can see it. The alternative is `SuccessExitStatus=3` if you would rather the exit be recorded as a clean stop instead of a failure.
- What does the default RestartSec=100ms mean for a service that crashes on startup?It respawns roughly ten times a second, which floods the journal and burns CPU until systemd's start rate limiter refuses further attempts. A few seconds is a better default for anything with an external dependency, and systemd 254 and newer can grow the delay automatically with RestartSteps= and RestartMaxDelaySec=.
- When is Restart=no the right choice for a production service?When a crash should be investigated rather than absorbed — batch jobs, data-migration units, and anything where an automatic restart could compound corruption. It is also right when a higher-level supervisor already owns recovery. The cost is that you must actually alert on the failed state, otherwise the unit sits dead unnoticed.
saying these in an interview costs you the question
- Thinks Restart=always restarts after systemctl stop
- Believes Restart= defaults to on-failure
- Says exit code 0 triggers a restart under on-failure
- Treats a restart policy as a substitute for fixing the crash
- Confuses on-abnormal with on-failure for nonzero exits