A systemd service configured with Restart=always is sitting in the failed state, and the journal shows "Start request repeated too quickly" followed by a start-limit-hit result. What has systemd decided, which settings control it, and how do you get the unit running again?
answer
- the supervisor gave up on purpose
- a burst counted inside a window
- five attempts in ten seconds
- the counter has to be cleared by hand
- a slower loop never trips it at all
basics
~20 ssystemd's start rate limiter has tripped: the unit was started more than StartLimitBurst= times (5 by default) within StartLimitIntervalSec= (10s by default), so systemd stopped honouring the restart policy and marked the unit failed. Clear the counter with systemctl reset-failed, then fix the crash.
solid answer
~40 s`Restart=always` did its job — the process kept dying and systemd kept respawning it. To stop an infinite respawn loop, systemd applies a start rate limit: if a unit is started more than `StartLimitBurst=` times (default 5) inside `StartLimitIntervalSec=` (default 10s), it refuses further attempts, logs "Start request repeated too quickly", records the result as `start-limit-hit`, and leaves the unit failed. Until the counter is cleared, even a manual `systemctl start` is refused, so the recovery command is `systemctl reset-failed <unit>` followed by `systemctl start`. That is only the cleanup, though — the limit is a symptom. The real work is `journalctl -u <unit> -b` and the `Process: ... (code=exited, status=...)` line in `systemctl status` to find why it dies, and raising `RestartSec=` so a genuinely slow dependency does not burn the budget.
code
bash · 5 linessystemctl status exampled.service
journalctl -u exampled.service -b --since '30 min ago'
systemctl show -p NRestarts -p ActiveEnterTimestamp exampled.service
systemctl reset-failed exampled.service
systemctl start exampled.servicego deeper
Recognise the phrase "start request repeated too quickly" as systemd refusing further restarts, and know that systemctl reset-failed is what clears it.
Explain the sliding window formed by StartLimitIntervalSec= and StartLimitBurst=, where they live in the unit file, and how RestartSec= interacts with the budget.
Reason about the loop you cannot see: a slow crash loop that never trips the limiter, what NRestarts and the Active-since timestamp reveal, and how to read the exit code before touching the limits.
Set the fleet policy on restart budgets and StartLimitAction=, and make crash-looping units observable so that automatic recovery never becomes a way of not learning about a defect.
## What the message means The journal sequence for this failure is distinctive: ``` exampled.service: Main process exited, code=exited, status=1/FAILURE exampled.service: Failed with result 'exit-code'. exampled.service: Scheduled restart job, restart counter is at 5. exampled.service: Start request repeated too quickly. exampled.service: Failed with result 'start-limit-hit'. Failed to start Example API. ``` systemd is not saying your service failed to start — it is saying it refuses to try again. The restart policy and the rate limiter are two separate mechanisms, and the second one overrides the first. ## The two directives that define the limit The limit lives in the `[Unit]` section, not `[Service]`, because it applies to any unit type: ```ini [Unit] StartLimitIntervalSec=30s StartLimitBurst=3 ``` The rule is a sliding window: more than `StartLimitBurst=` start attempts inside `StartLimitIntervalSec=` and the unit is refused. The system-wide defaults are 5 attempts in 10 seconds, configurable as `DefaultStartLimitIntervalSec=` and `DefaultStartLimitBurst=` in `/etc/systemd/system.conf`. Setting `StartLimitIntervalSec=0` disables the limiter for that unit entirely — occasionally right for a service whose upstream dependency is genuinely flaky, and usually the wrong instinct, because the limiter is the thing that saves the machine from a service respawning ten times a second. There is a third directive worth knowing: `StartLimitAction=`, whose default is `none` but which can be set to `reboot`, `reboot-force`, `poweroff` and similar. On appliance-style systems, "if this service cannot be kept alive, reboot the box" is a real policy. ## Why the default gets hit so easily `RestartSec=` defaults to `100ms`. A service that crashes on startup therefore consumes the entire default budget of five attempts in about half a second. That is by design — the limiter exists to catch exactly this — but it means the default configuration gives a dependency almost no time to become available. A service that starts before its database is reachable will burn its budget and land in `failed` while the database is still coming up. Raising `RestartSec=` to 5 or 10 seconds, or widening the window, is the usual fix for that specific shape of problem. ## Recovering the unit Because the counter persists, a plain `systemctl start` after the limit is hit is refused with the same message. The sequence is: ```bash systemctl reset-failed exampled.service systemctl start exampled.service ``` `reset-failed` clears both the failed state and the rate-limit counter. If you have edited the unit file to change the limits, `systemctl daemon-reload` first. Note that a `systemctl restart` of a unit that has hit the limit does not implicitly reset it on all versions — reaching for `reset-failed` explicitly is the habit worth building. ## The judgment the question is really testing Hitting the start limit is good news, not bad news. The alternative — a service with `Restart=always` and a short `RestartSec=` that crashes every few minutes rather than every few hundred milliseconds — never trips the limiter at all. It reports `active (running)` almost every time anyone looks, its uptime silently resets, and the outage shows up only as intermittent errors on the client side. That is the trap in the leaf: a restart policy is a recovery mechanism, and recovery without visibility is how a crashing process stays crashing for weeks. The things to check for that hidden case: ```bash systemctl show -p NRestarts exampled.service systemctl status exampled.service # 'Active: active (running) since ...' — how recent? journalctl -u exampled.service -b --since '1 hour ago' ``` `NRestarts` is a counter systemd keeps per unit, and a value in the hundreds on a service nobody has touched is the whole diagnosis. The `Active: ... since` timestamp being minutes old on a service that has been deployed for a week says the same thing. ## Fixing the underlying crash The rate limit tells you nothing about the cause, so go to the evidence: the `Process:` line in `systemctl status` carries the exit code and whether the process died by signal (`code=killed, signal=KILL` points at the kernel's out-of-memory killer or an external kill; `code=exited, status=1` points at the program's own error path). Then read the unit's log for the last attempt. Once the cause is understood, decide deliberately whether the right restart policy is still `always`, or whether this unit should fail loudly and page someone.
- A service crash-loops every three minutes instead of every 100 ms. Why is that more dangerous operationally?It never trips the start rate limiter, so the unit is almost always reported as active (running) and nothing lands in the failed state for anyone to alert on. The evidence is indirect: `systemctl show -p NRestarts` climbing, and an `Active: ... since` timestamp that is always only minutes old. Slow crash loops routinely survive for weeks.
- When is it defensible to set StartLimitIntervalSec=0?When the service depends on something genuinely and repeatedly slow to appear — a network mount, a remote database during a maintenance window — and you would rather it keep retrying than land in failed. Pair it with a generous RestartSec= so the retries are paced, otherwise you have removed the only guard against a process respawning many times a second.
- How do you tell from systemctl status whether the process exited on its own or was killed?The Process: line spells it out. `code=exited, status=1/FAILURE` means the program returned that exit status itself, so look at its own logs. `code=killed, signal=KILL` means something terminated it — typically the kernel out-of-memory killer or a stop escalation — and the next step is the kernel log rather than the application log.
saying these in an interview costs you the question
- Thinks the service failed to start rather than being refused
- Says systemctl start alone will recover a rate-limited unit
- Looks for StartLimitBurst in the [Service] section
- Treats raising the limit as the fix for a crashing process
- Assumes any crash loop will show up as a failed unit