skip to content

On a slow-booting systemd host, systemd-analyze blame reports that a unit took 25 seconds, but disabling it does not shorten the boot at all. Why can blame mislead, and which systemd-analyze subcommand answers "what is actually delaying the boot"?

level: seniorimportance: nice to knowfreq 40%

answer

  1. parallel start breaks a flat ranking
  2. slowest is not always blocking
  3. waiting time counts as duration
  4. the tree respects the ordering graph
  5. @ is when, + is how long

basics

~20 s

blame ranks units by how long each took to initialise, ignoring the dependency graph, so a slow unit nothing waits for costs nothing. systemd-analyze critical-chain walks the ordering chain into the boot target and shows which delays actually accumulated.

solid answer

~50 s

`systemd-analyze blame` is a flat, sorted list of per-unit initialisation times with no notion of the graph. Because systemd starts everything it can in parallel, a unit can burn 25 seconds off to one side while the boot target is reached without ever waiting for it — removing it changes nothing. It is also misleading in the other direction: a unit can appear slow simply because it was blocked waiting for something else. The subcommand that respects the graph is `systemd-analyze critical-chain`, which prints the time-critical chain of ordered units into the default target, with `@` marking when each unit became active and `+` how long it took to start. Plain `systemd-analyze` gives the firmware/loader/kernel/initrd/userspace split first, which tells you whether to look at units at all, and `systemd-analyze plot` renders the whole parallel timeline as an SVG when the chain does not explain it.

code

bash · 5 lines
bash
systemd-analyze
systemd-analyze blame | head -20
systemd-analyze critical-chain
systemd-analyze critical-chain sshd.service
systemd-analyze plot > boot.svg

go deeper

for a junior

Know that systemd-analyze exists, that plain systemd-analyze prints the phase breakdown, and that blame lists units by how long they took. Do not present the top of blame as the cause.

for a middle

Explain why a flat ranking cannot answer the question when units start in parallel, and read a critical-chain tree correctly — @ as the moment the unit went active, + as its own start duration.

for a senior

Demonstrate the whole path: split the boot by phase first, use critical-chain to find the real chain, then journalctl -u for that boot to explain the unit. Note that the fix is usually an ordering change, not a faster unit.

for a principal

Decide whether boot time is worth engineering at all — it matters for autoscaling and rolling reboots and rarely for a static fleet. Push the durable fix upstream into the image, and treat boot-time work that belongs on a timer as a pattern to eliminate fleet-wide.

## Start by splitting the boot Before blaming any unit, find out which phase the time is in: ```bash $ systemd-analyze Startup finished in 4.1s (firmware) + 2.3s (loader) + 3.8s (kernel) + 1.9s (initrd) + 12.4s (userspace) = 24.6s ``` If the userspace number is small, no amount of unit tuning will help — the time is in firmware (server BIOS memory training and option ROMs routinely cost tens of seconds), the boot loader menu timeout, or the kernel and initrd. Candidates who jump straight to `blame` skip this and end up optimising the wrong phase. ## What blame measures — and does not ```bash $ systemd-analyze blame 25.104s cloud-final.service 8.221s dnf-makecache.service 3.019s NetworkManager-wait-online.service ... ``` This is a per-unit initialisation duration, sorted descending. Two structural blind spots: 1. **No graph.** Units start in parallel. If nothing in the chain to `default.target` is ordered after `cloud-final.service`, its 25 seconds overlap with everything else and contribute nothing to the total. Disabling it changes the list and not the boot. 2. **Blocking is counted as work.** A unit that spends its time waiting for a dependency still shows the wall-clock duration. So the slowest entry is sometimes a *victim* of the real culprit rather than the culprit itself. blame is still useful — it is the right tool for "which unit is doing something expensive" — but it does not answer "what made the boot long". ## critical-chain: the graph-aware view ```bash $ systemd-analyze critical-chain multi-user.target @14.882s └─sshd.service @14.101s +780ms └─network-online.target @14.098s └─NetworkManager-wait-online.service @11.077s +3.019s └─NetworkManager.service @10.462s +603ms └─dbus.service @10.455s └─basic.target @10.301s ``` Read the two markers carefully, because they mean different things: - `@` is the point in time at which the unit became active — an absolute offset from boot. - `+` is how long that unit itself took to start. A chain where the `@` values jump but the `+` values are small means the delay is *between* units — usually something waiting on a target. A single large `+` is a unit that is genuinely slow. In the example above, three of the fifteen seconds are `NetworkManager-wait-online.service` doing its job; the rest accumulated earlier in the chain. You can also aim it at a specific unit — `systemd-analyze critical-chain sshd.service` — to see the chain that gated *that* service rather than the default target. Its own caveat, which the man page states: the output can still mislead, because socket activation and parallel execution mean the displayed chain is not necessarily the whole story of what a unit was waiting for. ## plot, when the chain is not enough ```bash systemd-analyze plot > boot.svg ``` This renders every unit as a bar on a shared timeline, so overlap is visible directly. It is the tool for "three things all finish at 11s — what were they waiting for?", and it makes serialisation obvious in a way a text tree does not. ## Then ask why, not just what Once you have a real culprit, the analysis tools stop and ordinary triage begins: read the unit's logs for that boot with `journalctl -u <unit> -b`, check what it is ordered after with `systemctl list-dependencies --after <unit>`, and confirm the unit file is what you think it is with `systemctl cat`. The usual findings are dull and repeatable — a wait-online service blocking on an interface that will never come up, a mount ordered against an unreachable NFS server, a `Type=oneshot` unit doing network I/O inside `ExecStart`, or a metadata refresh running at boot that belongs on a timer. ## The fixes that actually shorten a boot Most of the time you are removing an ordering edge rather than making a unit faster: move work off the critical chain by not gating other units on it, convert a boot-time job into a timer unit that runs a minute after boot, or narrow a wait-online configuration to the interface that matters. Making the slow unit itself faster is usually the last option, not the first — and it is worth nothing at all if the unit was never on the chain.

  • critical-chain points at a wait-online service holding the boot for twelve seconds. What are your options?
    Either make the wait honest or stop waiting on it. Narrow its configuration so it waits only for the interface that matters rather than every managed one — a second, unplugged NIC is the usual cause. Or remove the gating: if the units ordered behind `network-online.target` can retry, drop the dependency and let `Restart=on-failure` with a `RestartSec=` handle a network that is not ready yet.
  • Where does journalctl fit into a slow-boot investigation?
    After the analysis tools name a unit. `journalctl -b -u <unit>` shows that unit's messages from the current boot, and `journalctl -b -p warning` surfaces the warnings from the whole boot — which is where timeouts, retries and DNS failures announce themselves. The analysis tools tell you *which* unit cost time; only the log tells you what it was doing.
  • Why check the firmware and kernel figures before touching any unit?
    Because they are often the majority of the time and no unit change affects them. Server firmware doing memory training and option ROM initialisation can cost tens of seconds, and a boot loader menu timeout is pure waiting. Plain `systemd-analyze` prints the firmware, loader, kernel, initrd and userspace split, so one command tells you whether unit-level work can possibly help.

saying these in an interview costs you the question

  • Treats the top of blame as the boot bottleneck
  • Sums blame durations to explain total boot time
  • Skips the firmware/kernel/userspace split entirely
  • Reads @ and + in critical-chain as the same measurement
  • Assumes disabling the slowest unit always shortens boot

context