A systemd unit declares both Requires= and After= on its database unit, and the ordering is confirmed correct, yet the service still fails its first connection at boot. What does After= actually wait for, and how can a unit declare that it is genuinely ready?
answer
- started is not the same as ready
- the guarantee depends on Type=
- forked, not initialised
- the daemon has to say so itself
- remote dependencies have no ordering edge
basics
~20 sAfter= waits for the other unit to be considered started, which is not the same as ready. For Type=simple that means only that systemd forked the process. A daemon signals real readiness with Type=notify and sd_notify, otherwise the consumer must retry.
solid answer
~50 sOrdering guarantees a state transition, and which transition depends entirely on the dependency's `Type=`. With `Type=simple` — still the most common — systemd considers the unit started the instant it has forked the main process, before a single line of the program's initialisation has run, so `After=` returns almost immediately. `Type=exec` tightens that to "execve succeeded", `Type=forking` to "the parent exited", `Type=oneshot` to "the process exited", and `Type=dbus` to "the bus name was acquired". The one that means what you want is `Type=notify`: systemd holds the unit in activating until the daemon itself sends `READY=1` over `sd_notify`, so downstream `After=` genuinely waits for the daemon to be serving. If you cannot change the dependency's type, ordering cannot express readiness at all, and the consumer has to retry with backoff — which is the right design anyway, because the ordering graph only covers processes on this one host.
code
ini · 11 lines[Unit]
Description=Orders API
Wants=postgresql.service
After=postgresql.service
[Service]
Type=notify
ExecStart=/usr/local/bin/orders-api
TimeoutStartSec=60
Restart=on-failure
RestartSec=5sgo deeper
Know that After= only waits for systemd to consider the other unit started, and that for the common Type=simple this happens as soon as the process is forked — long before it can serve.
Map the types to their transitions: simple forks, exec execs, forking waits for the parent to exit, oneshot waits for exit, dbus waits for the bus name, notify waits for READY=1 via sd_notify.
Show the triage — check the dependency's Type=, interleave both units' logs for that boot with journalctl — and give the durable answer: the consumer retries, because ordering cannot cover a restart at 3 a.m. or a database on another host.
Set the standard that services tolerate absent dependencies. Boot ordering is a local optimisation that stops working the moment anything moves off the host, so readiness and retry belong in the application contract, with unit ordering merely reducing noise at boot.
## Ordering waits for a state, and the state is type-dependent `After=b.service` means: do not run my start job until b.service's start job has completed. The question interviewers are really asking is when systemd decides that job is complete, and the answer is set by `Type=` in b.service's `[Service]` section: - **`Type=simple`** — considered started as soon as the main process has been forked. systemd does not wait for `execve` to succeed, let alone for the program to open a socket. Anything ordered after it starts essentially immediately. - **`Type=exec`** — considered started once `execve()` has succeeded. Strictly better than `simple` for catching a missing binary or a bad `User=`, but still says nothing about initialisation. - **`Type=forking`** — the traditional daemon: started when the original process exits, having forked its long-lived child. Correct for old daemons, and the reason `PIDFile=` exists. - **`Type=oneshot`** — started when the process *exits*, which makes it the one type where ordering does imply the work finished. This is why migrations, provisioning steps and firewall loads belong in oneshot units. - **`Type=dbus`** — started when the service acquires the bus name given by `BusName=`. - **`Type=notify`** — started when the daemon sends `READY=1` on the notification socket. So the same `After=` line is nearly meaningless behind a `Type=simple` unit and a strong guarantee behind a `Type=notify` one. ## The readiness protocol systemd passes a Unix socket path in the `NOTIFY_SOCKET` environment variable. The daemon writes newline-separated `KEY=VALUE` messages to it; `READY=1` is the one that ends activation: ```ini [Service] Type=notify ExecStart=/usr/local/bin/orders-api TimeoutStartSec=60 ``` ```c /* after listeners are bound and the pool is warm */ sd_notify(0, "READY=1"); ``` The daemon can also send `STATUS=` strings that show up in `systemctl status`, `RELOADING=1`, and `WATCHDOG=1` heartbeats. Note the trade-off you have taken on: with `Type=notify`, a daemon that never signals hangs in `activating` until `TimeoutStartSec=` fires and the unit fails. That is usually what you want — an unready service should not be declared up — but it must be a deliberate choice, and the timeout must be long enough for a genuine cold start. Many widely deployed daemons already ship `Type=notify` units. Checking `systemctl show -p Type <unit>` on the dependency before you assume anything is the two-second version of this whole investigation. ## When you cannot change the dependency Three honest options, in descending order of preference: 1. **Make the consumer retry.** Connect in a loop with backoff, or let the unit fail and be restarted: `Restart=on-failure` with a `RestartSec=` of a few seconds. This is correct at boot and, crucially, still correct at 3 a.m. when the database restarts and no ordering graph is involved at all. 2. **Gate with an ExecStartPre probe.** A short command that polls until the dependency answers, bounded by `TimeoutStartSec=`, moves the wait into systemd where it is visible in `systemctl status` and the journal. It is a crutch, and an unbounded probe is worse than no probe, but it beats a fixed sleep. 3. **Fix the dependency's unit.** A drop-in changing `Type=simple` to `Type=notify` only works if the daemon actually implements the protocol — otherwise you have turned a fast start into a guaranteed timeout. `Type=exec` is a safe tightening; `Type=notify` is not, unless the program supports it. What is *not* an option is `ExecStartPre=/bin/sleep 10`. It is a guess that is simultaneously too long on every healthy boot and too short on the one slow boot that matters, and it hides the real dependency from anyone reading the unit later. ## The limit of the whole mechanism Ordering dependencies are a property of one host's unit graph. The moment the database is on another machine — which it usually is — there is nothing to order against, and no `Type=` on earth helps. That is the real reason the retry answer is the senior one: it is the only approach that covers the local case, the remote case, and the post-boot case with a single mechanism. Ordering is a boot-time optimisation that avoids a burst of pointless failures; it is not a correctness guarantee, and a service that treats it as one is fragile in exactly the situations where you need it not to be. (Socket activation offers a different way to remove the race entirely, and it is a topic in its own right.) ## Diagnosing it on a live host ```bash systemctl show -p Type -p NotifyAccess postgresql.service systemctl list-dependencies --after orders-api.service journalctl -b -u orders-api.service -u postgresql.service ``` Interleaving the two units' logs from the same boot with a single `journalctl` invocation shows the actual sequence: your service's connection error will be timestamped *before* the database's "ready to accept connections" line, which settles the argument about whether the ordering was the problem.
- What is the risk of changing a unit from Type=simple to Type=notify with a drop-in?If the daemon does not implement the protocol it will never send `READY=1`, so the unit sits in `activating` until `TimeoutStartSec=` expires and then fails — converting a fast, if imprecise, start into a guaranteed failure. Confirm the program supports `sd_notify` first. `Type=exec` is the safe tightening when it does not: it only waits for `execve()` and cannot hang.
- Why is Type=oneshot the type where ordering genuinely means the work is done?Because a oneshot unit is considered started only when its process has exited, so anything ordered after it runs when the work has actually completed. That makes it the right shape for schema migrations, provisioning steps and firewall rule loading. Pair it with `RemainAfterExit=yes` if you want the unit to stay shown as active once the job has finished.
- How do you prove from the logs that the ordering, rather than the ordering being ignored, was the problem?Interleave both units in one query for the current boot: `journalctl -b -u consumer.service -u dependency.service`. The entries are merged in timestamp order, so you can see whether the consumer's connection error precedes the dependency's own "ready" message. If it does, ordering was honoured and readiness was the gap; if the consumer started before the dependency's start job at all, the ordering edge itself is missing.
saying these in an interview costs you the question
- Assumes After= waits until the service can serve requests
- Adds sleep to ExecStartPre and calls it fixed
- Thinks Type=simple waits for the program to initialise
- Switches a dependency to Type=notify without the daemon supporting it
- Believes unit ordering can cover a database on another host