skip to content

On a systemd host, `systemctl stop myapp.service` takes about ninety seconds every time and the journal then reports that systemd killed the process. Walk through what systemd does during a stop job, and which unit settings you would change.

level: seniorimportance: should knowfreq 50%

answer

  1. nothing is hanging, something is waiting
  2. a polite signal first, then a rude one
  3. ninety seconds is a default, not a coincidence
  4. which processes get signalled is configurable
  5. shortening the wait does not fix the cause

basics

~20 s

The service is not reacting to SIGTERM, so systemd waits out TimeoutStopSec= — 90 seconds by default — and then sends SIGKILL. A stop job runs ExecStop= if present, signals the unit's processes with KillSignal=, waits, then escalates. Fix the signal handling rather than shortening the timeout.

solid answer

~40 s

A stop job has a fixed shape. systemd runs `ExecStop=` if the unit defines one, then sends `KillSignal=` — SIGTERM unless changed — to the processes in the unit, chosen by `KillMode=`, which defaults to `control-group` and therefore signals every process the unit spawned. It then waits up to `TimeoutStopSec=`, 90 seconds by default on most distributions, and if anything is still alive it escalates to SIGKILL because `SendSIGKILL=yes` is the default. The journal spells this out: "State 'stop-sigterm' timed out. Killing." followed by "Killing process NNN with signal SIGKILL". A ninety-second stop every time means the process is ignoring or mishandling SIGTERM. The correct fix is a SIGTERM handler in the service, or `KillSignal=` if the daemon shuts down on a different signal — shortening `TimeoutStopSec=` only makes the hard kill arrive sooner.

code

ini · 6 lines
ini
[Service]
ExecStart=/usr/sbin/exampled --foreground
KillSignal=SIGQUIT
KillMode=mixed
TimeoutStopSec=45s
SendSIGKILL=yes

go deeper

for a junior

Know that systemd asks a service to stop with SIGTERM first, waits, and only then kills it outright, and that the wait has a default of about ninety seconds.

for a middle

Lay out the stop job in order — ExecStop=, KillSignal= to the set chosen by KillMode=, TimeoutStopSec=, then SIGKILL — and name the default value of each.

for a senior

Separate a broken SIGTERM handler from a genuinely slow drain from a process stuck on unresponsive I/O, and change the setting that matches the cause instead of shortening the timeout.

for a principal

Define what graceful shutdown means for your services — how long a drain may take, what happens to in-flight work when it is cut short — and make the unit timeouts and the application's shutdown path agree with that contract.

## The stop sequence, in order When you run `systemctl stop foo.service`, systemd performs these steps: 1. **`ExecStop=`**, if the unit defines one. This is an optional, unit-supplied command — typically a graceful-shutdown CLI call. It is *not* how the service is terminated; it is a hook that runs first. 2. **The stop signal.** systemd sends `KillSignal=` (default SIGTERM) to the processes selected by `KillMode=`. 3. **The wait.** systemd waits up to `TimeoutStopSec=` for every process to be reaped. The unit sits in `deactivating` throughout. 4. **The escalation.** If anything survives the wait, systemd sends SIGKILL, because `SendSIGKILL=yes` is the default, and waits again briefly. The log lines from steps 3 and 4 are the ones to recognise: ``` foo.service: State 'stop-sigterm' timed out. Killing. foo.service: Killing process 3172 (exampled) with signal SIGKILL. foo.service: Main process exited, code=killed, status=9/KILL foo.service: Failed with result 'timeout'. ``` ## KillMode: which processes get signalled `KillMode=` decides the set of processes, and the default matters: - `control-group` (default) — every process still associated with the unit is signalled, not just the one systemd tracks as the main process. This is why a wrapper script or a service that spawns workers still shuts down cleanly under systemd even when the wrapper forwards nothing. - `mixed` — SIGTERM to the main process only, but SIGKILL at the escalation step to all of them. This is right when the main process is a supervisor that must coordinate its own children's shutdown and would be upset by them receiving SIGTERM independently. - `process` — only the main process is signalled at every stage. Surviving children are left running, which is almost always a mistake unless you have a specific reason. - `none` — signal nothing; run `ExecStop=` and hope. Strongly discouraged. ## Where the ninety seconds comes from `TimeoutStopSec=` defaults to the manager-wide `DefaultTimeoutStopSec=` set in `/etc/systemd/system.conf`, which is 90 seconds on typical distributions. If the unit's own `TimeoutStopSec=` is unset, that is the value in force. You can set it per unit, and `TimeoutSec=` sets the start and stop timeouts together. Passing `infinity` disables the escalation entirely — appropriate for a database that must be allowed to finish flushing however long it takes, and dangerous everywhere else because a stuck stop job will block shutdown of the whole machine. ## Diagnosing the actual cause A reliable ninety-second stop points at one of a small number of causes: - **The service has no SIGTERM handler.** The default disposition of SIGTERM is termination, so this normally cannot happen — unless the program installed a handler that does nothing, blocked the signal, or handles it by starting a shutdown that never completes. - **Shutdown is genuinely slow.** Draining long-lived connections, finishing an in-flight batch, flushing a large write buffer. Here the timeout is not the bug; the timeout is too short, and you should raise `TimeoutStopSec=` rather than accept a hard kill that loses work. - **The daemon expects a different signal.** Some daemons reserve SIGTERM for a fast shutdown and use another signal for the graceful one — nginx, for instance, treats SIGQUIT as its graceful shutdown. `KillSignal=SIGQUIT` in the unit tells systemd to use it. - **A child is stuck in an uninterruptible state.** If a process is blocked on unresponsive I/O, even SIGKILL will not reap it until the I/O completes, and the stop job will overrun both timeouts. The evidence to collect: `journalctl -u foo.service -b` for the exact timeout line, `systemctl status foo.service` during the stop to see the `deactivating` state and which processes remain listed under the unit, and the application's own log for whether it acknowledged a shutdown request at all. ## What to change In order of preference: implement or fix SIGTERM handling in the service; set `KillSignal=` if the daemon's graceful signal is something else; raise `TimeoutStopSec=` when the shutdown work is real; set `KillMode=mixed` when the main process must own its children's shutdown. Lowering `TimeoutStopSec=` to make the stop "faster" only moves the SIGKILL earlier and guarantees the abrupt termination you were trying to avoid. ```ini [Service] ExecStart=/usr/local/bin/exampled KillSignal=SIGQUIT KillMode=mixed TimeoutStopSec=45s ``` When you need to end it now rather than wait, `systemctl kill -s SIGKILL foo.service` sends the signal immediately instead of waiting for the escalation.

  • When is raising TimeoutStopSec= the right answer rather than a workaround?
    When the shutdown work is real — draining long-lived connections, finishing an in-flight batch, flushing a large buffer — and the current default cuts it off mid-way. In that case the SIGKILL is destroying work, and giving the service the time it genuinely needs is the fix. It is a workaround only when the service is not shutting down at all.
  • What does KillMode=mixed change compared with the default?
    The default control-group sends SIGTERM to every process in the unit. Mixed sends SIGTERM only to the main process, then SIGKILL to all of them at the escalation step. Use it when the main process is a supervisor that must sequence its children's shutdown itself and would misbehave if they received SIGTERM independently.
  • Why might SIGKILL fail to end a stop job at all?
    A process blocked in an uninterruptible wait cannot be reaped until that wait completes — a hung network filesystem is the usual cause. The signal is recorded but not acted on, so the unit stays in deactivating past both timeouts. The diagnosis moves off systemd entirely and onto whatever I/O the process is stuck on.

saying these in an interview costs you the question

  • Lowers TimeoutStopSec to make the stop faster
  • Thinks ExecStop is what terminates the service
  • Assumes only the main process receives SIGTERM
  • Believes SIGKILL always ends the process immediately
  • Treats a slow drain and a broken handler as the same problem

context