skip to content

On a Linux server, a process sits in state D and `kill -9` has no effect on it. The host mounts a network filesystem whose server has stopped responding. What does state D mean, and why can the process not be killed?

level: seniorimportance: should knowfreq 50%

answer

  1. blocked, but refusing to be woken
  2. signals are checked at the exit door
  3. the task never reaches the door
  4. it is waiting on hardware, not on you
  5. fix the storage, not the process

basics

~20 s

State D is uninterruptible sleep: the task is blocked inside a kernel operation that cannot be unwound safely, usually waiting on storage. SIGKILL is recorded as pending but only acted on when the task heads back to user space, which it cannot do until the operation finishes.

solid answer

~50 s

`D` means the task is sleeping in kernel mode in a way that declines to be woken by signals — classically while waiting for block I/O or a network filesystem reply. Signals are not injected into a running task; they are queued and inspected at the boundary where the kernel returns control to user space. A task in uninterruptible sleep never reaches that boundary, so your `SIGKILL` sits pending and does nothing visible. That is deliberate: the sleep usually straddles a half-finished operation holding locks and buffers, and unwinding it arbitrarily risks corrupting kernel state. The way out is to fix what it is waiting for — restore or fail the storage — after which the syscall completes or errors, the task returns toward user space, and the pending kill takes effect instantly. Modern kernels also offer a killable variant of the sleep, used by some paths including parts of NFS, where fatal signals do break the wait.

go deeper

for a junior

Know that D means the process is blocked inside the kernel waiting for I/O, that it is normal for short periods, and that no signal — SIGKILL included — will move it while it is there.

for a middle

Explain that signals are examined only when a task returns toward user space, so a pending SIGKILL simply cannot be acted on yet, and contrast interruptible S sleep, which a signal does wake.

for a senior

Reason from the symptom to the cause: persistent D usually means storage or a network filesystem is not answering, the load average is inflated by those tasks, restarts only pile up more of them, and the fix is at the storage layer or a reboot.

for a principal

Own the design consequence: hard-mounted remote storage converts a dependency outage into unkillable processes on every client, so decide deliberately where you accept that blocking, where soft failure semantics are safe, and how the platform contains the blast radius.

## What the state letters mean The Linux kernel keeps every task in one of a small number of states, which tools surface as single letters: - **R** — running or runnable: on a CPU, or queued waiting for one. - **S** — interruptible sleep: waiting for an event, and willing to be woken early by a signal. The overwhelming majority of idle processes on any machine. - **D** — uninterruptible sleep: waiting inside the kernel, refusing signal wakeups. - **Z** — zombie: terminated, awaiting collection of its status by its parent. - **T** — stopped, typically by a job-control signal; **t** — stopped by a tracer. - **I** — idle, used for idle kernel threads so they do not inflate load figures. `S` and `D` are both "blocked", and the difference between them is the whole question. ## Why signals cannot reach a D-state task A signal is not an interrupt delivered into arbitrary code. Sending one sets a bit in the target's pending mask and tries to wake it. The target only *acts* on pending signals at a well-defined point: when it is about to return from kernel mode to user mode. That is the only moment at which the kernel knows the task holds no half-modified internal state. An interruptible (`S`) sleep cooperates: a signal wakes it, the syscall unwinds — commonly returning `EINTR` — and the task heads for the user-space boundary where the signal, including a fatal one, is handled. An uninterruptible (`D`) sleep does not accept that wakeup at all. The task stays parked until the thing it is waiting for completes. Your `SIGKILL` really was delivered; it is queued, and it will fire the instant the task can leave the kernel. Note that this is *not* the same as blocking a signal: `SIGKILL` cannot be blocked or caught. It is a matter of when the task next looks. ## Why the kernel does this at all Uninterruptible sleep is chosen for operations where waking early would leave things inconsistent: a page being read into the page cache with buffers and locks pinned around it, a filesystem transaction mid-flight, a device driver waiting for hardware that will complete whether or not anyone still cares. Allowing an arbitrary signal to abandon that would mean writing a correct unwind path for every such point in the kernel — and getting one wrong corrupts memory or a filesystem. A stuck process is the far cheaper failure. This is also why brief `D` states are entirely normal. Every disk read passes through one. Only a *persistent* `D` is pathological, and it almost always means the underlying storage is not answering: a dead network filesystem server, a SAN path that has gone away, a failing disk retrying, or a device driver wedged. ## The consequences an operator sees On Linux the load average counts uninterruptible tasks as well as runnable ones — a deliberate historical choice — so a handful of processes hung on dead storage can drive a load figure into double digits on an otherwise idle box. That mismatch between "load is 40" and "the CPUs are doing nothing" is one of the standard fingerprints of a storage hang. The processes also cannot be cleaned up. They will not exit, they will not release their references to the mount, and anything that touches the same mount joins them. A supervisor that tries to restart the service simply adds another `D`-state task. ## The killable middle ground Since Linux 2.6.25 there is a third option, an uninterruptible sleep that *does* respond to fatal signals. Kernel paths that can unwind safely when the process is going to die anyway — parts of NFS among them — use it, which is why on modern kernels `kill -9` sometimes does free a task hung on a network filesystem. It is not guaranteed, and the classic block-I/O paths are not among them. This is also the reason the old NFS `intr`/`nointr` mount options are ignored on modern kernels: fatal signals can already interrupt the waits those options were invented for. ## How the situation actually ends There is no way to force a task out of an uninterruptible wait from user space — no signal, no priority change, no scheduler intervention. What is left is to address the wait itself: - restore the unresponsive server or storage path, so the pending operation completes or fails and every parked task drains at once (pending kills then take effect immediately); - for a network filesystem, mount with `soft` so operations give up with an error after retries instead of waiting forever — accepting that returning errors mid-write can corrupt application data, which is why `hard` is the default; - failing all that, reboot the host. Saying "I'd send a stronger signal" is the answer that fails this question. There is no stronger signal than `SIGKILL`, and the problem is not the signal's strength.

  • Was the SIGKILL lost, or is it waiting?
    It is waiting. The signal was recorded in the task's pending set; SIGKILL cannot be blocked or caught, so nothing rejected it. Pending signals are only acted on when the task is about to return to user space, which an uninterruptible sleep prevents. The moment the underlying operation completes or errors out, the task heads for that boundary and dies immediately.
  • Why does a Linux box with several D-state processes report a high load average while the CPUs are idle?
    Because Linux counts uninterruptible tasks in the load average alongside runnable ones. A handful of processes parked on dead storage can push the figure into double digits with no CPU work happening at all. That contradiction — high load, idle processors — is a strong fingerprint of a storage or network-filesystem hang rather than a compute problem.
  • If uninterruptible sleep is so painful, why not make every wait interruptible?
    Because each interruptible wait needs a correct unwind path for whatever state is held at that point — pinned buffers, locks, an in-flight filesystem transaction, a device command already issued. Getting one wrong corrupts memory or on-disk data. The kernel does offer a killable variant for paths that can unwind safely when the task is dying anyway, and network filesystem code uses it.

saying these in an interview costs you the question

  • Says SIGKILL was blocked by the process
  • Suggests kill -9 harder or a stronger signal exists
  • Thinks D means the process is dead or defunct
  • Believes renice or SIGSTOP can free the task
  • Treats any D state as a fault rather than normal brief I/O

context