On a Linux host you send SIGKILL to a hung process, and minutes later `ps` still shows it in state D. Why has SIGKILL not taken effect, and what does the kernel do with the signal in the meantime?
answer
- the signal is not lost
- delivery happens at a boundary
- the task never returns to user space
- fix the I/O, not the process
- check SigPnd in /proc
basics
~20 sThe signal is recorded as pending, not lost. A task in state D is in uninterruptible sleep inside a kernel call — usually blocked on I/O — and signals are only acted on when a task heads back toward user space, which it never reaches.
solid answer
~50 sSIGKILL is not lost, it is pending. Signals are acted on only when a task returns from kernel mode to user mode, and a task in state D is in uninterruptible sleep inside a kernel call — classically waiting on a stalled disk, an unreachable NFS server or a wedged FUSE daemon — so it never reaches that check. You can see the pending bit in `/proc/<pid>/status` under `SigPnd`. Nothing you send from user space will move it: the process dies the instant that I/O completes or errors out, and not a moment before. So you fix the storage, not the process — restore the server, recover the device — and if the resource is genuinely gone, a reboot is the only lever left. Some kernel paths use killable sleeps so fatal signals do work, and NFS `soft` mounts fail the operation rather than waiting forever.
code
bash · 2 linesgrep -E '^(State|SigPnd|SigBlk|SigIgn):' /proc/4242/status
cat /proc/4242/wchan; echogo deeper
Know that a process shown in state D is stuck waiting on I/O and that kill -9 will not clear it, and that the usual culprit is a disk or network filesystem rather than the program itself.
Explain that signals are acted on when a task returns toward user space, that state D means uninterruptible sleep inside a kernel call, and that the SIGKILL sits pending until the operation ends. Be able to point at SigPnd in /proc/<pid>/status as proof.
Demonstrate the diagnosis: correlate several processes stuck on one mount, read the hung-task message and the kernel stack, identify the failing storage, and fix that rather than the process. Be ready to explain the hard versus soft NFS trade-off you would set on that mount.
Own the design consequence: any dependency that can block a thread inside the kernel indefinitely is an availability risk that timeouts in application code cannot rescue. Be ready to argue for bounded-failure storage semantics, isolation of hosts that mount remote filesystems, and a documented reboot policy for wedged nodes.
## Delivery is not instantaneous Sending a signal does not run anything in the target. `kill(2)` sets a bit in the target task's **pending** set and, if the task is in an interruptible sleep, wakes it. The signal is only *acted on* at a specific point: when the task is about to return from kernel mode to user mode. That is when the kernel examines the pending set, subtracts the blocked mask, and either runs a handler or applies the default action. SIGKILL is special in that its action is applied by the kernel rather than by the target's own code, but it is not exempt from that scheduling point. The task still has to be in a state where the kernel can act on it. ## What state D means `ps` reports Linux task states with single letters — `R` runnable, `S` interruptible sleep, `D` uninterruptible sleep, `Z` zombie, `T` stopped. `D` is `TASK_UNINTERRUPTIBLE`: the task is blocked inside a kernel call at a point where the code path cannot safely be unwound. It holds kernel state — a half-finished filesystem operation, a page fault against a mapped file, a device request submitted and awaiting completion — that has no correct "abandon here" answer. So the kernel deliberately refuses to wake it for signals. The SIGKILL sits in `SigPnd` and waits. When the underlying operation finally completes or returns an error, the task unwinds, reaches the return-to-user-space boundary, sees the pending fatal signal, and dies instantly. From the outside it looks like your `kill -9` finally worked ten minutes later; in reality it took effect the microsecond it became possible. ## Where these come from in production The classic sources are all storage: a disk that has stopped answering and is quietly retrying, a hard-mounted NFS export whose server is unreachable, a network block device that lost its target, a FUSE filesystem whose user-space daemon crashed while requests were outstanding, or a device driver bug. A useful signature is a whole cluster of unrelated processes in D at once — every one that touched the same mount point — rather than a single unlucky program. The kernel will usually tell you. When a task stays blocked past `kernel.hung_task_timeout_secs` (120 seconds by default), the hung-task detector logs `INFO: task <name>:<pid> blocked for more than 120 seconds` along with a stack trace. Filesystem and driver layers log their own complaints — for NFS, repeated messages about a server not responding. ## What evidence to collect `/proc/<pid>/status` shows the state and the signal masks: `SigPnd` (pending for this thread), `SigBlk`, `SigIgn`, `SigCgt`. A set bit for signal 9 in `SigPnd` is direct proof that your SIGKILL arrived and is waiting. `/proc/<pid>/wchan` names the kernel function the task is sleeping in, and with sufficient privilege `/proc/<pid>/stack` gives the kernel stack — which layer is stuck is usually obvious from the symbol names. Correlate that with the kernel log, and you have your answer without ever touching the process again. ## What actually resolves it Nothing you send from user space. Not a second SIGKILL, not signalling the parent, not changing priority, not a debugger. The only real fixes address the thing being waited on: - restore or restart the unreachable server, or repair the network path to it; - recover the failing device, or let its error handling finally time out and return an error; - restart the crashed FUSE daemon so its outstanding requests are completed with an error; - as a last resort, when the resource is permanently gone, reboot the host. Mount options matter here as prevention. An NFS mount with the `soft` option gives up after its retries and returns an error to the application, which unblocks the task at the cost of applications seeing I/O errors they must handle. A `hard` mount — the default, and the right choice when data integrity matters — waits indefinitely, which is precisely the behaviour you are looking at. ## The killable middle ground Because permanently unkillable processes are so unpleasant to operate, the kernel gained a third sleep state, `TASK_KILLABLE`: the task ignores ordinary signals but *does* wake for a fatal one. Several paths, notably in the NFS client, were converted to use it. That is why `kill -9` sometimes does clear a stuck NFS process and sometimes does not — it depends on which wait the task happens to be sitting in. ## The look-alike case One other process survives `kill -9` and appears in `ps`: a zombie, state `Z`. That is the opposite situation. A zombie is already dead — there is no task left to signal, only an exit status held for a parent that has not collected it. Signals do nothing there because there is nothing to signal, whereas a task in D is very much alive and merely unreachable.
- Both a task in state D and a zombie survive `kill -9`. What is the difference in what the signal does?A task in D is alive but blocked inside the kernel, so the signal is recorded as pending and takes effect the moment the task can unwind. A zombie is already dead — nothing remains to signal but an exit status waiting to be collected — so a signal has no target at all and simply does nothing. One is a delivery problem, the other is a bookkeeping entry.
- What is TASK_KILLABLE and why was it introduced?It is a sleep state that ignores ordinary signals but wakes for a fatal one. It exists because unkillable processes blocked on remote storage were a persistent operational problem, so paths that could safely abandon their work — notably in the NFS client — were converted to it. It explains why `kill -9` clears some stuck NFS processes and not others.
- How would you get supporting evidence from the kernel that this is an I/O stall rather than an application bug?Check the kernel log for the hung-task detector's message about a task blocked for more than 120 seconds, tuned by `kernel.hung_task_timeout_secs`, together with any filesystem or driver complaints. Then read `/proc/<pid>/wchan` and, with privilege, `/proc/<pid>/stack` to see which kernel layer the task is sleeping in. Several unrelated processes stuck on the same mount is the strongest hint.
saying these in an interview costs you the question
- Claims SIGKILL was dropped and must be resent
- Says SIGKILL can force a syscall to abort
- Thinks a bigger signal number would work
- Confuses uninterruptible sleep with a zombie
- Reaches for a debugger instead of the storage layer