skip to content

questions

4

On Linux, the `kill` command sends SIGTERM by default and `kill -9` sends SIGKILL. What is the difference between the two signals, and why can a process never handle SIGKILL?

level: juniorimportance: must knowfreq 85%

answer

  1. one is a request, one is not
  2. the kernel refuses two signals
  3. handlers are code inside the target
  4. sigaction returns EINVAL for it
  5. escalate: TERM, grace period, KILL

basics

~20 s

SIGTERM is a polite request: the process can catch it and shut down cleanly. SIGKILL is enforced by the kernel — the target never runs code for it, so buffers, lock files and in-flight work are abandoned as they are.

solid answer

~40 s

`kill` sends SIGTERM by default, and SIGTERM is a request: the process can install a handler and use it to stop accepting work, finish what is in flight, flush buffers and exit on its own terms. `kill -9` sends SIGKILL, which the kernel refuses to let any process catch, block or ignore — `sigaction` on it fails with `EINVAL`. The target never runs another instruction; the kernel simply tears the task down. So SIGKILL always works, but it is never clean: unflushed writes are lost, temporary and lock files stay behind, requests die mid-flight, and any shutdown hook the program has is skipped. The right pattern is escalation — send SIGTERM, allow a bounded grace period, and only then send SIGKILL. SIGSTOP is the other signal a process cannot intercept.

code

bash · 3 lines
bash
kill 4242                                  # SIGTERM by default: ask it to stop
sleep 5                                    # bounded grace period
kill -0 4242 2>/dev/null && kill -9 4242   # still there? escalate to SIGKILL

go deeper

for a junior

Know that plain kill sends SIGTERM and kill -9 sends SIGKILL, and be able to say plainly that the first can be handled and the second cannot. Mentioning that SIGKILL skips any cleanup the program would have done is enough at this level.

for a middle

Explain the mechanism: dispositions installed with sigaction, default actions, and the fact that SIGKILL and SIGSTOP are rejected outright by the kernel. Be ready to describe what a good SIGTERM handler actually does — drain, flush, release, exit.

for a senior

Show you have operated this. Talk about how you pick a grace period from the longest real unit of work, what state is left behind by a hard kill in your particular stack, and how you diagnose a process that does not die on SIGTERM rather than reflexively escalating to -9.

for a principal

Own the policy: where the shutdown deadline is defined, whether it is consistent across every supervisor in the estate, and what your systems are allowed to lose when the deadline expires. Be ready to argue for crash-only design — make hard kills survivable — instead of ever-longer grace periods.

## What a signal actually is A signal is a one-number notification the kernel delivers to a process to interrupt whatever it was doing. Every signal has a name (`SIGTERM`), a number (15 on Linux), and a **default action** the kernel takes when the process has expressed no preference: terminate, terminate and write a core dump, ignore, stop, or continue. A process expresses a preference by installing a **disposition** with `sigaction(2)`: either a handler function of its own or the explicit constant `SIG_IGN`. `kill -l` prints the full list on the machine in front of you. Two consequences of that model explain this whole question. A handler is *code inside the target process*, so it only runs if the kernel lets the target run. And the kernel, not the process, decides which signals a process is allowed to have an opinion about. ## SIGTERM: the request SIGTERM means "please stop." It is what `kill <pid>` sends when you name no signal, what service supervisors send first, and what a well-behaved program is expected to act on. Its default action, if no handler is installed, is to terminate the process — so a program that ignores signals entirely still dies. The point of SIGTERM is that the program *may* interpose: ```c signal handler -> stop accepting new work -> let in-flight requests finish -> flush buffers, close files, release locks -> exit(0) ``` That window is where every clean-shutdown behaviour lives: draining connections, committing a transaction, removing a PID file, deregistering from a load balancer. SIGINT (2, what Ctrl+C sends to the foreground process group) and SIGQUIT (3, which also writes a core dump) are catchable in exactly the same way; SIGTERM is simply the one conventionally used by tooling rather than by a human at a keyboard. ## SIGKILL: the order SIGKILL cannot be caught, blocked or ignored. This is not a convention — it is enforced: calling `sigaction(2)` for SIGKILL fails with `EINVAL`, and adding it to a blocked signal mask silently has no effect. The kernel does not ask the process to do anything; it removes the task itself. The process executes no further user-space instruction after the signal takes effect. What the kernel *does* clean up is everything the kernel owns: the address space is freed, file descriptors are closed, kernel-held file locks are released, the exit is reported to the parent. What nobody cleans up is everything user space owned: - data still sitting in the program's own buffers is gone; - temporary files, sockets on disk and stale lock/PID files remain; - in-flight requests are dropped with no response; - a data store that needs a shutdown to be consistent must now run crash recovery on next start; - a mutex held in shared memory can be left locked, wedging surviving processes. This is why "just `kill -9` it" is a real habit to break: it works every time, which is precisely why people reach for it before finding out why SIGTERM was not enough. ## Why the asymmetry exists If a process could refuse SIGKILL, there would be no way to remove a buggy, wedged or hostile program short of a reboot. So the kernel reserves exactly two signals as uninterceptable: SIGKILL, which always kills, and SIGSTOP, which always suspends. Everything else — including SIGTERM — is negotiable by the program, which is the entire reason SIGTERM is useful as a shutdown request. Sending is still privileged: an unprivileged process may signal a target only when its real or effective user ID matches the target's real or saved set-user-ID; otherwise it needs `CAP_KILL`. And `kill -0 <pid>` sends nothing at all — it runs only the existence and permission check, which is the standard way to ask "is this process still there?" ## The escalation pattern Because SIGTERM may be handled badly or not at all, every supervisor implements the same escalation: send SIGTERM, wait a bounded grace period, then send SIGKILL to whatever is left. Choosing that grace period is a real design decision — it must exceed the longest legitimate in-flight unit of work, or the graceful path is a fiction and you always pay the hard kill anyway. A process can also appear to ignore SIGTERM without any handler. It may have **blocked** the signal with `sigprocmask(2)`, in which case the signal stays pending until unblocked. It may be **stopped** (after SIGSTOP), so nothing is delivered until SIGCONT. Or it may be in uninterruptible sleep inside a kernel call, where even a pending SIGKILL waits. In the first two cases SIGKILL still wins immediately; in the third, nothing you send helps until the I/O resolves.

  • How can a process appear to ignore SIGTERM even though it has installed no handler?
    Three ways. It may have blocked SIGTERM with `sigprocmask`, so the signal sits pending until unblocked. It may be stopped after SIGSTOP, so nothing is delivered until SIGCONT arrives. Or it may be in uninterruptible sleep inside a kernel call, where even pending fatal signals wait. In the first two cases SIGKILL still takes effect at once.
  • Is there any way for a program to protect itself from SIGKILL?
    Not from user space — the disposition simply cannot be changed. The only real protection is the permission check on the sender: you may signal a process only if your real or effective UID matches its real or saved set-user-ID, or you hold `CAP_KILL`. Anything else, such as sitting in uninterruptible sleep, delays the kill rather than preventing it.
  • What does SIGQUIT do that SIGTERM does not?
    SIGQUIT's default action is to terminate the process *and* write a core dump, which makes it the signal of choice when you want the post-mortem rather than a clean exit. It is what Ctrl+\ sends to the foreground process group. Like SIGTERM it is catchable, and some runtimes install their own handler for it to dump diagnostic state instead of dying.

saying these in an interview costs you the question

  • Says SIGKILL lets the program run cleanup first
  • Claims kill -9 is the normal way to stop a service
  • Thinks SIGTERM cannot be ignored by a process
  • Believes kill -9 always removes the process instantly
  • Confuses SIGKILL with SIGSTOP as 'the forceful one'

context

open as a page

A long-running command you started in an interactive SSH session keeps dying whenever the connection drops. Which signal kills it, what makes the kernel send that signal, and how would you start the job so it survives?

level: middleimportance: should knowfreq 60%

basics

~20 s

SIGHUP kills it. When the SSH connection drops the kernel hangs up the controlling terminal and signals the session, and SIGHUP's default action is to terminate. Start the job under nohup or setsid, or inside a terminal multiplexer.

open as a page

On a Linux host you send SIGKILL to a hung process, and minutes later `ps` still shows it in state D. Why has SIGKILL not taken effect, and what does the kernel do with the signal in the meantime?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The signal is recorded as pending, not lost. A task in state D is in uninterruptible sleep inside a kernel call — usually blocked on I/O — and signals are only acted on when a task heads back toward user space, which it never reaches.

open as a page

In a C program on Linux, what may a signal handler installed with `sigaction()` safely do, and why can calling something like `printf()` from a handler deadlock the process?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Only async-signal-safe functions, listed in the signal-safety(7) manual page — chiefly write() and _exit() — plus setting a volatile sig_atomic_t flag. printf() takes locks the interrupted code may already hold, so re-entering it can deadlock.

open as a page