A Linux server with 8 CPUs shows a 1-minute load average of 30, yet CPU utilisation is near idle. What does Linux's load average actually count, and what does this combination point to?
answer
- it is not a CPU percentage
- two very different states are summed
- Linux counts a state Unix did not
- D means waiting in the kernel
- idle CPUs plus deep queues means storage
basics
~20 sLinux counts both runnable tasks and tasks in uninterruptible sleep (D state) in its load average, unlike traditional Unix. High load with idle CPUs therefore means processes are blocked in the kernel waiting on storage or a hung mount, not competing for CPU.
solid answer
~50 sLinux's load average is not a CPU-utilisation figure. It is an exponentially damped moving average, over 1, 5 and 15 minutes, of the number of tasks in state **R** (running or runnable) *plus* those in state **D** (uninterruptible sleep) — Linux deliberately includes uninterruptible waits, which classic Unix did not, so the number reflects general system demand rather than run-queue length alone. Load 30 on 8 CPUs with idle CPUs therefore says roughly 30 tasks are waiting, but almost none of them want a CPU: they are blocked in the kernel on I/O. In practice that means a saturated or failing block device, a stalled network filesystem mount, or a storage path that has stopped completing requests. The next step is to find which tasks are in D state and what device they are waiting on — the load number itself tells you demand exists, not where it is.
code
bash · 2 lines# the three averages, then runnable/total tasks, then the last PID
cat /proc/loadavggo deeper
Know that the three numbers are 1-, 5- and 15-minute averages of waiting tasks, not percentages, and that they must be read against the machine's CPU count.
Explain that Linux counts runnable tasks plus uninterruptible (D state) sleepers, and that this is exactly why load can be high while the CPUs are idle.
Show the diagnostic path from the symptom to the device: identify D-state tasks and the mount or block device behind them, and explain why iowait is idle time rather than busy time.
Own the metric-choice argument — why load average is a poor SLO signal, what saturation and latency measures you would put in its place, and how you keep teams from capacity-planning against it.
## What the three numbers are `/proc/loadavg` exposes five fields; the first three are the load averages over 1, 5 and 15 minutes. They are **not** percentages, not normalised per CPU, and not sampled from CPU busy time. The kernel samples the count of qualifying tasks every few seconds and feeds it through an exponentially damped moving average, which is why the 1-minute figure moves quickly and the 15-minute one lags — comparing the three tells you whether the situation is arriving or receding. The remaining fields are the number of currently runnable tasks over the total number of tasks, and the most recently allocated PID. ## The Linux-specific part: D state counts Traditional Unix load average counted only the run queue: tasks running or waiting for a CPU. Linux additionally counts tasks in **`TASK_UNINTERRUPTIBLE`** — the D state you see in a process listing — which are tasks blocked inside the kernel in a sleep that cannot be interrupted by a signal, overwhelmingly waiting for a block-device or network-filesystem operation to complete. This single design choice is why Linux load averages behave the way they do and why the number is best read as **"system demand"** rather than **"CPU demand"**. It also explains the two classic confusions: - **High load, idle CPUs.** Tasks are piling up in D state on storage. The CPUs have nothing to do because everyone is waiting on a device. - **High load, busy CPUs.** Genuine CPU saturation, with tasks queued for CPU time. The same number, two completely different problems. You cannot tell which from the load average alone, which is the point of the question. ## Reading the magnitude A load average of *N* means roughly *N* tasks were in R or D during the window. Against 8 CPUs: - Load around 8 with busy CPUs is a fully-used machine, not necessarily an unhealthy one. - Load 30 with busy CPUs means about 22 tasks were queued waiting for a CPU — real run-queue saturation, with latency consequences. - Load 30 with idle CPUs, the case here, means the waiting is elsewhere: in the kernel, on I/O. The number is also not comparable between machines with different CPU counts, and the count includes *threads*, not just processes, so a heavily-threaded application can produce a large load figure with modest process counts. ## Diagnosing the idle-CPU case Once you know the load is D state rather than run queue, the question becomes *which device*. The signature findings are: - Many tasks parked in D state, often all belonging to the same application or all touching the same mount point. - A large fraction of time accounted as **iowait** — CPU time where the CPU was idle *and* at least one I/O was outstanding. Note the trap: iowait is a flavour of idle. High iowait does not mean the CPU is busy, and low iowait does not prove storage is healthy, because a CPU with other work to do will run it instead of accounting the wait. - A device with deep queues and high service times, or a network filesystem mount whose server is unreachable. The pathological version is a hung network mount: every process that touches it enters D state and never leaves, so load climbs indefinitely, the CPUs stay idle, and the tasks cannot even be killed with SIGKILL because an uninterruptible sleep does not process signals. That last detail — D-state tasks ignore signals until the operation completes or the kernel gives up — is the practical reason "the process will not die" and "the load average is 300" so often show up in the same incident. ## What it is not - Not a percentage. Load 1.0 is not "100% busy". - Not divided by CPU count. Interpreting it requires knowing how many CPUs the machine has. - Not instantaneous. The damping means a spike that just started barely moves the 1-minute figure, and a problem that ended minutes ago still inflates the 15-minute one. - Not container-aware in the usual case: a containerised process reading the load average generally sees the host's figure, because it is a global kernel statistic rather than a per-cgroup one. ## What an interviewer is checking That you do not treat load average as CPU utilisation; that you specifically know Linux folds uninterruptible sleep into it; and that from "load 30, CPUs idle" you immediately reach for storage or a stalled mount rather than for CPU capacity. Reaching for "add more CPUs" here is the wrong answer, because more CPUs would sit idle exactly like the current ones.
- Why does Linux include uninterruptible sleep in the load average when traditional Unix counted only the run queue?Because Linux wanted the number to express overall system demand rather than CPU demand alone. Tasks blocked in the kernel on a device are genuinely waiting on the system and their latency is real, so folding D state in makes the metric detect storage saturation as well as CPU saturation. The cost is ambiguity: one number now covers two different bottlenecks.
- A process is stuck in D state and SIGKILL does not remove it. Why not?An uninterruptible sleep does not check for pending signals — that is precisely what makes it uninterruptible. The task will only leave that state when the kernel operation it is blocked in completes or times out, which is why a hung network mount produces unkillable processes. Recovering usually means fixing or force-unmounting the underlying storage path, not sending stronger signals.
- Is a load average of 8 on an 8-CPU machine a problem?Not on its own. It means demand roughly matched capacity over that window, which can be a healthy, fully-utilised batch machine. Whether it is a problem depends on what the tasks were waiting for and whether request latency actually suffered — the load figure is a demand indicator, not a service-level one.
- How should you compare the 1-, 5- and 15-minute load averages?Read them as a trend. A 1-minute value well above the 15-minute one means load is arriving now; the reverse means an episode is receding and you may be looking at the tail of an incident that has already ended. Because they are exponentially damped, none of them reflects the instantaneous state.
saying these in an interview costs you the question
- Treats load average as a CPU utilisation percentage
- Says load 1.0 means the machine is fully busy
- Thinks Linux load counts only tasks waiting for CPU
- Recommends adding CPUs when load is high but CPUs are idle
- Believes high iowait means the CPU is busy working