skip to content

Explain how the kubelet decides that a Kubernetes node is under resource pressure: which node conditions it sets, and the difference between its hard and soft eviction thresholds.

level: middleimportance: must knowfreq 55%

answer

  1. signals: memory.available, nodefs/imagefs.available, inodesFree, pid.available
  2. hard = immediate, grace period 0
  3. soft = must persist for grace period, graceful kill capped
  4. conditions MemoryPressure / DiskPressure / PIDPressure → NoSchedule taints
  5. minimum-reclaim + 5m pressure-transition-period = no flapping

basics

~20 s

The kubelet watches signals like memory.available and nodefs.available. Crossing a hard threshold evicts pods immediately with no grace period; a soft threshold must stay crossed for a configured grace period first, and pods get a bounded termination grace period. Crossing sets node conditions MemoryPressure or DiskPressure.

solid answer

~50 s

The kubelet samples **eviction signals** — `memory.available`, `nodefs.available`, `nodefs.inodesFree`, `imagefs.available`, `pid.available` — every housekeeping interval and compares them to configured thresholds (absolute or percentage). - **Hard thresholds** (`--eviction-hard`, default `memory.available<100Mi`, `nodefs.available<10%`, `imagefs.available<15%`) trigger **immediately** on breach and kill pods with **no graceful termination** — grace period 0. - **Soft thresholds** (`--eviction-soft`) must remain breached for `--eviction-soft-grace-period` before eviction starts, and victims get up to `--eviction-max-pod-grace-period` to shut down. They exist to give an early, gentler reclaim before the hard line. Breaching a memory signal sets the node condition `MemoryPressure=True`; disk/inode signals set `DiskPressure=True`. Those conditions are also a scheduling input: `MemoryPressure` makes the scheduler avoid placing BestEffort pods, and `DiskPressure` blocks new pods entirely via the default taints. `--eviction-minimum-reclaim` makes each round free a bit more than the threshold so the node does not oscillate, and `--eviction-pressure-transition-period` (default 5m) damps condition flapping.

code

yaml · 18 lines
yaml
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
evictionHard:
  memory.available: "200Mi"
  nodefs.available: "10%"
  imagefs.available: "15%"
  nodefs.inodesFree: "5%"
evictionSoft:
  memory.available: "800Mi"
  nodefs.available: "15%"
evictionSoftGracePeriod:
  memory.available: "1m30s"
  nodefs.available: "2m"
evictionMaxPodGracePeriod: 60
evictionMinimumReclaim:
  memory.available: "300Mi"
  nodefs.available: "2%"
evictionPressureTransitionPeriod: "5m"

go deeper

for a junior

Recall that the kubelet watches free memory and disk against thresholds and evicts pods when they are crossed, and that this shows up as MemoryPressure/DiskPressure on the node.

for a middle

Name the signals, state hard = immediate with zero grace, soft = must persist for a grace period with a capped graceful shutdown, and connect the conditions to the NoSchedule taints.

for a senior

Add the operational layer: image/container garbage collection before eviction, minimum-reclaim and transition-period as anti-flap controls, and how you would tune soft thresholds so hard eviction is rare.

for a principal

Discuss node headroom policy across the fleet — how reserved resources plus eviction thresholds define usable allocatable capacity, and the cost/reliability tradeoff of running nodes closer to full.

## The problem the kubelet is solving A node is a real machine with finite memory, disk, and process IDs. Containers can use more than they asked for. If the machine actually runs out, the Linux kernel's out-of-memory killer or a full filesystem takes down whatever it likes — possibly the container runtime or the kubelet itself, turning one bad pod into a dead node. The kubelet therefore polices the node's resources itself and reclaims them *before* the kernel has to, by killing pods. This is **node-pressure eviction**. ## Eviction signals The kubelet computes a small set of signals from cgroup and filesystem statistics: | Signal | Meaning | |---|---| | `memory.available` | Node memory available for new work (derived from the root cgroup: capacity minus working set) | | `nodefs.available` / `nodefs.inodesFree` | Free space/inodes on the filesystem holding the kubelet root — pod logs and `emptyDir` volumes | | `imagefs.available` / `imagefs.inodesFree` | Free space/inodes on the filesystem holding container images and writable layers, when the runtime uses a separate one | | `pid.available` | Free process IDs on the node | Each signal has a threshold expressed absolutely (`memory.available<500Mi`) or as a percentage of capacity (`nodefs.available<10%`). ## Hard versus soft thresholds **Hard eviction thresholds** are the emergency line. The moment an observed signal is below the threshold, the kubelet begins evicting. There is **no grace period at all**: victim pods are terminated with grace period zero, so applications do not get to flush state or drain connections. The kubelet's built-in defaults are roughly `memory.available<100Mi`, `nodefs.available<10%`, `imagefs.available<15%`, `nodefs.inodesFree<5%`. **Soft eviction thresholds** are the early-warning line, and they are opt-in. A soft threshold has three parts: 1. the threshold itself (`--eviction-soft=memory.available<1Gi`), 2. a **grace period** the breach must persist for before anything happens (`--eviction-soft-grace-period=memory.available=1m30s`), 3. an optional cap on how long victims may take to exit (`--eviction-max-pod-grace-period`). So a soft threshold gives the node a chance to recover on its own (a batch job finishes, a cache shrinks) and, if it does not, evicts pods *gracefully* — preStop hooks and SIGTERM handling run, bounded by the max grace period. A well-tuned node sets soft thresholds meaningfully above the hard ones so most reclaim is graceful and the hard line is rarely reached. ## Node conditions When a memory signal is breached the kubelet reports node condition `MemoryPressure=True`; disk and inode signals report `DiskPressure=True`; PID exhaustion reports `PIDPressure=True`. `Ready` is a separate condition about kubelet health, not pressure. These conditions feed back into scheduling. The kubelet applies matching taints — `node.kubernetes.io/memory-pressure` and `node.kubernetes.io/disk-pressure` (both `NoSchedule`) — so the scheduler stops sending new work. Memory pressure only repels **BestEffort** pods (there is a default toleration for the others), whereas disk pressure repels everything, because any new pod needs image and log space. ## Anti-flapping controls Two settings stop the node from thrashing: - **`--eviction-minimum-reclaim`** — after evicting, keep going until the signal is above the threshold *plus* this margin (e.g. `memory.available=500Mi`). Without it, a node hovering exactly at the threshold evicts one pod, dips again, evicts another, forever. - **`--eviction-pressure-transition-period`** (default 5 minutes) — the kubelet will not clear a pressure condition until the signal has been healthy for this long, which stops the taint and condition from oscillating and the scheduler from repeatedly refilling a sick node. ## Reclaim before eviction For disk signals the kubelet first tries to reclaim without touching workloads: it garbage-collects dead containers and then unused images. Only if that is not enough does it start evicting pods. For memory there is nothing to garbage collect, so eviction begins right away. ## What to say in an interview Name at least two signals, state the hard-versus-soft distinction in terms of *grace period* (0 vs configured) and *dwell time* (immediate vs must persist), name the resulting node conditions and their taints, and mention minimum-reclaim/transition-period as the reason a healthy node does not flap. That is a complete answer.

  • Why would you configure a soft threshold at all when a hard threshold already protects the node?
    A hard threshold kills pods with grace period zero, so applications cannot flush buffers, deregister from load balancers, or finish in-flight requests. A soft threshold set above the hard line gives the node an early, graceful reclaim path plus a dwell time in which a transient spike can resolve itself, so the violent hard eviction becomes the rare fallback rather than the normal mechanism.
  • What is `evictionMinimumReclaim` for?
    It makes each eviction round free more than the bare minimum — the kubelet keeps reclaiming until the signal is above the threshold plus the configured margin. Without it a node sitting exactly at the threshold would evict one pod, immediately dip below again, evict another, and thrash. It buys headroom so the node stabilises after one round.

saying these in an interview costs you the question

  • Saying soft evictions are "less forceful" without naming the grace period and dwell time as the actual difference
  • Believing hard eviction honours the pod's terminationGracePeriodSeconds — it uses 0
  • Thinking eviction thresholds are set per pod or in the pod spec rather than in kubelet configuration per node
  • Confusing the node `Ready` condition with pressure conditions
  • Assuming DiskPressure only counts image space — kubelet root filesystem (logs, emptyDir) is a separate `nodefs` signal

context