skip to content

Pods on a Kubernetes node are being killed with the message "The node was low on resource: ephemeral-storage". What counts toward a container's ephemeral storage usage, and how would you stop this from recurring?

level: seniorimportance: should knowfreq 45%

answer

  1. ephemeral-storage = writable layer + emptyDir + pod logs
  2. nodefs vs imagefs; also inodesFree
  3. image/container GC runs before eviction
  4. pod-level limit → evict that pod alone, polled not enforced by cgroup
  5. emptyDir medium: Memory counts as memory, not disk

basics

~20 s

Ephemeral storage is the container writable layer, emptyDir volumes, and container logs on the node's disk. The kubelet evicts pods when nodefs/imagefs run low, or when a pod exceeds its ephemeral-storage limit. Fix with log rotation, sized emptyDir, ephemeral-storage requests/limits, and image garbage collection.

solid answer

~50 s

**Ephemeral storage** = the container's writable layer, any `emptyDir` volumes, and the pod's log files on the node — everything written to local disk that is not a persistent volume. Two distinct kills produce this message: 1. **Node-level**: the `nodefs.available` or `imagefs.available` eviction signal is breached, `DiskPressure=True`, and the kubelet evicts pods — ranked by usage over their ephemeral-storage request. 2. **Pod-level**: a pod exceeds its own `limits.ephemeral-storage`; the kubelet evicts that single pod regardless of node health. Remediation, in order of leverage: - Set `requests`/`limits` for `ephemeral-storage` so the scheduler accounts for it and one pod cannot fill the disk. - Use `emptyDir.sizeLimit` to cap scratch volumes (including `medium: Memory`, which counts against memory). - Fix logging: bound container log rotation (`containerLogMaxSize`/`containerLogMaxFiles`), stop apps writing large files inside the container, ship logs off-node. - Tune image garbage collection (`imageGCHighThresholdPercent`) and give nodes a bigger or separate `imagefs`. - Mount a PVC for genuinely large working data.

code

yaml · 16 lines
yaml
spec:
  containers:
    - name: worker
      image: registry.example.com/worker:1.9
      resources:
        requests:
          ephemeral-storage: "1Gi"
        limits:
          ephemeral-storage: "4Gi"
      volumeMounts:
        - name: scratch
          mountPath: /tmp/work
  volumes:
    - name: scratch
      emptyDir:
        sizeLimit: 2Gi

go deeper

for a junior

Know that ephemeral storage means the container's writable layer, emptyDir, and logs, and that filling the node's disk gets pods evicted.

for a middle

Separate the node-level DiskPressure signal from the per-pod ephemeral-storage limit, and know that requests/limits and emptyDir sizeLimit exist.

for a senior

Drive the full diagnosis — describe pod/node, df and df -i on the node, locate the consumer — then contain it with requests/limits, log rotation, image GC tuning, and PVCs for real data.

for a principal

Set platform defaults: mandatory ephemeral-storage requests via LimitRange, node disk sizing and a separate imagefs, centralised log shipping, and admission policy so no workload can take a node down through its root disk.

## What "ephemeral storage" means Every container gets a **writable layer** on top of its read-only image layers. Anything the process writes that is not on a mounted volume lands there and disappears when the container is deleted. Kubernetes groups three things under the resource name `ephemeral-storage`: 1. the container **writable layer**, 2. **`emptyDir` volumes** (disk-backed by default), 3. the pod's **container logs** written by the kubelet on the node filesystem. What does *not* count: PersistentVolumes, `configMap`/`secret`/`downwardAPI` volumes, `hostPath` mounts (they consume the host's disk but are not attributed to the pod), and read-only image layers, which are accounted to the image filesystem rather than to the pod. ## Two filesystems, two signals The kubelet tracks up to two filesystems: - **`nodefs`** — where the kubelet root lives: `emptyDir` volumes and pod logs. - **`imagefs`** — where the container runtime stores images and writable layers, when the runtime is configured with a separate filesystem. If there is no separate one, everything is `nodefs`. Each has an `available` and an `inodesFree` signal. Breaching either raises `DiskPressure=True`, which taints the node `node.kubernetes.io/disk-pressure:NoSchedule` so no new pods land there. Before evicting anything, the kubelet tries **garbage collection**: it removes dead containers and then unused images, driven by `imageGCHighThresholdPercent` (start collecting) and `imageGCLowThresholdPercent` (stop). Only when that fails to bring the signal back does it evict pods. Ranking works like memory: pods whose ephemeral-storage usage exceeds their `requests.ephemeral-storage` go first, then by pod priority, then by size of the overage. A pod with no ephemeral-storage request is over request by definition. ## The other kill: per-pod limits Separately from node pressure, if a pod declares `resources.limits.ephemeral-storage` and its combined usage across writable layers, `emptyDir`s, and logs exceeds it, the kubelet evicts **that pod alone**, even on a perfectly healthy node. Detection is periodic (the kubelet polls volume and filesystem stats, typically every ~10–15 seconds), so a burst can briefly exceed the limit before the eviction lands — this is not a hard cgroup ceiling like memory. An `emptyDir` with `sizeLimit` behaves similarly: exceed it and the pod is evicted. ## Diagnosing it 1. `kubectl describe pod` — the eviction message names `ephemeral-storage` and often the offending container and its usage versus request. 2. `kubectl describe node` — confirm `DiskPressure` and read the events (`EvictionThresholdMet`, `ImageGCFailed`). 3. On the node: `df -h` and `df -i` (inodes can run out with disk free), then find the consumer — `/var/lib/kubelet/pods/*/volumes/kubernetes.io~empty-dir` for scratch space, `/var/log/pods` for logs, and the runtime's image directory. 4. `kubectl exec` plus `du -xh --max-depth=1 /` inside a suspect container to find files written into the writable layer. ## Fixing it durably **Declare it.** `requests.ephemeral-storage` makes the scheduler reserve disk the way it reserves memory, so the node stops being overcommitted on disk; `limits.ephemeral-storage` converts "one pod fills the node and everyone is evicted" into "the offending pod is evicted". That containment is the single biggest win. **Cap scratch space.** `emptyDir.sizeLimit: 1Gi` bounds a temp directory. Note `medium: Memory` emptyDirs are tmpfs and count against the pod's *memory*, not disk — a common surprise that turns a disk problem into an OOM. **Fix logging.** An app that logs to stdout at high volume fills `nodefs` through the kubelet's log files; set `containerLogMaxSize` and `containerLogMaxFiles` in kubelet config. An app that writes its own log file inside the container fills the writable layer invisibly — redirect it to stdout or to a sized volume. Ship logs off-node so retention is not the node's problem. **Manage images.** Nodes that pull many large images churn `imagefs`; tune GC thresholds, provision a bigger disk, or give the runtime a dedicated filesystem so image churn cannot evict workloads through `nodefs`. **Use real storage.** If a workload genuinely needs tens of gigabytes of working data, that is a PersistentVolumeClaim (or a generic ephemeral volume, which gives per-pod dynamically provisioned storage that is deleted with the pod), not the node's root disk. ## Interview framing Define the three things that count, separate the node-level signal from the per-pod limit, describe the diagnosis path (`describe pod` → `describe node` → `df`/`df -i` on the node), and land on containment: requests/limits, emptyDir sizeLimit, log rotation, image GC, and PVCs for real data.

  • A node shows plenty of free bytes but pods are still evicted for disk. What would you check?
    Inodes. The kubelet also watches `nodefs.inodesFree` and `imagefs.inodesFree`, and a workload that creates millions of tiny files — session files, cache entries, per-request temp files — can exhaust the inode table while gigabytes remain free. `df -i` shows it; the fix is the same containment plus stopping the file churn.
  • Does setting `limits.ephemeral-storage` enforce the cap the way a memory limit does?
    No. Memory limits are enforced synchronously by cgroups, so a write past the limit fails or the container is OOMKilled instantly. Ephemeral storage is measured by the kubelet polling filesystem statistics every few seconds, so a pod can briefly exceed its limit before the kubelet notices and evicts it. It is containment, not a hard ceiling.

saying these in an interview costs you the question

  • Counting PersistentVolume usage as ephemeral storage
  • Not knowing that container logs written by the kubelet count against the pod
  • Assuming an ephemeral-storage limit is enforced instantly like a memory cgroup limit
  • Overlooking inode exhaustion when bytes look fine
  • Thinking `emptyDir: {medium: Memory}` consumes disk — it is tmpfs and counts against memory

context