skip to content

In Kubernetes, what does an emptyDir volume with medium: Memory give you, and how do sizeLimit and memory limits bound it?

level: middleimportance: should knowfreq 46%

answer

  1. tmpfs, not node disk
  2. smallest of three ceilings
  3. kernel says ENOSPC, not eviction
  4. pages charged to the writer's cgroup
  5. RSS plus tmpfs versus limit

basics

~20 s

medium: Memory mounts the emptyDir as RAM-backed tmpfs, which is fast and never touches the node disk. The kubelet caps it at the smallest of sizeLimit, the pod's memory limit and node allocatable, and every byte written counts as memory use.

solid answer

~40 s

Setting `emptyDir.medium: Memory` makes the kubelet mount a **tmpfs** instead of a directory on the node's disk. Reads and writes are fast, and the data is gone when the pod ends or the node reboots. The tmpfs size is the smallest of the volume's `sizeLimit`, the pod's memory limit (when its containers declare one) and the node's allocatable memory. Once the volume is full, writes fail with *no space left on device*. The trap is accounting: tmpfs pages count toward the memory cgroup of the container that wrote them. An OCR container with a `3Gi` limit, `1.9Gi` of resident memory and `1.2Gi` of rasterised pages in tmpfs is at `3.1Gi`, and it is OOM-killed even though the volume itself is not full. Budget the memory limit to cover the process plus the tmpfs contents.

code

yaml · 21 lines
yaml
apiVersion: v1
kind: Pod
metadata:
  name: ocr-worker-2b8q
spec:
  containers:
  - name: ocr
    image: registry.example.com/ocr/engine:5.3.0
    resources:
      requests:
        memory: 3Gi
      limits:
        memory: 3Gi
    volumeMounts:
    - name: scratch
      mountPath: /scratch
  volumes:
  - name: scratch
    emptyDir:
      medium: Memory
      sizeLimit: 1536Mi

go deeper

for a junior

Recall that medium: Memory turns the emptyDir into RAM-backed tmpfs: fast, never on disk, and gone when the pod ends.

for a middle

Explain the sizing rule (the smallest of sizeLimit, the pod memory limit and node allocatable), that the kernel enforces it with ENOSPC, and that written pages count toward the writing container's memory.

for a senior

Diagnose the OOM kill that is really resident memory plus tmpfs contents. Budget limits and requests for both, and choose disk-backed or memory-backed scratch per workload.

for a principal

Decide when RAM-backed scratch is worth its memory cost at fleet scale, and set platform guidance so teams do not trade disk-pressure incidents for OOM incidents.

## What `medium: Memory` changes An **emptyDir** is a scratch directory whose lifetime is tied to its pod. Its `medium` field has two accepted values: - `""`, the default, puts the directory on the node's default storage, normally the disk that holds the kubelet's data. - `Memory` makes the kubelet mount a **tmpfs**, a Linux filesystem whose contents live in RAM. What tmpfs gives you: - **Speed.** No disk I/O, which helps hot scratch files such as rasterised page images in a document-OCR pipeline. - **No disk residue.** Nothing sensitive is left on the node's disk. Contents are lost when the pod leaves the node or the node reboots. - **No node-disk pressure.** The volume never counts toward the node's disk usage. ## How big the tmpfs is The kubelet sets the tmpfs size when it mounts the volume. It uses the **smallest** of: 1. the node's **allocatable memory**, 2. the **pod's memory limit**, which exists only when the pod's containers declare memory limits, 3. the volume's **`sizeLimit`**, if one is set. The size is passed to the kernel as a mount option, so the **kernel** enforces it. A write past the size fails with `ENOSPC` (*no space left on device*) inside the application, and the pod is not evicted. With no `sizeLimit` and no memory limits, the tmpfs can grow to the node's whole allocatable memory. Disk-backed emptyDir behaves differently. Nothing stops the write itself. Instead, the kubelet periodically measures usage, and when it goes over `sizeLimit` the kubelet **evicts** the pod with a message of the form `Usage of EmptyDir volume "scratch" exceeds the limit "6Gi"`. | | Disk-backed emptyDir | `medium: Memory` emptyDir | |---|---|---| | Backing | Node disk | RAM (tmpfs) | | Enforcing `sizeLimit` | Kubelet evicts the pod | Kernel returns ENOSPC | | Counts against | Ephemeral storage | Memory of the writing container | | Survives node reboot | Not something to rely on | No | | Blocks Cluster Autoscaler scale-down by default | Yes | No | ## The accounting trap: tmpfs is memory Pages written to tmpfs are **charged to the memory cgroup of the container that wrote them**, so they count toward that container's memory limit and toward the pod's. The OCR worker below shows the effect: - the `ocr` container has `limits.memory: 3Gi` - its process holds about `1.9Gi` resident - it has written `1.2Gi` of page images into `/scratch` That totals `3.1Gi`, which is over the `3Gi` limit, so the container is **OOM-killed**, even though `/scratch` (`sizeLimit: 1536Mi`) had room left. From the outside this looks like a memory leak. The fix is to budget for both: - raise the memory limit to cover the process **plus** the worst-case tmpfs contents, or - lower `sizeLimit` so the volume cannot push the container over its limit, or - delete scratch files as soon as they are used, because a deleted tmpfs file frees its memory. Memory requests matter too. The scheduler places pods by requests, so a pod that plans to hold `1.5Gi` in tmpfs should request that memory. Otherwise it can land on a node that has no room for it. ## When to use it - Small, hot scratch data: lock files, sockets, decoded page buffers. - Secrets or intermediate data that should never reach a disk. - Workloads on nodes with small or slow boot disks. Avoid it for large scratch data. Every gigabyte comes out of the memory budget, and on a busy node that raises the chance of OOM kills and memory-pressure evictions. ```yaml apiVersion: v1 kind: Pod metadata: name: ocr-worker-2b8q spec: containers: - name: ocr image: registry.example.com/ocr/engine:5.3.0 resources: requests: memory: 3Gi limits: memory: 3Gi volumeMounts: - name: scratch mountPath: /scratch volumes: - name: scratch emptyDir: medium: Memory sizeLimit: 1536Mi ``` ## Version note Sizing the tmpfs from `sizeLimit` came with the `SizeMemoryBackedVolumes` feature: alpha in 1.20, beta in 1.22, GA in 1.32, and always on in current releases. On clusters from before that feature, the tmpfs was sized to node memory whatever `sizeLimit` said.

  • The pod has no memory limits and no sizeLimit on its memory-backed emptyDir. How large can the tmpfs grow?
    Up to the node's **allocatable memory**, which the kubelet uses as the ceiling when nothing smaller applies. One pod can then fill a large share of the node's RAM with scratch files. That pushes the node toward memory pressure and puts other pods at risk of eviction. Always set `sizeLimit`, a memory limit, or both.
  • Why is a pod whose only emptyDir is memory-backed not treated as blocking Cluster Autoscaler scale-down, while a disk-backed one is?
    Cluster Autoscaler's drain check counts `hostPath` volumes and emptyDirs whose medium is not `Memory` as **local storage**. With `skip-nodes-with-local-storage` at its default of `true`, it will not remove a node running such pods unless the pod carries `cluster-autoscaler.kubernetes.io/safe-to-evict-local-volumes` naming those volumes. A memory-backed emptyDir is not counted, so it does not block scale-down.

saying these in an interview costs you the question

  • Thinks tmpfs usage is free and does not count toward memory limits
  • Expects the kubelet to evict the pod when a tmpfs emptyDir fills
  • Believes sizeLimit is ignored for medium: Memory on current clusters
  • Uses a memory-backed emptyDir for many gigabytes of scratch data
  • Reads an OOM kill caused by tmpfs contents as an application leak