skip to content

A document-OCR pipeline on a Kubernetes cluster that autoscales between 20 and 80 nodes needs 140Gi of per-pod scratch. When would you choose a generic ephemeral volume over emptyDir, and what can go wrong?

level: seniorimportance: nice to knowfreq 24%

answer

  1. scratch off the node disk
  2. pod name dash volume name
  3. owner reference, then garbage collection
  4. provisioning adds to scale-up time
  5. quota, attach limits, reclaim policy

basics

~20 s

Choose a generic ephemeral volume when per-pod scratch is too large for the node disk. A storage driver provisions a PVC for each pod, and it is deleted with the pod. The costs are provisioning latency, name collisions, quota and attach limits.

solid answer

~40 s

An `emptyDir` of 140Gi lands on the node's disk. It competes with images and logs, risks ephemeral-storage evictions, and by default blocks Cluster Autoscaler from removing the node. A generic ephemeral volume (`volumes[].ephemeral.volumeClaimTemplate`) makes kube-controller-manager's ephemeral volume controller create a PVC named `<pod name>-<volume name>`, owned by the pod. The storage driver provisions a volume sized to the claim, and garbage collection deletes the PVC when the pod goes away. What goes wrong: provisioning and attach add minutes on top of an already slow **19-minute** node scale-up. An existing PVC with the same name that the pod does not own blocks the pod from starting. The claims count against the namespace's ResourceQuota and against per-node volume attach limits. And if the StorageClass does not delete the backing volume, orphaned disks pile up.

code

yaml · 24 lines
yaml
apiVersion: v1
kind: Pod
metadata:
  name: ocr-worker-5c9d2
spec:
  containers:
  - name: ocr
    image: registry.example.com/ocr/engine:5.3.0
    volumeMounts:
    - name: scratch
      mountPath: /scratch
  volumes:
  - name: scratch
    ephemeral:
      volumeClaimTemplate:
        metadata:
          labels:
            app: ocr-worker
        spec:
          accessModes: ["ReadWriteOnce"]
          storageClassName: ocr-scratch
          resources:
            requests:
              storage: 140Gi

go deeper

for a junior

Recall that a generic ephemeral volume is a PVC template inside the pod spec: a real provisioned volume that is deleted together with the pod.

for a middle

Explain the ephemeral volume controller, the pod-name-dash-volume-name claim name, the owner reference and garbage collection, and why an unrelated claim with that name blocks the pod.

for a senior

Weigh node-disk emptyDir against provisioned scratch on a real autoscaling fleet: scale-down blocking, provisioning latency, attach limits, quota and reclaim policy, and back the choice with measurements.

for a principal

Decide which scratch-storage options the platform offers and at what cost to each team, balancing node shapes, driver latency and storage spend across bursty batch workloads.

## The problem A document-OCR pipeline runs one worker pod per large scanned archive. Each worker unpacks and rasterises up to **140Gi** of pages, and the scratch data is worthless once the pod finishes. The cluster autoscales between **20 and 80 nodes**, and a scale-up currently takes about **19 minutes** from Pending pods to Ready nodes. The question is where those 140Gi should live. ## Option 1: disk-backed emptyDir `emptyDir` puts the scratch on the node's own disk. It is simple and has no provisioning step, but at this size it causes trouble: - **Node disk contention.** Scratch shares the disk with container images, logs and other pods. Without `sizeLimit` and `resources.requests.ephemeral-storage`, a few workers can push the node into disk pressure, and the kubelet starts evicting pods. - **Node sizing.** Every node that may run a worker needs a boot disk big enough for several workers, even when it runs none. - **Autoscaler scale-down.** Cluster Autoscaler treats pods with a non-memory emptyDir, or a hostPath, as having **local storage**. With `skip-nodes-with-local-storage` at its default of `true`, it will not remove their nodes unless the pod carries `cluster-autoscaler.kubernetes.io/safe-to-evict-local-volumes` listing those volumes. A fleet of workers can hold the cluster near 80 nodes. ## Option 2: generic ephemeral volume A **generic ephemeral volume** embeds a PVC template in the pod spec: ```yaml apiVersion: v1 kind: Pod metadata: name: ocr-worker-5c9d2 spec: containers: - name: ocr image: registry.example.com/ocr/engine:5.3.0 volumeMounts: - name: scratch mountPath: /scratch volumes: - name: scratch ephemeral: volumeClaimTemplate: metadata: labels: app: ocr-worker spec: accessModes: ["ReadWriteOnce"] storageClassName: ocr-scratch resources: requests: storage: 140Gi ``` The mechanics: 1. The **ephemeral volume controller** in kube-controller-manager sees the pod and creates a PVC named `ocr-worker-5c9d2-scratch`, that is `<pod name>-<volume name>`. 2. The PVC's `ownerReferences` point to the pod, with `controller: true`. 3. The usual provisioning path runs: the StorageClass's driver creates a volume, and it is attached and mounted on the pod's node. 4. When the pod is deleted, the **garbage collector** deletes the PVC. The backing volume then follows the PV's reclaim policy, which is `Delete` by default for dynamically provisioned volumes. The scratch is now **off the node disk**, it is sized per pod, and it does not count as local storage for Cluster Autoscaler. ## What can go wrong | Failure | Cause | Mitigation | |---|---|---| | Workers take far longer than 19 minutes to start during a burst | Provisioning and attaching each volume comes on top of node scale-up | Measure the driver's provision and attach time; use a class that binds on first consumer so volumes are created where pods land | | Pod never starts | A PVC with the same name exists and is not owned by this pod | Delete the stray claim; avoid reusing pod names with leftover claims | | Pod rejected at creation | `<pod name>-<volume name>` is not a valid PVC name, for example too long | Shorten volume or pod names | | Pods Pending on nodes with free CPU | Per-node volume attach limit reached | Size node shapes for attach counts; spread workers | | Namespace suddenly refuses new pods | ResourceQuota on PVC count or `requests.storage` | Budget quota for peak concurrency: 140Gi per worker | | Cost creeps up | The class's reclaim policy keeps volumes | Use `Delete` for scratch classes | The `volumeClaimTemplate` is copied once. Kubernetes does not update the PVC afterwards, so editing a running pod's template changes nothing. ## How to decide - **Small, short-lived scratch** (a few GiB): use emptyDir with `sizeLimit` and ephemeral-storage requests. - **Large, variable scratch on an autoscaling fleet**: use a generic ephemeral volume, and pick a storage class with good provisioning latency. - **Fast local NVMe that is worth the node coupling**: use a driver that exposes local disks as volumes, still through a generic ephemeral volume, rather than `hostPath`. - **Data that must outlive the pod**: neither. Use a regular PVC or an external store.

  • Why not simply mount the node's local NVMe into each OCR pod with hostPath?
    A `hostPath` volume exposes an existing node path. Its `type` (`Directory`, `DirectoryOrCreate`, `File`, `Socket`, `BlockDevice`, and so on) only checks what is at that path. Nothing is cleaned up when the pod ends, pods on the same node can see each other's files, and a writable host mount lets a container reach node files. Pod Security admission policies commonly forbid it. hostPath also blocks Cluster Autoscaler scale-down by default. Expose local disks through a storage driver instead.
  • A worker pod stays in ContainerCreating and its events mention a PVC that already exists. What happened, and how do you fix it?
    A claim named `<pod name>-<volume name>` already exists but is not owned by this pod, for example one left over from an earlier pod with the same name. Kubernetes will not use a claim it cannot tie to this pod, to avoid mounting unrelated data. The pod waits until that claim is deleted, or an administrator deliberately adds an owner reference to the pod.

saying these in an interview costs you the question

  • Assumes generic ephemeral PVCs block Cluster Autoscaler scale-down like emptyDir
  • Expects the generated PVC to survive the pod like a StatefulSet claim
  • Ignores the provisioning and attach time added on top of node scale-up
  • Reaches for hostPath on local NVMe as the simple scratch answer
  • Forgets ResourceQuota counts the generated claims and their storage
  • Edits the pod's volumeClaimTemplate expecting the live PVC to change