skip to content

After a 210-node multi-tenant Kubernetes cluster starts requiring readOnlyRootFilesystem: true in container securityContext, a ticket-booking checkout crashes at startup. How do you make the app ready for it?

level: seniorimportance: nice to knowfreq 32%

answer

  1. volumes keep their own mode
  2. inventory hidden write paths
  3. scratch goes to emptyDir
  4. tmpfs pages cost pod memory
  5. mount hides image contents

basics

~10 s

Find every path the app writes, redirect those writes to explicitly mounted emptyDir volumes (disk-backed by default, memory-backed only when small and justified), set sizeLimit, and point the runtime's temp and cache directories there.

solid answer

~40 s

`readOnlyRootFilesystem: true` mounts the container's image filesystem read-only, so any write outside a mounted volume fails. The checkout usually writes to a temp directory, a web server's working directory, a cache, or a PID or lock file. I list those paths by running the image read-only in a test environment, then mount an `emptyDir` at each one, or one `emptyDir` with the runtime's temp and cache settings pointed into it. Each volume gets a `sizeLimit` so a runaway writer is evicted instead of filling the node. I keep `medium: Memory` for small, hot scratch data only, because its pages count against the pod's memory and can trigger an OOM kill. Anything that must survive a restart goes to a PersistentVolumeClaim or an external store, not an `emptyDir`.

code

yaml · 33 lines
yaml
apiVersion: v1
kind: Pod
metadata:
  name: checkout
spec:
  securityContext:
    runAsNonRoot: true
    fsGroup: 10001
  containers:
    - name: checkout
      image: registry.example.com/booking/checkout:4.12.3
      env:
        - name: JAVA_TOOL_OPTIONS
          value: -Djava.io.tmpdir=/scratch/tmp
      securityContext:
        readOnlyRootFilesystem: true
        allowPrivilegeEscalation: false
      resources:
        limits:
          memory: 3Gi
      volumeMounts:
        - name: scratch
          mountPath: /scratch
        - name: hot-cache
          mountPath: /var/cache/checkout
  volumes:
    - name: scratch
      emptyDir:
        sizeLimit: 1Gi
    - name: hot-cache
      emptyDir:
        medium: Memory
        sizeLimit: 64Mi

go deeper

for a junior

Know that readOnlyRootFilesystem blocks writes to the image filesystem but not to mounted volumes, and that emptyDir is the usual scratch space.

for a middle

Explain emptyDir lifetime, the difference between disk and memory media, and how sizeLimit caps usage for each.

for a senior

Show the procedure: inventory writes, delete the unnecessary ones, mount bounded scratch, budget tmpfs in the memory limit, and handle directories the mount would hide.

for a principal

Decide how the platform rolls the requirement out: templates with scratch volumes built in, an exception process, and default sizeLimit values that protect other tenants.

## What the setting actually does `securityContext.readOnlyRootFilesystem` is a container-level field whose default is `false`. When set to `true`, the runtime mounts the container's root filesystem, the image layers plus the usual writable layer, **read-only**. Volumes mounted into the container keep their own mode, so a mounted `emptyDir` stays writable. The field cannot be set for Windows pods. Platform teams require it because an attacker who gets code execution cannot drop binaries or modify application files in place, and because it stops apps from silently filling node disk through their writable layer. Note that the built-in Pod Security Standards do not check this field; a cluster that requires it enforces it through its own admission policy. ## Why the checkout crashes Most applications write somewhere without anyone having decided so: - the runtime's temp directory (`/tmp`), used by JVM features, upload handling and native libraries; - an embedded web server's working directory for multipart uploads or compiled pages; - local caches, such as a template cache or a downloaded certificate bundle; - PID files, lock files or a SQLite file for a local queue; - log files that should have been stdout in the first place. Each of those write attempts now fails with a read-only filesystem error, often during startup, so the pod lands in `CrashLoopBackOff`. ## The readiness procedure 1. **Inventory the writes.** Run the image read-only outside the shared cluster and collect every failing path. Finding them is a container-level exercise; the output you need is a short list of directories. 2. **Remove the ones that should not exist.** Logs go to stdout; durable data goes to a PersistentVolumeClaim or an external store. 3. **Mount scratch space for the rest.** Mount an `emptyDir` at each remaining directory, or a single `emptyDir` with the application configured to use it, for example by setting the temp directory through an environment variable or a JVM system property. 4. **Bound it.** Set `sizeLimit` on every `emptyDir` so a runaway writer gets the pod evicted instead of exhausting node ephemeral storage for other tenants. 5. **Test the shutdown and restart path.** An `emptyDir` is created empty when the pod is scheduled onto a node and deleted when the pod is removed; a container restart inside the same pod keeps it. The app must not assume anything survives a reschedule. ## Disk-backed versus memory-backed emptyDir | Aspect | `medium: ""` (default) | `medium: Memory` | |---|---|---| | Backing | node's local storage | tmpfs in RAM | | Counted against | pod ephemeral storage | the pod's memory usage | | Size cap | `sizeLimit` enforced by eviction | lower of `sizeLimit` and the sum of container memory limits | | Speed | disk speed | memory speed | | Risk | node disk pressure if unbounded | OOM kill when scratch plus heap exceed the limit | On a checkout with a 3Gi memory limit and a JVM already using most of it, a 512Mi memory-backed `/tmp` that fills during a sale can push the container over its limit. Choose `Memory` only for small, performance-sensitive files, and budget it in the memory limit. ## Things that still bite - **Pre-populated directories**: mounting an `emptyDir` over a path hides whatever the image put there. If the app ships default files in that directory, copy them in with an init container or move them elsewhere. - **Several writers**: two containers in the pod sharing one `emptyDir` must agree on file names. - **Ownership**: with a non-root user, the mounted directory must be writable by that user; `fsGroup` in the pod security context is the usual lever. ## How to tell the fix worked - the pod reaches Ready with the policy enforced, not just in a permissive namespace; - a load test that exercises uploads and cache warm-up completes without read-only errors in the logs; - the scratch volumes stay well under their `sizeLimit` at peak, so no eviction is waiting to happen. ## Why this is worth doing early Retro-fitting read-only root filesystems across a multi-tenant platform is slow because every team finds its own hidden writes. Treating "runs read-only with declared scratch volumes" as part of the migration checklist makes the platform rule a non-event.

  • Does data in an emptyDir survive a container restart in the same pod?
    Yes. An `emptyDir` belongs to the pod, not the container, so a crashed container that the kubelet restarts sees the same files. The volume is deleted when the pod is removed from the node, so a reschedule or rollout starts empty. Anything the checkout must keep across pods belongs in a PersistentVolumeClaim or an external store.
  • You mount an emptyDir at a directory where the image ships default templates. What happens?
    The mount hides the image's directory contents, so the app starts with an empty directory and fails to find its templates. Either keep the templates in a separate read-only path and mount scratch elsewhere, or copy them into the volume with an init container before the main container starts.

saying these in an interview costs you the question

  • Believes readOnlyRootFilesystem also makes mounted emptyDir volumes read-only
  • Says the Pod Security restricted profile already forces readOnlyRootFilesystem
  • Treats a memory-backed emptyDir as free RAM outside the pod's memory usage
  • Leaves emptyDir volumes without sizeLimit on a shared multi-tenant cluster
  • Stores durable booking data in an emptyDir because it survives container restarts