skip to content

On a managed container platform, why does a service that caches files on the container filesystem break?

level: seniorimportance: should knowfreq 52%

answer

  1. How long does that filesystem live
  2. Who else can see what you wrote
  3. Count the replicas, count the caches
  4. There is a quota on that scratch space
  5. docker diff tells you what you wrote

basics

~20 s

Writes land in the container's writable layer, which is private to that container and thrown away when it is replaced. Every replica keeps its own cache, every deploy empties it, and it grows until the platform's ephemeral-storage limit kills the container.

solid answer

~50 s

A container's filesystem is the image's read-only layers plus one thin writable layer on top. That writable layer belongs to a single container: replicas do not share it, and it is discarded when the container is removed -- which on a managed platform happens on every deploy, every scale event and every crash restart. So a thumbnail cache under `/var/cache/thumbs` gives you a cache hit rate that collapses as you scale out, silent data loss at each release, and growth that eventually trips the platform's ephemeral-storage limit and gets the container killed for what looks like an unrelated reason. Some platforms also run the root filesystem read-only, so the write fails outright at startup. Anything that must outlive one container instance belongs in external storage; local disk is for genuinely transient, per-request scratch that you size and clean up.

code

bash · 11 lines
bash
docker run -d --name thumbs-1 thumbnailer:9.4
# ... traffic renders and caches thumbnails ...
docker exec thumbs-1 sh -c 'ls /var/cache/thumbs | wc -l'   # 3412

# a second replica of the same image starts empty
docker run -d --name thumbs-2 thumbnailer:9.4
docker exec thumbs-2 sh -c 'ls /var/cache/thumbs | wc -l'   # 0

# and a replacement of the first one starts empty too
docker rm -f thumbs-1 && docker run -d --name thumbs-1 thumbnailer:9.4
docker exec thumbs-1 sh -c 'ls /var/cache/thumbs | wc -l'   # 0

go deeper

for a junior

Know that a container's filesystem starts as a copy of the image every time and that whatever you write is thrown away when the container is replaced. Keep anything that must persist outside the container.

for a middle

Explain the mechanics: image layers are read-only, one thin writable layer sits on top, it is private to that container, and it is created and destroyed with it. Explain why that makes a local cache useless across replicas.

for a senior

Show that you recognise the symptoms in production -- a hit rate that falls as you scale out, a latency spike after every release, a container killed for disk with no application error -- and describe the migration to external storage without hand-waving the transitional state.

for a principal

Own the constraint as a platform rule rather than a code review comment: images that hold no local state, a default path to shared storage teams do not have to design themselves, and an explicit exception process for the rare workload that genuinely needs local disk.

### What the container filesystem actually is A running container sees a union of the image's read-only layers with a single thin writable layer stacked on top. Reads fall through to the image; the first write to an existing file copies it up into the writable layer and modifies the copy. That layer is created when the container is created and destroyed when the container is removed. It is not shared with any other container, not even another instance of the same image, and it is not a volume -- volumes exist precisely because this layer is unsuitable for anything you want to keep. On a single developer machine none of this is obvious, because the container often lives for days. On a managed platform, container lifetime is measured in hours and every deploy replaces the whole set. ### The four failures, concretely Take a Spring Boot photo-thumbnail service that renders a resized image on first request and stores it under `/var/cache/thumbs`, keyed by source hash. **The cache is per-replica, so it does not work.** With six replicas behind a round-robin router, a given source image is rendered up to six times, once per replica that happens to receive a request for it. The hit rate you measured with one instance locally does not survive scaling out, and the CPU cost you thought you had amortised comes back roughly in proportion to replica count. Worse, if a client can see the rendered result's metadata, two requests can return objects created seconds apart with different timestamps -- an inconsistency with no obvious cause. **Every deploy is a full flush.** After a release, all 3,412 cached thumbnails are gone and the service faces a thundering herd of re-renders on top of a cold JVM. If the deploy also triggered an autoscale, the herd arrives while the fleet is at its smallest. This is the classic "our p99 spikes for ten minutes after every release and we cannot explain it" incident. **It counts against a disk quota you did not think about.** Managed platforms cap the ephemeral storage a container may consume, and an unbounded cache walks straight into that cap. The container is then terminated for a disk reason, restarts clean, refills, and is terminated again -- a slow crash loop whose period is however long it takes to fill the quota. Because the app never logs an error, the cause is easy to miss. **The write may not be allowed at all.** Platforms increasingly run application containers with a read-only root filesystem and grant one small writable path such as a `tmpfs` at `/tmp`. An image that writes elsewhere fails at startup with a permissions or read-only-filesystem error -- and the same image works perfectly on the developer's laptop, which imposes no such restriction. ### The same argument applies to more than caches Uploaded files parked on local disk before processing vanish mid-flight when the container is replaced. Session state on disk forces sticky routing and breaks on any replacement. An embedded database file gives each replica its own divergent copy of the truth. Even a lock file used to make a scheduled job run once is worthless -- it is per-container, so N replicas run the job N times, which is how a nightly cleanup ends up executing six times in parallel. ### What local disk is still good for Transient, per-request scratch is fine, and pretending otherwise leads to silly designs. Decoding a source image to a temporary file, writing it out, streaming the result and deleting it is exactly what `/tmp` is for. The rules are: assume nothing survives the request, size the working set against the platform's ephemeral-storage allowance, delete in a `finally` block rather than trusting a cleanup timer, and prefer memory when the object is small enough that the round trip through the filesystem buys nothing. ### The fix Push anything that must outlive one container instance out of the container: rendered thumbnails into object storage fronted by a CDN, session and cache entries into a shared cache service, uploads written straight through to durable storage rather than staged locally, and scheduled work coordinated by something all replicas can see rather than by a file only one of them can see. The rewrite is usually smaller than teams fear, because the code path already has a "cache miss" branch -- what changes is where the store lives. ### How to check before it bites you Run the image, exercise it, and diff what it wrote: `docker diff` lists every path added, changed or deleted relative to the image layers, and anything in that list which is not obviously scratch is a question to answer. Then start the container a second time from the same image and confirm the service behaves identically on an empty filesystem, because that is the state every replica the platform starts will be in.

  • Could you keep the local cache and just mount a volume for it?
    On a single host, yes -- that is what volumes are for. On a managed platform it usually is not an option, because containers are placed on hosts you do not choose and a host-local volume does not follow a replacement to a different machine. Even when shared storage is offered, a cache is rarely worth the coupling; external object storage or a shared cache service fits the access pattern better.
  • Is local disk ever the right answer inside a container?
    For per-request scratch, yes: decode, transform, stream, delete. Size the working set against the platform's ephemeral-storage allowance, clean up in a `finally` block rather than trusting a timer, and never let a request's correctness depend on a file written by an earlier request.
  • Image size is often described as something you pay for at deploy time. How does that connect to this?
    Both are consequences of the image being the entire deploy unit. A 618 MB JRE image is fetched again on every new instance, so the size becomes latency in exactly the moments you are already short of capacity -- a scale-out or a restart storm. Trimming the image and keeping state outside it attack the same problem: making a fresh instance cheap to create.
  • A nightly cleanup job runs inside the service container, guarded by a lock file. What goes wrong?
    The lock file is in the container's private writable layer, so each replica sees an unlocked state and every replica runs the job. Coordination has to happen somewhere all replicas can see -- a shared lock in a datastore, or moving the job out of the serving container entirely.

The writable layer is a whiteboard in a hot-desk booth, not a filing cabinet. It is genuinely useful while you are sitting there, nobody in the next booth can read it, and it is wiped the moment you stand up.

saying these in an interview costs you the question

  • Thinks replicas of one image share a filesystem
  • Expects the writable layer to survive a redeploy
  • Uses a container-local lock file for a scheduled job
  • Ignores the platform's ephemeral-storage limit
  • Stages uploads on local disk before processing
  • Assumes a volume is always available on a PaaS

context