skip to content

questions

4

Why does a file a container writes inside itself vanish when that instance is replaced by a new one?

level: juniorimportance: must knowfreq 84%

answer

  1. ask where the write actually landed
  2. the image stays read-only
  3. one thin surface per instance
  4. created empty at start
  5. discarded with the instance, never archived

basics

~20 s

The write landed in that instance's own thin writable layer, a private area created empty on top of the read-only image. Replacing the instance discards the layer, and nothing written only there is kept or recoverable.

solid answer

~50 s

An image is read-only, and many instances share it, so a running container is given a **thin writable layer** of its own: a private, initially empty area stacked over the image's content. Every file the process creates or modifies is captured there, which is why an ordinary path like `/var/lib/records` looks like a normal filesystem from the inside. That layer belongs to one instance. It is created empty when the instance starts and thrown away when the instance is destroyed — no copy is archived, and there is nothing to recover from afterwards. So a data store that wrote to a plain path inside the container was writing to scratch space the whole time, and the replacement simply threw the scratch away. Two replicas of the same workload also have two separate layers, which is why one replica never sees what the other wrote.

code

pseudocode · 10 lines
pseudocode
on start(instance):
    instance.root = image.readOnlyContent + newEmptyWritableLayer()
    # every write from the process is captured in the writable layer
    process.write("/var/lib/records/data.db")   -> instance.writableLayer

on replace(instance):
    discard(instance.writableLayer)             # data.db goes with it
    newInstance = start(instance.spec)
    # newInstance.root is the same image plus an empty layer
    newInstance.read("/var/lib/records/data.db") -> not found

go deeper

for a junior

Remember one sentence: a container writes to a thin private layer over a read-only image, and that layer is thrown away with the instance. Any path inside the container is scratch space unless something was deliberately attached there.

for a middle

Be able to explain why the layer exists at all — the image must stay read-only so many instances can share it — and to separate a restart in place, which keeps the layer, from a replacement, which discards it.

for a senior

Show the diagnostic habit: for each write a workload performs, say where it lands and how long it must survive. Expect to be asked why a crash dump or a cache written at an internal path was missing when someone went looking.

for a principal

The trade-off you own is which workloads are allowed to write anything meaningful inside the instance at all, and what the platform guarantees by default. A default that silently accepts durable-looking writes is a data-loss incident waiting on a routine rollout.

## The image is read-only, so something else has to take the writes A container starts from an **image**: a set of files the platform treats as strictly read-only. That is what makes an image safe to share — many instances can start from the same one, on the same host or on different hosts, and none of them can modify it for the others. Real processes write anyway. They write temporary files, caches, upload spools, lock files, database files, application log files. To make that work without touching the image, the runtime gives each instance a **thin writable layer**: a private area, empty at start, presented on top of the image's content so the process sees one ordinary filesystem. Every create, modify and delete the process performs is captured in that private area rather than in the image. Three properties follow, and between them they are the whole of this subject: - **It is per instance.** Two replicas of the same workload have two different writable layers. A file one replica wrote does not exist for the other, and neither one can see it. - **It starts empty.** A new instance begins with only what the image contained. Nothing carries forward from the instance it replaced. - **It is discarded with the instance.** The platform does not archive it, back it up, or hand it to the replacement. Once the instance is gone, so is everything that was only there. ## Why the data "disappeared" The canonical version of this story is a data store started inside a container with its files left at whatever path it uses by default. It runs, it accepts writes, and everything works — because a writable layer is a real filesystem and behaves like one. Then the instance is replaced: a new build is rolled out, the host is drained, the process is moved, the workload is scaled down and back up. The replacement starts from the same image, and the image never contained those records. The new instance is therefore correct and empty, and the old layer is already gone. Nothing failed. The write went exactly where the model says it goes. The mistake was treating a path inside the container as storage rather than as scratch space with the lifetime of one instance. ## Restart in place is not the same as replacement This is where candidates trip, because the observed behaviour differs: | event | what happens to the writable layer | |---|---| | the process deletes a file it wrote | the bytes are released back to the host's shared disk pool | | the same container is restarted in place | the same layer is still there, with its contents | | the instance is replaced by a new one | the old layer is discarded; the new one starts empty | | the workload is scaled up | each added instance gets its own empty layer | | the image is pulled on another host | the image is identical, and no writable layer travels with it | Platform designs genuinely differ here: a single-host runtime typically restarts the same container in place, keeping its layer, while a cluster scheduler usually replaces the instance — often on another machine entirely — which is a fresh layer every time. That is why "it kept my files when I restarted it on my laptop" and "it lost them in production" can both be true reports of the same image. ## How to reason about it, and where the fix lives The useful habit is to ask, for every write a workload performs, **where does this land and how long does it need to survive?** Three answers cover almost everything: 1. **It only needs to live as long as this instance.** The writable layer will do, though a declared and sized scratch area is the better home for anything large. 2. **It must survive the instance.** It needs storage that is separate from the container and outlives it — a different subject with its own mechanisms, and the whole point of this one is that a plain path inside the container is not it. 3. **It is diagnostic output.** Anything written to a file inside the instance is gone the moment the instance is, which is why a crash dump left at an internal path is usually missing by the time someone goes looking. The layer is also worth understanding for the opposite reason: because it accepts writes silently and cheaply, a process can fill it with temporary files for hours, and nothing in the image's size hints at how much host disk the running instance is consuming. ## What this is not It is not a cache the platform manages for you, not a snapshot, and not a place the replacement can read. And the write never modified the image: the image is exactly the same bytes for the next instance and for every other host that pulls it.

  • A teammate restarted the same container and the file was still there. Does that contradict the rule?
    No. A restart in place reuses the same instance and therefore the same writable layer, so its contents survive. Replacement creates a new instance with a new, empty layer. Platforms differ in which of the two they do by default, which is why the same image can look like it keeps files on one setup and loses them on another.
  • Two replicas of one workload run from the same image. Do they share a writable layer?
    No — one per instance. Each replica has its own private layer, so a file written by one is invisible to the other. That is why a cache kept at an internal path gives inconsistent hit rates across replicas, and why "it works on the instance I shelled into" tells you nothing about the others.
  • Does writing a file change the image the instance was started from?
    Never. The image is read-only and identical for every instance and every host that pulls it; the write is captured in that one instance's writable layer. Building a new image from a running instance's state is a separate deliberate act, not something a write performs.

The image is a printed manual handed to every reader unchanged; the writable layer is a whiteboard wheeled in beside it. The manual goes to the next reader untouched, and the whiteboard is wiped before it leaves the room.

saying these in an interview costs you the question

  • Thinks the write goes into the image and becomes part of it
  • Believes a container keeps its files simply because it has a filesystem
  • Says the lost data can be recovered from the replaced instance afterwards
  • Confuses restarting the same container with getting a new instance
  • Assumes the platform quietly backs the writable layer up somewhere
open as a page

A worker unpacks each upload into a temporary directory inside the container — why declare a scratch area for that instead?

level: middleimportance: should knowfreq 55%

basics

~20 s

A declared scratch area gives the intermediate files a named path, an explicit size ceiling and a chosen backing — memory or disk — that the platform can account for. The writable layer gives none of that: it is unbounded and drawn silently from shared host storage.

open as a page

A long-running worker's writable layer has grown to tens of gigabytes of temporary and rotated files — what actually clears it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Two things clear it: the process deleting files it wrote, which releases those bytes immediately, and replacing the instance, which discards the layer whole. Restarting the same container in place keeps it, and nothing inside prunes it on a schedule.

open as a page

A team backs a worker's scratch directory with memory instead of disk — what does that speed cost them?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Every byte written into a memory-backed scratch area is charged against the workload's own memory budget, as if the process had allocated it. The area's ceiling must therefore fit inside that budget, on top of what the code actually needs.

open as a page