skip to content

Why is a container's first write to a large inherited file costlier than the writes that follow it?

level: middleimportance: should knowfreq 52%

answer

  1. lower layers never change
  2. modify means copy upward first
  3. cost scales with file size
  4. one byte can copy everything
  5. after copy-up, writes are ordinary

basics

~20 s

The image's layers are read-only, so the union filesystem cannot modify the file in place. On the first modification it copies the file up into the container's writable top layer and applies the write there; every later write hits that copy directly.

solid answer

~40 s

Everything a container inherits from its image sits in read-only layers that nothing may edit. When the process first modifies such a file, the union filesystem performs a **copy-up**: it copies the file into the writable layer on top of the stack, then applies the write to that copy. Because the copy is of the whole file, the cost is proportional to the **file's size, not to the number of bytes written** — changing one byte of a 400 MB file can move 400 MB. It also consumes that much space in the writable layer. After the copy-up the path resolves to the upper copy, so subsequent writes are ordinary in-place writes with no extra cost. Reads never trigger it: they are served straight from the lower layer.

code

pseudocode · 12 lines
pseudocode
write(path, offset, bytes):
    if writableLayer.has(path):
        writableLayer.apply(path, offset, bytes)   # ordinary write, no copy
        return

    source = topmost readOnlyLayer holding path
    if source is none:
        writableLayer.create(path, bytes)          # brand-new file, no copy
        return

    writableLayer.copyEntireFile(from = source, path = path)   # cost = file size
    writableLayer.apply(path, offset, bytes)

go deeper

for a junior

Know the shape: the image's layers cannot be changed, so writing to a file that came from the image makes a private copy first, and that copy is what you are editing.

for a middle

Explain copy-up precisely — which layer the copy lands in, that it is lazy, that it copies the whole file, and that the cost is therefore the file's size rather than the write's.

for a senior

Bring the operational consequence: a first-write stall that does not appear at start-up, writable-layer space consumed by whole files, and a design that moves hot mutable data out of the layer stack.

for a principal

Own the trade-off: read-only shared layers buy cheap start-up and cross-image reuse, and copy-on-write is the bill. Decide where mutable state is allowed to live so that bill is never paid on a hot path.

## Why nothing below can be edited A container's filesystem is the image's ordered stack of **read-only layers** with one **writable layer** on top. The lower layers are read-only by design, not by permission setting: one stored copy of a layer may be serving many containers and many images at once on the same host, so letting any one container modify it would corrupt every other user of it. The format's whole economy — start a container without copying anything, share layers between images — rests on those layers never changing. That leaves a question the union filesystem has to answer: what happens when a process opens an inherited file for writing? ## Copy-up on the first modification The answer is **copy-on-write**. The first time a container modifies a file that came from a lower layer, the union filesystem: 1. Locates the file in the topmost read-only layer that holds it. 2. Copies it, in full, into the container's writable layer. 3. Applies the modification to that new copy. After that, top-down resolution finds the file in the writable layer first, so the process is working on its own copy and the lower one is simply shadowed. The behaviour is **lazy**: nothing is copied when the container starts, and a file that is only ever read is never copied at all — reads are served in place from the lower layer, which is why hundreds of containers can share one copy of a large read-only asset with no duplication. ## The cost is the file's size, not the write's This is the part that surprises people, and it is the reason the question gets asked: - Appending one kilobyte to a 400 MB inherited file typically copies **400 MB** before the kilobyte lands. - The writable layer then grows by the size of the file, not by the size of the change. - The latency lands on the **first** write, so a service can start cleanly, serve traffic, and then stall for seconds the first time it touches a large inherited file. - The second and every later write to that file are ordinary writes with no copy, which is exactly why the first one looks anomalous in a latency graph. - Changing a file's metadata — its owner or permission bits — counts as a modification on typical implementations, so a start-up step that adjusts ownership on a large inherited tree can trigger copy-ups of files whose contents nobody touched. | Operation on an inherited file | Copy-up? | Cost driver | |---|---|---| | Read the contents | No | Served in place from the lower layer | | First write, of any size | Yes | The whole file's size | | Later writes to the same file | No | The bytes written | | Metadata change on typical implementations | Yes | The whole file's size | | Create a brand-new file | No | It is written straight into the writable layer | ## Where implementations differ Union filesystems do not all pay the same price. Some copy whole files and nothing else; some can share unchanged regions with the layer below, so a large file's first write costs far less than its size; some can perform the copy cheaply when the writable layer and the lower layer live on the same backing store. Designs genuinely differ here, so the honest statement is: **assume whole-file copy-up unless you know the platform does better**, and measure rather than assume. What does not differ is the direction of the rule — lower layers are never modified, so a modification always produces a copy somewhere above them. ## Designing around it The fix is not to make copy-up faster; it is to keep frequently written data out of the layer stack entirely: - Have large, mutable files supplied from outside the image at run time, so writes go straight to that storage and no inherited copy exists to duplicate. - Where a file must be rewritten wholesale at start-up, have the process write a fresh file rather than modify an inherited one — creating a new path costs only what it writes. - Treat a big, hot, inherited file as a smell in image design: it is simultaneously paid for in transfer, in on-disk size, and again in copy-up on the first write. ## Answering it in an interview Lead with the constraint, then the mechanism, then the number: lower layers are read-only because they are shared, so the first modification copies the file up into the container's writable layer and writes there; the cost therefore scales with the size of the file rather than the size of the write, which is why one byte into a huge inherited file can move hundreds of megabytes while every write after it is ordinary.

  • Does reading an inherited file trigger the same copy?
    No. Reads are served in place from the layer that holds the file, which is what lets many containers share one stored copy of a large read-only asset. Only a modification forces a copy-up — and on typical implementations a metadata change, such as adjusting ownership, counts as a modification even though no content changes.
  • How would you avoid the penalty for a file the workload rewrites constantly?
    Keep it out of the layer stack. Have the path supplied from outside the image at run time so writes land directly on that storage and there is no inherited copy to duplicate. Alternatively, have the process create a fresh file instead of modifying an inherited one, since a new path is written straight into the writable layer with no copy-up.
  • Why does a service sometimes stall seconds after it starts serving, rather than at start-up?
    Copy-up is lazy. Nothing is copied when the container starts, so the cost lands the first time the process actually modifies an inherited file — which may be minutes in, on the first request that touches it. The second such write is fast, so the symptom looks like a one-off anomaly rather than a systematic cost.

saying these in an interview costs you the question

  • Thinks the runtime edits the file in the lower layer
  • Expects the cost to match the number of bytes written
  • Believes the copy happens when the container starts
  • Says reading an inherited file also copies it upward
  • Claims a second write to the same file copies it again