skip to content

Ten workers share one mounted tree holding millions of tiny frame files, and throughput collapses — what is the dominant cost?

level: seniorimportance: nice to knowfreq 32%

answer

  1. two budgets, not one
  2. the link is idle, the workload crawls
  3. metadata calls outnumber the bytes
  4. one hot directory serialises
  5. make the unit bigger, not the pipe

basics

~20 s

Round trips, not bytes. On a shared mount every open, stat, create and directory lookup crosses the network, and with tiny files those metadata operations vastly outnumber the data transferred. The fix is to make the unit of work bigger.

solid answer

~50 s

Throughput on any networked storage has two budgets — **bytes per second** and **operations per second** — and tiny files spend the second one. A shared mount preserves filesystem semantics by sending each metadata operation to a service across the network: open, stat, create, rename, directory lookup, close. For a one-kilobyte frame, those round trips dwarf the kilobyte itself, so you are measuring latency multiplied by file count and the link is nearly idle. Sharing makes it worse: a hot directory that ten workers all create into serialises on the server's updates to that one directory. The answer is not a faster mount, it is a bigger unit: keep per-frame work on the worker's local block volume, pack the frames into a few large files or one container, and touch the shared tree or the object store once per batch instead of once per frame.

go deeper

for a junior

Take away one habit: many tiny files over a network is slow because of the number of operations, not the amount of data. Ask how many round trips a job makes before asking how much it transfers.

for a middle

Explain which operations cost the round trips — path resolution, create, stat, open, close — and why caching does not rescue a workload that touches each file exactly once.

for a senior

Diagnose it: operations per second saturated against an idle link, contention on one hot directory, and a fix that changes the unit — local scratch, packing, one published artefact — rather than buying bandwidth.

for a principal

The angle is setting the storage contract other teams build against. If the platform makes per-item storage easy and packing hard, every team writes the pathological shape; the standard, and the tooling for it, are the real deliverable.

## Throughput has two budgets Every storage system is bounded by two separate things, and engineers habitually reason about only the first: - **Bytes per second** — how much data the link and the device can move. - **Operations per second** — how many discrete requests can complete, which is set by latency and by how many you can have in flight. When the average file is large, the byte budget dominates and everything behaves the way intuition expects. When the average file is tiny, the operation budget dominates completely: the link sits nearly idle while the workload crawls. A fleet writing millions of individual frames is the clearest case of this there is. ## Where the round trips come from A shared mount is not a disk. It keeps filesystem semantics by asking a service across the network, and the operations that cost you are mostly not reads and writes of content: 1. **Path resolution.** Reaching `frames/job-77/scene-4/000512.png` may consult each component, and a deep tree costs more than a shallow one. 2. **Create.** Making a new file means allocating it and updating the containing directory. 3. **Stat.** Every size check, existence check and listing entry is a request. A tool that walks a tree to see what is there issues one per entry. 4. **Open and close.** Each carries state the service must track. 5. **The actual bytes.** For a one-kilobyte frame, this is the cheapest part of the whole sequence. Caching helps a single reader that revisits the same paths and helps almost nothing here, because these files are each touched once. That is the property that makes this workload pathological: no locality to amortise anything against. ## Sharing makes it worse, not better The second effect is contention on shared structure. Ten workers creating into one directory are all updating the same object on the service, which has to serialise those updates to keep the directory consistent. Adding workers past that point adds queueing, not throughput — the classic shape where the eleventh worker makes everyone slower. Splitting the work across many directories relieves it; pointing all ten at one hot directory does not. | the habit | what it costs here | the fix | |---|---|---| | one file per frame | a create, an open and a close per frame | pack frames into one file per shot or batch | | a deep directory tree | path resolution on every access | flatten, or address by a single name | | walking the tree to find work | a stat per entry, repeatedly | keep the work list outside the filesystem | | all workers in one directory | serialised directory updates | shard by worker or by job | | scratch on the shared mount | a round trip per operation on hot data | scratch on the worker's attached volume | ## Making the unit bigger The durable fix is always the same shape: **do less per unit of value**. 1. **Keep per-frame work local.** Frames are written and read by one worker, so they belong on that worker's attached block volume, where an operation is local-device latency rather than a network round trip. 2. **Pack before you publish.** Roll a shot's frames into a single container or archive and write that once. Millions of operations become thousands. 3. **Publish the packed result to an object store.** One key per finished artefact, fetched by whoever needs it, with no mount and no shared tree to contend on. 4. **Keep the index out of the filesystem.** If something must know which frames exist, that is a list in a database or a manifest object, not a directory you walk. ## The same trap on the other shapes This is not a shared-mount-specific flaw; it is the general shape of networked storage, and it shows up everywhere: - **In an object store**, every object is at least one request, so millions of tiny objects means millions of requests. A store is superb at few-and-large and mediocre at many-and-tiny, and listing a store holding an enormous number of keys is itself a long paged scan. - **On a block volume**, the device has its own operations-per-second ceiling, often provisioned separately from capacity and from throughput. A workload of tiny scattered writes can saturate it while moving very little data. The diagnostic instinct to build is: when something is slow, ask **how many operations per unit of useful work** before asking how many bytes. If the answer is "one or more per tiny file", no amount of extra bandwidth, extra workers or a faster storage tier will rescue it — only changing the unit will. And the measurement that proves it is simple: watch operations per second and bytes per second side by side. A saturated operation count against an idle link is the whole diagnosis.

  • The team proposes a faster mount tier to fix it. Why is that usually the wrong purchase?
    A faster tier raises the byte budget, and bytes were never the constraint — the link is already idle. It may raise the operation ceiling somewhat, but the workload issues one or more round trips per tiny file, so the cost per unit of value is unchanged. Packing frames into larger units changes the arithmetic instead.
  • Why does adding more workers past a point make the job slower rather than faster?
    Because they contend on shared structure. Creating into one directory means updating one object on the service, and those updates serialise to keep it consistent, so extra workers queue behind each other. Sharding the output across many directories, or across separate trees per worker, restores the parallelism.
  • Does moving the same millions of tiny files to an object store fix this?
    Not by itself. Each object is at least one request, so millions of tiny objects is still millions of requests, and listing them is a long paged scan. What the object store does remove is the shared-tree contention and the provisioned size. The real fix in both cases is packing.

It is the difference between shipping a thousand envelopes and one pallet. The postage per envelope, not the weight of the paper, is what ruins you.

saying these in an interview costs you the question

  • Blames link bandwidth when the link is nearly idle
  • Buys a faster storage tier for a metadata-bound workload
  • Adds workers to a contended directory and expects scaling
  • Assumes a mount behaves like a local disk per operation
  • Believes moving the tiny files to an object store fixes it alone
  • Walks the whole tree to discover work on every run