skip to content

Why does a build context take two minutes to package and transfer before ten seconds of actual build work?

level: middleimportance: should knowfreq 47%

answer

  1. cost is charged before step one
  2. the tree grew, not the build
  3. dependencies, output, fixtures, history
  4. split the wall clock at instruction one
  5. measure what was packaged, not the rules

basics

~20 s

Because the build was pointed at a working tree full of things no instruction needs — installed dependencies, earlier build output, test fixtures, version-control metadata — and the packaging step takes all of it. The lever is the set that is sent, not the machine.

solid answer

~50 s

The packaging and transfer step is charged on the **whole** context: everything under the path the build was pointed at, minus the ignore rules. In a working tree that has been used, that is dominated by files no instruction will ever read — the directory an install step fills with third-party dependencies, output from earlier local runs, fixture and sample data, exports, and version-control metadata that carries every historical version of every file. None of it is needed and all of it is walked, packaged and moved before step one. The fix is to shrink the set, in this order: add ignore rules for the offenders, point the build at a narrower subdirectory when the repository holds several components, and move genuinely large data out of the tree. Confirm by asking the builder what it actually packaged, not by reading the rules file.

go deeper

for a junior

Know that the time before the first instruction's output is packaging and transferring the directory you pointed at — not building. Look at what is actually in that directory before blaming the build.

for a middle

Explain both halves of the cost, walking files and moving bytes, and name the usual occupants of a working tree. Then name the levers in order: ignore rules first, a narrower root second, moving data out third.

for a senior

Show the diagnosis, not the fix list: split the wall clock at the first instruction, compare tree size against the size the builder reports for what it packaged, and resist assuming the rules work because they exist.

for a principal

Frame it as a cost paid by every engineer and every automated run, and as the same set that bounds exposure — which is why where teams point their builds is a standard worth setting once rather than a per-project preference.

## What the two minutes actually is It is not compilation, it is not dependency resolution and it is not anything your instructions asked for. Before the first instruction runs, the tool walks the path the build was pointed at, applies the ignore rules, packages what remains and moves it to whatever performs the build. Two costs live in that sentence: - **The walk.** Every file under the path is stat-ed and matched against the rules. A tree with hundreds of thousands of small files is slow to walk even if almost nothing survives the filter. - **The transfer.** Whatever survives is moved across a boundary — to another process, another machine, or a shared build service. As a rough anchor: a 2 GB context over a link that sustains 20 MB/s is about 100 seconds before anything begins. State your own numbers when you reason about this; the arithmetic is the argument. How much crosses on a *second* build depends on the builder — some move the full bundle every time, some synchronise only what changed since last time — but the walk and the filtering are redone every run, and the scope is decided from scratch every run. ## What is actually in there A working tree that has been used for real work usually carries, in rough order of size: - **the directory an install step fills with third-party dependencies** — frequently the single largest thing on disk, and always rebuilt inside the image anyway; - **output from earlier local runs** — compiled artifacts, packaged bundles, coverage reports; - **test fixtures and sample data** — the multi-hundred-megabyte file someone committed once to reproduce a bug; - **prior exports and generated reports** that accumulate because nothing deletes them; - **version-control metadata**, which holds the compressed history of every file the project ever had, so it can dwarf the checkout itself in an old repository; - **editor, tooling and local cache directories** nobody thinks of as project files. None of these is exotic. The reason a build "suddenly" got slow is almost never that the build changed; it is that the tree grew. ## The levers, in the order to try them | lever | effect on the set | cost of getting it wrong | |---|---|---| | ignore rules | subtracts matched paths from every build using that context root | a rule that never matched silently does nothing, so verify rather than assume | | a narrower context root | everything outside the new root becomes unreachable | an instruction that needed a file outside it now fails loudly, which is at least an honest failure | | move the data out of the tree | the file cannot be in any context | whoever needed it locally has to fetch it another way | The first is the reviewable one and should be the default: it is a file next to the code, it is read in review, and it binds every caller that uses that root. The second is worth doing when one repository holds several independently built components, because then each build's honest scope really is a subdirectory. ## How to confirm it rather than guess The diagnostic mistake is to read the ignore rules and conclude the context must be small. Rules go stale, patterns mis-anchor and new directories appear. Measure the two halves separately: 1. **Measure the tree.** Total size and file count under the context root, then per-subdirectory, so the offender names itself. 2. **Measure what was sent.** Most builders will report the size of the context they packaged; compare that number against the tree. A tree of 2 GB and a context of 1.9 GB means the rules are doing nothing. 3. **Watch where the wall-clock time goes.** Time spent before the first instruction's output appears is packaging and transfer; time after it is the build itself. They have completely different fixes and conflating them wastes an afternoon. ## Why this is worth fixing beyond the minutes The same oversized set is also the exposure surface: an instruction that copies the whole context can only take files that were sent, so trimming the set both shortens the build and shrinks what any careless copy could sweep in. And the cost is paid by everyone — every engineer, every automated run, every retry — which is what turns a "minor annoyance" into a real number when you multiply it by a team. One honest caveat: not every slow build is a context problem. If the time is being spent *after* instructions start running, the scope of the context is not your culprit and shrinking it will change nothing. That is exactly why step three above — splitting the wall clock at the first instruction — comes before any fix.

  • The context is small but packaging is still slow — what else explains it?
    The walk itself. Scope cost has two components: bytes transferred and files examined. A tree with an enormous number of tiny files is slow to walk and filter even when almost nothing survives the rules, so the bundle can be small while the step before it is not.
  • Does shrinking the build context make the finished image smaller?
    Only indirectly. Image size is decided by what instructions write into it, not by what was available to them. A smaller context makes a broad copy instruction take less, so in practice the two often move together — but a tiny context with a careless copy still produces a fat image.
  • Why do two engineers see very different packaging times on the same project?
    Because the context is their working tree, not the repository. One has run the build locally a dozen times and has an install directory, stale output and generated reports on disk; the other has a fresh checkout. Same rules, same instructions, very different sets.

saying these in an interview costs you the question

  • The slow part must be the build steps, so give the builder more CPU.
  • A second build is free because the context was already sent once.
  • Only large files matter; file count has no effect on packaging.
  • The ignore rules are written, so the context must already be small.
  • Shrinking the context is pointless because the image stays the same size.