skip to content

Compiling a separate body for each type argument makes calls faster - what does the build pay for that speed?

level: middleimportance: must knowfreq 52%

answer

  1. the speed is bought, not free
  2. count distinct arguments, not call sites
  3. build minutes, artifact bytes, cache lines
  4. growth is multiplicative, not additive
  5. buys a direct, inlinable call

basics

~20 s

Per-argument body generation costs build time (a fresh compile per distinct argument), artifact size (every generated body ships), and instruction-cache pressure (more distinct code competing for the same cache). The gain is a direct, inlinable call with no run-time indirection.

solid answer

~50 s

Generating a body per type argument moves work from run time to build time, and the build is where the price lands. Each distinct argument is a fresh compilation: the body is instantiated, optimised and emitted again, so build time follows the number of distinct type arguments, not the number of call sites. Every body that survives to link time is also bytes in the shipped artifact, and the growth is multiplicative rather than additive, because a generic routine instantiates the generic routines it calls for the same argument. The third cost is the one people forget: more distinct code doing the same work means a larger instruction footprint competing for a finite cache. What you buy is a body in which the argument is concrete, so operations on it are direct calls the optimiser can inline, with no run-time indirection left at all.

code

pseudocode · 13 lines
pseudocode
// written once
function largest<T>(items, greater)
    best = items[0]
    for each item in items
        if greater(item, best) then best = item
    return best

// what the artifact holds after per-argument generation
largest_for_A(items)   // greater is a known target here, inlined
largest_for_B(items)   // a second full copy of the loop
largest_for_C(items)   // a third

// three call sites that all pass A still share largest_for_A

go deeper

for a junior

Recall that generic source written once can be compiled many times, and that each compiled copy is bytes in the shipped build as well as work for the build machine.

for a middle

Explain that the count follows distinct type arguments rather than call sites, and name all three costs: compile time, artifact bytes and instruction-cache footprint, against a direct inlinable call.

for a senior

Show how you would find the cost in a real build - which instantiations expanded most, how much of the artifact they own, and whether the hot path measurably got faster in return.

for a principal

Frame it as a budget with two currencies, bytes and build minutes, and set where the team is allowed to spend them: on measured hot paths, re-checked each release, not by a blanket rule.

## What gets compiled, and how many times Generic code is written against a placeholder type rather than a concrete one. Under a strategy that generates a body per type argument, the toolchain does not ship that generic source as one routine: for every distinct argument the program actually uses, it produces a separate compiled body in which the placeholder has been replaced by that argument throughout. The routine you wrote once exists many times in the compiled output - once per distinct argument. The first thing to get right is what the count follows. It is **not** the number of call sites. A thousand calls that all pass the same type argument share a single generated body. Two calls passing two different arguments produce two bodies. The unit of cost is the **distinct type argument**, and the set of distinct arguments is exactly the thing that grows as a codebase and its team grow. ## The bill, in three parts - **Build time.** Each generated body is a fresh trip through the optimiser and the emitter. Several standard optimisation passes are superlinear in body size, so duplicating a large generic routine across a dozen arguments costs noticeably more than twelve times a trivial one. This is the cost the team pays, on every build, for the whole life of the code. - **Artifact bytes.** Every body that survives to link time occupies space in the shipped output. For code delivered over a network under a size budget, this is the cost the end user pays, on first load, on every device, whether or not that code path ever runs. - **Instruction-cache pressure.** A processor executes out of a finite instruction cache. Ten bodies doing the same work have roughly ten times the instruction footprint of one shared body, so a loop that touches several of them can evict its own code between iterations. The removed indirection is a saving of a few cycles per call; a cache miss is worth far more than that. ## Why the growth is multiplicative Generic code composes, and that is what turns a linear-looking cost into an expansion: 1. A generic container is instantiated for an element type. 2. The generic routine that sorts it is instantiated for that element **and** for the comparison it was handed. 3. The generic pipeline wrapping both is instantiated for its input, its output and its error type. Each level instantiates the level below it for the same arguments, so one new argument introduced at the top can produce a fresh body at every level underneath. The expansion follows the product of the arguments along the chain rather than their sum, which is why a single new type threaded through a widely used generic utility can cost more artifact bytes than a hundred new ordinary functions. Toolchains push back on this. Where two generated bodies end up with identical instructions, they can be merged so the artifact carries the code once. That is a genuine rescue for size, but it happens after the bodies have been generated and optimised, so the build time is already spent. ## What the bill buys - Inside a generated body the placeholder is a concrete type, so every operation on it has a known target rather than one selected through a run-time indirection. - A known target can be inlined, and inlining is what unlocks the rest: constant folding across the boundary, removal of branches that are dead for this argument, better register allocation. - Values of the argument can be laid out and moved as that argument, without a uniform representation imposed to keep one shared body honest. - Nothing about the type argument has to be consulted while the program runs. ## Who feels which cost, and when | Item | Who feels it | When it is felt | |---|---|---| | Extra compile work | the team | on every build | | Extra artifact bytes | the end user | on first load, before anything runs | | Larger instruction footprint | the running program | in the hot loop | | Removed indirection, inlined calls | the running program | at every call (this is the gain) | ## The honest conclusion None of the three costs makes per-argument generation wrong, and the gain does not make it free. The trade is real in both directions, and it is settled per hot path with measurements rather than declared once for a codebase: find which instantiations own the bytes, find which paths are actually hot, and spend the duplication where a measurement says it pays. A team that cannot say which generic routines dominate its artifact has not made the trade at all - it has inherited it.

  • Does the cost follow the number of call sites or the number of distinct type arguments?
    Distinct arguments. A thousand call sites that all pass the same argument share one generated body, while two call sites passing two different arguments produce two. That is why introducing one new type argument into a widely used generic utility can cost more than adding a hundred ordinary calls.
  • Why is the growth described as multiplicative rather than additive?
    Because generic code composes. A routine instantiated for an argument instantiates, for that same argument, the generic routines it calls, so one new argument at the top can produce a fresh body at every level below. The total follows the product of the arguments along the chain, not their sum.
  • If the toolchain merges identical generated bodies, is the build time saved too?
    No. Merging happens after the bodies have been generated and optimised, so the compile work is already spent. Deduplication rescues artifact bytes; it does not give back build minutes. That is why build time keeps climbing even in a codebase whose shipped size looks well behaved.

It is the trade a print shop makes with plates: a dedicated plate per poster prints faster and cleaner than one adjustable rig, but each plate takes time to cut and shelf space to keep, so you cut plates for the runs you actually print in volume.

saying these in an interview costs you the question

  • Thinks generated bodies are free because the source was written once
  • Counts call sites instead of distinct type arguments
  • Assumes the artifact grows by a fixed amount whatever the argument count
  • Believes a larger binary always runs faster because everything is inlined
  • Treats build time as a nuisance rather than a real recurring cost
  • Assumes the expansion is additive across composed generic routines