Ten type arguments each generate a body, yet the artifact grows far less - which toolchain step explains that?
answer
- bodies are instructions, not types
- same representation, no per-argument operation
- merged late, near link time
- rescues bytes, never build minutes
- generated count is an upper bound
basics
~20 sDeduplication of identical generated bodies. Where two arguments compile to the same instructions, the toolchain can emit the code once and point both instantiations at it, so the artifact carries one copy instead of ten.
solid answer
~40 sA body is generated per type argument, but two arguments frequently produce instructions that are byte-for-byte the same - typically when the arguments share a machine representation and the body only moves values around rather than performing operations specific to each. A merging step, usually late in the build or at link time, recognises those duplicates and keeps one copy that every matching instantiation refers to. That is why a routine instantiated ten ways can cost far less than ten bodies in the shipped artifact. It is only a partial rescue, though: the merge happens after the bodies were generated and optimised, so none of the build time comes back, and arguments with different layouts or genuinely different operations produce different instructions and cannot be merged at all.
code
pseudocode · 11 lines// generated for two arguments that share a machine representation
swap_for_RefA(a, b) { t = a; a = b; b = t } // moves one machine word
swap_for_RefB(a, b) { t = a; a = b; b = t } // identical -> one copy kept
// generated for an argument laid out as four words
swap_for_Wide(a, b) { t = a; a = b; b = t } // moves four words: different
// instructions, cannot merge
// a body that does per-argument work rarely coincides at all
max_for_RefA(x, y) { if compareA(x, y) > 0 ...} // inlines compareA
max_for_RefB(x, y) { if compareB(x, y) > 0 ...} // inlines compareBgo deeper
Recall that a compiled body is instructions, and that two different type arguments can end up with instructions that are exactly the same, in which case only one copy needs shipping.
Explain the conditions under which bodies coincide - same machine representation, no operation that differs per argument - and that merging happens late, on finished output.
Show that you attribute shipped bytes to instantiations rather than predicting them, and that you know the merge leaves build time untouched and spares none of the per-argument work.
Treat merging as a variable you observe, not a guarantee you plan on: it lowers the ceiling for routines that move values and does almost nothing for the ones your hot paths depend on.
## Why two arguments can produce the same instructions A generated body is instructions, not types. Once the placeholder has been replaced and the body optimised, what remains is a sequence of machine operations, and two different type arguments can arrive at exactly the same sequence. That happens most often when both conditions hold: - the two arguments have the same machine-level representation, so values of either are moved, copied and stored with the same instructions; and - the body does not perform any operation whose implementation differs between them, so nothing argument-specific is baked into the code. A routine that stores, reorders or forwards values of its type parameter without interpreting them is the clearest case: the instructions depend on the size and shape of what is moved, not on what it means. A routine that compares, formats or arithmetically combines values of the parameter usually is not, because the operation it inlines differs per argument. ## What the merging step does The toolchain compares generated bodies and, where their instructions coincide, keeps one and redirects the other instantiations to it. Nothing about the program's meaning changes: the type checking that distinguished the arguments has already happened, at a much earlier phase, and cannot be undone by observing that two compiled bodies came out the same. What is removed is duplicated bytes in the output. This is why the relationship between distinct type arguments and artifact growth is not the clean multiplication a first analysis suggests. The generated count is an upper bound. The shipped count is that bound minus whatever merged. ## What it does not rescue | Cost of per-argument generation | Does merging help? | |---|---| | Artifact bytes | Yes, this is exactly what it reclaims | | Instruction-cache footprint of the merged code | Yes, one resident copy instead of several | | Build time | No - the bodies were generated and optimised before anything merged | | Bodies that differ in layout or in operations | No - different instructions cannot be merged | The build-time row is the one that surprises people. Merging is a late step operating on finished output, so every merged body was still instantiated and optimised at full cost. A codebase can therefore have a well-behaved artifact size and a build that keeps getting slower, and those two facts are not in conflict. ## Why it is a partial rescue The fraction of bodies that merge is a property of the code, not a dial anyone sets: 1. The more the generic routine actually **does** with values of its type parameter, the fewer of its bodies coincide, because the inlined operations differ per argument. 2. The more the arguments differ in representation - especially arguments laid out directly against arguments handled through a uniform reference-shaped representation - the fewer coincide. 3. Merging can only reclaim what duplication created; it never makes a heavily specialised routine cheaper than the shared-body alternative would have been. So the honest statement is that duplication is the ceiling on the cost and merging pulls the real figure somewhere below it, by an amount you cannot predict from the number of arguments alone and can only observe by looking at the built artifact. ## Why this matters when you are budgeting If you are reasoning about whether a generic routine is affordable under a size budget, two wrong models are available and both are common. The first counts distinct arguments, multiplies, and concludes the routine is unaffordable when in practice most of its bodies merged. The second hears that merging exists, assumes it handles the problem, and stops attributing bytes at all - which fails exactly on the routines that matter, the ones doing real per-argument work, because those are the ones whose bodies do not coincide. The usable habit is to measure the artifact rather than model it: attribute shipped bytes back to instantiations, and let that tell you which generic routines genuinely dominate. Merging then becomes a pleasant part of the result rather than an assumption underneath it. And because the build-time cost is untouched by any of this, a routine that merges beautifully can still be the reason the build is slow.
- Which kinds of generic routine merge well, and which barely merge at all?Routines that move values around without interpreting them merge well, because their instructions depend on the size and shape being moved rather than on the argument's meaning. Routines that compare, format or combine values of the parameter inline a different operation per argument, so their bodies differ and stay separate.
- Does deduplication reduce instruction-cache pressure as well as artifact size?For the bodies that actually merged, yes - one resident copy replaces several, so the hot path touches less distinct code. It does nothing for bodies that could not merge, which are precisely the ones doing per-argument work, and those are usually the ones in the hot loop.
saying these in an interview costs you the question
- Assumes every generated body merges, so size never needs attributing
- Believes merging gives back build time as well as bytes
- Thinks merging two bodies weakens the type guarantees
- Confuses deduplicating bodies with switching to one shared body
- Predicts artifact growth from argument count alone