Your client app ships under a hard download-size budget - how do you decide where per-argument body generation is worth it?
answer
- budget first, then attribute
- per hot path, not per codebase
- ask which instantiations own the bytes
- measure both strategies, real workload
- gate the size in the build
basics
~20 sPer hot path, with measurements, rather than once for the codebase. Budget the artifact, attribute its bytes to instantiations, generate bodies only where a measured hot path pays for them, and re-check both the size and the speed every release.
solid answer
~50 sTreat it as a budget with two currencies and one enforcement point. The artifact has a hard ceiling that every user pays on first load, so the first job is attribution: which generic routines and which type arguments own the bytes. Against that, find which paths are actually hot on the real workload, because only there does a direct inlined call buy anything a user can perceive. The defensible policy is mixed - generate bodies where measurement says the path is hot and the routine is small enough that duplication stays cheap, keep a shared body for the long tail, and set a size gate in the build so a regression is caught by the build rather than by a user. What you do not do is declare one strategy for the whole codebase, because the inputs to the decision differ per routine and change every release.
go deeper
Recall that code shipped over a network costs every user on first load, so how much compiled code a generic routine turns into is a product concern, not only a build detail.
Explain the two sides concretely: bytes paid by all users up front against per-call speed returned only on paths that run, and why that asymmetry favours a shared body by default.
Show the working: attribute shipped bytes to instantiations, measure the candidate path both ways on the real workload, and defend a mixed strategy rather than a blanket one.
Own the budget and its enforcement - a default, a measured exception process, a size gate in the build, and a re-attribution each release so the decision cannot quietly go stale.
## Why there is no codebase-wide answer Per-argument body generation buys a direct, inlinable call with no run-time indirection, and pays in build time, artifact bytes and instruction footprint. Whether that trade is good depends on quantities that differ per routine: how large the generated body is, how many distinct arguments it is used with, how far the expansion propagates into composed generic routines, whether those bodies coincide and merge, and whether the path is hot enough that a per-call saving is observable at all. Those quantities are not properties of the codebase, so a rule stated at codebase level is guaranteed to be wrong somewhere - either paying bytes for cold code or giving up speed on the path that matters. Under a hard download budget the asymmetry sharpens. Bytes are paid by every user, on first load, whether or not the code runs. Speed is paid back only on paths that actually execute, often enough to matter. A blanket generate-everything policy therefore charges everyone for a benefit that a minority of the code delivers. ## What to measure before deciding 1. **Attribute the artifact.** Break shipped size down to instantiations, not to source files. The question you need answered is which generic routines, at which type arguments, own the bytes - and duplication is the ceiling on that, since some bodies will have merged. 2. **Find the hot paths on the real workload**, not on a benchmark that exercises one argument. A component benchmark cannot see instruction-cache contention with the rest of the application. 3. **Measure the candidate path both ways.** The per-call saving and the footprint cost are in different units, and only a measurement on the assembled application converts between them. 4. **Watch build time as its own series.** It is the cost the team pays daily, it is not reclaimed by any merging step, and it degrades quietly. ## The shape of a defensible policy | Input | What it pushes toward | |---|---| | Path is measurably hot | generating bodies for the dominant argument | | Routine is small, few distinct arguments | generating - duplication is cheap here | | Wide routine, many arguments, composed deeply | one shared body - the expansion is multiplicative | | Code is cold or rarely reached | one shared body, whatever its call cost | | Budget already close to the ceiling | shared body by default, generation by exception | A policy other teams can follow needs three parts and not more: a default (shared body), an exception process (generate where a measurement on the real workload shows the path is hot and the win survives), and an enforcement point (a size gate in the build that fails when the artifact crosses its ceiling, so a regression is caught by the build rather than by a user on a slow connection). ## Keeping the choice reversible The worst outcome is not picking wrongly; it is picking in a way that cannot be revisited. Keep the decision reversible: - Keep the strategy choice out of the calling code, so switching a routine does not ripple into its callers. - Resist letting type arguments proliferate at the top of a composed chain, since each one multiplies bodies through every level below. - Record why each exception was granted, with the measurement attached, so a later engineer can re-test the claim instead of inheriting it as folklore. - Re-run the attribution each release. New arguments arrive with new features, and the artifact drifts without anyone deciding anything. ## What to say out loud in the interview The honest conclusion is the one to state plainly: this is decided per hot path with measurements, not declared once for a codebase, and the decision has a shelf life. Two failure modes are worth naming as the things you are protecting against. The first is optimising for a benchmark loop instead of the shipped artifact, which produces a fast component inside an application nobody can download. The second is deferring size work to a later release, which never comes, because every release adds arguments and none removes them. A lead's real deliverable here is not a preference between two compilation strategies. It is the budget, the attribution that makes the budget actionable, the gate that enforces it, and a rule narrow enough that the engineers applying it do not have to re-derive the trade-off every time they write a generic routine.
- Why is a blanket generate-everything policy particularly bad under a download budget?Because the two costs fall on different populations. Bytes are paid by every user on first load, whether the code runs or not, while the speed is returned only on paths that actually execute often. Generating everywhere charges the whole user base for a benefit a small fraction of the code delivers.
- What makes the decision go stale, and how do you catch that?New features bring new type arguments, and each one multiplies bodies through every composed level beneath it, so the artifact drifts without anyone choosing. Catch it by re-running the size attribution each release and by putting a size ceiling in the build, so a regression fails the build rather than reaching a user.
- How do you keep the choice reversible once it has been made per routine?Keep the strategy out of the calling code so switching a routine does not ripple into its callers, avoid letting type arguments proliferate at the top of a composed chain, and record the measurement behind each exception. Then a later engineer can re-test the claim instead of inheriting it as folklore.
saying these in an interview costs you the question
- Declares one strategy for the whole codebase and never revisits it
- Optimises for a benchmark loop rather than the shipped artifact
- Treats download size as something to clean up in a later release
- Assumes the cost grows linearly with the number of generic routines
- Picks a strategy from a rule of thumb instead of an attribution
- Leaves the budget unenforced, so regressions reach users first