skip to content

How do you assign per-layer bit-widths across a network under one total size budget?

level: principalimportance: nice to knowfreq 30%

answer

  1. size is parameters times bits
  2. flat layers donate, steep layers spend
  3. equalize accuracy lost per bit
  4. sweep a multiplier to hit the budget
  5. independent drops do not add

basics

~20 s

Spend bits where the sensitivity curve is steep and take them from where it is flat, equalizing the accuracy lost per bit saved until the size budget is met. Then validate the joint configuration, because per-layer curves were measured independently.

solid answer

~50 s

I treat it as an allocation problem over a measured sensitivity profile. Each layer's size cost is its parameter count times its bit-width, and its sensitivity curve gives an estimated accuracy drop at each candidate width. The allocation rule is to equalize the marginal cost: keep taking bits from whichever layer loses the least accuracy per parameter-bit saved, until the total fits. Mechanically it is a small knapsack: pick per layer the width minimizing drop plus lambda times size, and sweep lambda to hit the budget. Three constraints keep it honest. Only a few widths are worth having, so the search space is small and discrete. Boundary layers usually start exempted. And the per-layer drops were measured in isolation, so the joint configuration must be evaluated as a whole, then retrained and re-measured. The leadership question is whether the search is worth its maintenance cost against a two-tier rule.

go deeper

for a junior

Know the basic accounting: a layer's contribution to model size is its parameter count times its bit-width, so the same budget can be met by many different assignments across layers.

for a middle

Explain the greedy rule — take precision from the layer that loses the least accuracy per bit saved — and why only a handful of discrete widths are worth searching over.

for a senior

Show that you validate the joint configuration rather than trusting summed per-layer estimates, and that you iterate allocate-cut-retrain-remeasure instead of optimizing once against a stale profile.

for a principal

Own the decision of whether the search happens at all. Weigh the accuracy won against a per-layer vector that must be regenerated on every architecture change, and set who owns it and what baseline justifies keeping it.

## Framing You are given a total size budget — the compressed model must fit some fixed number of megabytes — and a network of, say, twenty layers. A uniform bit-width satisfies the budget but ignores everything you know: the sensitivity sweep already told you that some layers collapse at 4 bit while others are indistinguishable from full precision. The job is to convert that profile into a per-layer bit-width vector that meets the same budget at a smaller accuracy cost. ## The allocation rule Write each layer's size as parameters times bit-width, and let its sensitivity curve give an estimated accuracy drop for each candidate width. Two equivalent ways to state the optimum: **Greedy / equal marginal cost.** Start everything at the highest width you would consider. Repeatedly remove one step of precision from whichever layer loses the least accuracy per bit of size saved. Stop when the budget is met. At the stopping point the marginal accuracy-per-bit is roughly equal across layers — if it were not, you could move a bit from a cheap layer to an expensive one and come out ahead. **Lagrangian sweep.** For a multiplier lambda, choose for each layer independently the width minimizing `drop_l + lambda * size_l`. Small lambda gives a large accurate model, large lambda a small degraded one. Sweep lambda until the total size lands on the budget. This is the same optimum, computed in one pass per lambda instead of step by step, and it makes the tradeoff curve — accuracy against total size — fall out for free. The practical effect is exactly what the leaf name suggests: the flat layers donate the bits that the steep layers spend. A twenty-layer network might end with a handful of layers at 8 bit, the bulk at 4, and a couple of very redundant wide layers lower still, at the same total size as a uniform assignment that was measurably worse. ## The constraints that make it tractable - **The width set is small and discrete.** There is no point optimizing over a continuum: only a few widths are actually representable and worth supporting, so each layer has perhaps three or four options and the search is tiny. - **Granularity may be coarser than one layer.** If a whole block or a repeated stage must share a setting, optimize over groups, not layers. Fewer knobs also means a vector a human can read. - **Boundary layers are usually pre-assigned.** The stem and the output head typically start at a higher width, which removes them from the search and costs little budget. ## Where the arithmetic misleads you The curves were measured one layer at a time. Two consequences: **Drops do not add.** When ten layers are degraded simultaneously, the perturbations propagate and interact, and the joint drop is typically worse than the sum of the individual drops. The sum-of-independent-drops objective is a ranking heuristic, not a prediction of the final accuracy. Any chosen vector must be evaluated as a complete configuration before it is believed. **The profile shifts after retraining.** Recovery fine-tuning changes which layers are fragile, because layers adapt to the perturbation their neighbours now carry. That argues for one or two outer iterations — allocate, cut, retrain, re-measure the layers you cut hardest, reallocate — rather than a single-shot optimization polished to four decimal places on a model that no longer exists after training. ## The judgment call an interviewer is really probing A per-layer bit vector is a liability as well as an asset. It has to be re-derived when the architecture changes, when the data drifts, when someone adds a block. Against it stands the two-tier default: boundary layers at 8 bit, everything else at 4, decided in an afternoon and understandable by anyone. So the leadership answer is conditional. Do the full search when the model is a long-lived asset shipped at scale, when the uniform assignment misses the accuracy bar and the two-tier rule does not close the gap, or when the sensitivity profile is strongly uneven — a nearly flat profile means there is little to win and uniform is correct. Skip it when the architecture is still moving, when the team cannot commit to re-running the sweep on every change, or when the accuracy headroom is comfortable. If you do adopt a per-layer vector, ship it as versioned configuration alongside the profile it came from and the evaluation that validated it, name an owner for regenerating it, and record the accuracy at the uniform baseline so anyone can see what the complexity bought. An unowned, unexplained bit vector inherited across three model generations is worse than the simple rule it replaced.

  • Why does the sum of independent per-layer drops understate the joint degradation?
    Each curve was measured with every other layer intact, so it captures a perturbation the rest of the network could still absorb. Degrade many layers at once and the errors propagate forward and interact — a layer that was compensating for its neighbour is now itself degraded. Treat the sum as a ranking signal and validate the full configuration end to end.
  • When would you refuse the search and ship a two-tier rule instead?
    When the sensitivity profile is nearly flat, so there is little to win; when the architecture is still changing and the vector would be stale within a month; or when no one will own re-running the sweep. Boundary layers at a higher width and everything else uniform is understandable, robust to change, and often within noise of the optimized vector.
  • How do you keep an optimized bit vector maintainable across model versions?
    Ship it as versioned configuration next to the sensitivity profile it was derived from and the evaluation that validated it, prefer per-block groups over per-layer knobs so a human can read it, name an owner responsible for regenerating it when the architecture changes, and always record the uniform-assignment baseline so the complexity can be justified or dropped.

saying these in an interview costs you the question

  • Assigns the same bit-width everywhere despite an uneven profile
  • Adds independent per-layer drops and trusts the total
  • Optimizes over continuous bit-widths that cannot be represented
  • Skips end-to-end validation of the chosen configuration
  • Ships a per-layer vector with no owner and no baseline comparison

context