A hot loop writes millions of elements into an array whose widening forces a per-store type check - what does that check cost, and when can a compiler drop it?
answer
- one guard, always on the write side
- predictable branch, unpredictable optimisation loss
- even code that never widened pays it
- removal needs a proof, not a heuristic
- exact element type visible at the store
basics
~20 sOne extra type test per store: individually cheap and well predicted, but it sits on the write path and constrains bulk-store optimisations. A compiler may drop it only where the array's exact element type is known at the store and the stored value provably fits.
solid answer
~40 sThe check is a small test against the element type the array carries, so per store it costs a load and a comparison that branch prediction handles well - measurable on a tight write loop, invisible almost everywhere else. The real cost is structural: an operation that could otherwise be a straight memory write now has a guard, which limits how aggressively stores are batched or vectorised, and it applies even to code that never widened anything. Removing it requires proof, not heuristics: the compiler must know the array's exact element type at that store - typically because the array was allocated nearby and never escaped as a wider reference - and know the value fits. Otherwise it can at best hoist one guard out of the loop rather than delete the check.
go deeper
Recall that the safety of a widened array is bought with work on every write, and that writes therefore cost slightly more than they would otherwise.
Explain what the guard compares and why a well-predicted branch is cheap per store yet still blocks turning a fill loop into a bulk transfer.
Show the diagnosis: measure before blaming the guard, then remove the widening from the hot write path so the exact element type is visible at the store.
Weigh the two designs for your platform - a per-store run-time tax on all code versus rejecting the convenient call and obliging authors to publish read-only surfaces.
## What the check actually is When a language allows an array of a narrower element to be reached through a reference typed for a wider one, every store has to be validated while the program runs. Concretely, the store site must obtain the element type the array carries, test the incoming value against it, and proceed or refuse. That is a load plus a comparison plus a branch that is taken essentially never in correct programs. ## Why it is usually cheap - The branch is **overwhelmingly one-directional**, so a predictor learns it immediately and the guard costs close to nothing in the steady state. - The element type is **already adjacent** to data the store touches, so the load is usually warm. - Most code stores at a rate that is dwarfed by the work producing the values, so the guard disappears into the noise. On an ordinary write path, arguing about this check is premature. The honest senior answer starts here, not with outrage. ## Where it is not cheap 1. **Tight fill and copy loops**, where the store is nearly the whole body. Here the guard is a real fraction of each iteration, and the fraction is visible in a microbenchmark. 2. **Bulk stores.** A guard per element blocks turning a copy into a wide block move or a vectorised write, unless the language can establish the property once for the whole run. This is often a bigger loss than the guard itself. 3. **Everyone pays.** The check is on the store, not on the widening, so code that never widened an array anywhere still carries it. There is no way to opt out by writing careful code. ## When a compiler may remove it Removal demands a proof, and there are only a few ways to get one: 1. **The exact element type is known at the store.** If the array was allocated in view of the optimiser and the reference reaching the store provably denotes that allocation, the element type is not a guess and the test is decided at compile time. 2. **The stored value provably fits.** When the value's type is exactly the array's element type - often because it came out of the same array, or from a factory the optimiser can see through - the comparison has a constant answer. 3. **One guard covers many stores.** Where neither of the above holds but the array and the value's type are loop-invariant, a run-time compiler can check once before the loop and run an unguarded body, falling back if the assumption is violated. This is not deletion; it is amortisation, and it needs a safe exit. What does **not** license removal: the storing routine being small, the loop being hot, or the code having no widening anywhere in the source the author can see. Separate compilation and dynamic loading mean the optimiser generally cannot assume what it has not been shown. ## The alternative design, and what it actually saves | | Widening allowed, stores checked | Container usable only at its own element | |---|---|---| | The convenient call | accepted | rejected at build time | | Cost per store | a guard | none | | When the mistake is found | at run time, at a blameless store | at compile time, at the call | | What the author must write | nothing extra | a read-only view, or a routine parameterised over the element | The second design is not free either: it pushes work onto the author, who must publish a read-only surface for readers or make the routine generic over the element type. It just moves the cost to build time, where it is paid once, instead of to every store for the life of the program. ## What this means for a system you own Do not redesign a data path because of this guard until you have measured it. When it does show up, the fixes in order are: stop routing bulk writes through a reference that could have been widened, so the exact element type is visible at the store; prefer a container parameterised over the element type, where the language can discharge the obligation at compile time; and only then consider restructuring the loop. Treat it as evidence about the shape of your write path rather than as a defect to be hand-optimised away.
- Why is blocking a bulk store often worse than the guard itself?A guard per element costs a predictable branch; losing the block move costs the difference between one wide transfer and n element-sized ones. The per-element tax is linear and small, while the missed transformation changes the constant factor of the whole loop, which is what shows up in a profile.
- Does a loop-invariant guard hoisted out of the loop count as removing the check?No. It amortises the check rather than deleting it, and it obliges the compiler to keep a correct exit for the case where the assumption fails. The guarantee is intact; only its frequency changed.
- If the check is usually cheap, why is the design still criticised?Because its real price is not throughput but safety: a store that is locally well typed can fail at run time, far from the widening that caused it. The guard converts silent corruption into a loud failure; it never converts the mistake into a build error.
saying these in an interview costs you the question
- Says the check scans the whole array on every store
- Claims it happens once when the array is created
- Assumes any hot loop lets the compiler drop it
- Thinks only code that widens an array pays the cost
- Calls the check free because the branch predicts well
- Proposes hand-optimising the write path before measuring