Why does a hot loop that built one small coordinate value per iteration without allocating start allocating heavily once each value is passed to a caller-supplied callback?
answer
- the computation did not change
- a proof was withdrawn, not a bug added
- unknown callee must be assumed to retain
- one block per iteration scales with throughput
- pass the components, not the composite
basics
~20 sThe value now escapes. A caller-supplied callback is a target the compiler cannot examine, so it must assume the reference is retained. The proof that kept the value in the frame is withdrawn and every iteration allocates a real block.
solid answer
~50 sIn the first version every reference to the per-iteration value died inside the iteration, so the compiler could prove non-escape and keep the value in the frame — or dissolve it into registers entirely, in which case nothing was allocated at all. Passing it to a callback supplied by the caller removes the evidence: the target is chosen outside the compiled code, so the analysis must assume the callee stores the reference somewhere long-lived. Conservative assumption means escape, escape means a dynamic block, and the loop now produces one block per iteration, so the volume scales with throughput. The source edit looks trivial because the change is not in what the code computes but in what can still be *proven* about it. Handing the callback the two plain numbers instead of the composite value usually restores the original placement.
code
pseudocode · 12 lines// version 1: every reference to p dies inside the iteration
sum = 0
for each row in rows:
p = make Point(row.x, row.y)
sum = sum + p.x * p.y
// version 2: p is handed to a target chosen by the caller
sum = 0
for each row in rows:
p = make Point(row.x, row.y)
observe(p) // unexamined callee: must assume p is retained
sum = sum + p.x * p.ygo deeper
Remember that handing a freshly built value to code the compiler cannot see forces that value into dynamic storage, even when the calculation is unchanged.
Explain the mechanism: an unexamined target could retain the reference, the analysis must assume the worst legal case, and the frame-local proof is therefore withdrawn.
Recognise the signature in production - a per-operation allocation figure that jumps after a harmless-looking hook or extension point - and fix it by reshaping the interface rather than by shrinking the value.
Weigh the extension point against its cost: an open callback on a hot path buys pluggability and spends a provable placement, and that trade should be made deliberately rather than discovered in a profile.
## The two versions Version one builds a small coordinate value each iteration, reads its two fields and drops it. Version two does the same, then hands the value to a callback that the caller supplied. Nothing else differs — same arithmetic, same iteration count, same result. Yet the first version allocates nothing per iteration and the second allocates once per iteration. ## Why the verdict flips Placement follows a proof, and the proof is about **reachability after the frame returns**, not about what the code appears to do. - In version one the compiler can see every use of the value. All of them are reads, all of them are inside the iteration, and none stores, returns or publishes a reference. Non-escape is proven, and the value gets a frame slot or is scalar-replaced into registers. - In version two the value is handed to a target chosen by whoever called this code. The compiler has no body to examine, so it cannot rule out that the callee writes the reference into a long-lived field, appends it to a shared structure or passes it to another thread. The analysis is a *may* analysis: what it cannot rule out, it must assume. The key asymmetry: a wrong 'escapes' verdict costs one allocation, while a wrong 'does not escape' verdict leaves a live reference pointing into storage the frame pop has already released. So the analysis always errs toward escape, and **absence of evidence is treated as evidence of escape.** ## What that costs under load 1. **One dynamic block per iteration.** The per-operation cost is small, but it is multiplied by the loop's trip count and again by the request rate, so it scales exactly with throughput. 2. **Downstream reclamation work.** Every block that is created is a block some reclamation policy must later account for, whichever policy the platform uses. 3. **Worse locality.** Registers and frame slots are as close to the processor as storage gets; a fresh block per iteration spreads the same data across addresses that the loop then has to chase. 4. **A second allocation, sometimes.** If the callback is itself a closure that captures state, that capture is its own escaping value, separate from the coordinate. ## Getting the placement back The options, roughly in order of how much they change the interface: - **Hand over the components.** Pass the two numbers rather than the composite value. Nothing is constructed, so nothing can escape. - **Narrow the callback's contract.** If the interface promises that the callee will not retain what it is given, the call can be shaped to pass a copy or plain values, and the composite need never be built. - **Move the work inside.** If the callback exists only to accumulate, pass the accumulator's update as a value rather than the object it applies to. - **Accept the allocation and bound it.** Sometimes the interface is worth the block; then the honest step is to know the per-operation figure rather than to hope. What does *not* work is asserting the intention in a comment, or making the value smaller. Size is not the test. ## How this shows up in a real investigation The signature is a per-operation allocation figure that jumps after a change whose diff looks harmless: a new hook, a listener, a metrics callback, an extension point added so another team could plug in. Nothing about the computation changed, and no bug was introduced — a proof was withdrawn. That is why this class of regression is hard to spot in review: the cause is not in the added line, it is in what the added line made unprovable about the lines around it. A related trap is the reverse conclusion. Seeing a dynamic block does not prove the value truly escapes; it proves only that non-escape could not be established here. The callee may never retain anything. The fix is therefore about giving the compiler evidence — or removing the need for it by not creating the value at all — rather than about persuading anyone that the callee is well behaved.
- The callback demonstrably never keeps the value. Why does that not help?Because the compiler is not reasoning about this particular callee, it is reasoning about every callee the caller could supply. The target is selected outside the compiled code, so no body is available to examine. Evidence a human has is not evidence the analysis has, and the analysis must cover the worst legal case.
- Would making the coordinate value smaller fix the allocation?No. Size affects how wide the block is, not whether one is needed. A two-field value that escapes still needs storage that outlives the frame, while a much larger value that provably does not escape needs none. The test is reachability after the frame returns, never byte count.
- If the callback is a closure that captures state, what is allocated?Potentially two things: the captured state, because the closure is handed across a boundary and everything it can reach must outlive the creating frame, and the per-iteration value the loop passes to it. They are separate escapes with separate fixes - shorten the capture for the first, stop constructing the composite for the second.
saying these in an interview costs you the question
- Blames the extra call itself rather than what it makes unprovable
- Says a value passed as an argument is always copied, so it cannot allocate
- Assumes the block appears because the loop runs too many times
- Insists no allocation can happen because the callee never keeps anything
- Proposes shrinking the value instead of not constructing it