skip to content

questions

5

Before a compiler may place a new object on the stack instead of the heap, what must it prove?

level: middleimportance: must knowfreq 68%

answer

  1. lifetime, not size
  2. who can still see it afterwards
  3. reachable after the frame returns
  4. unproven is treated as escaping
  5. frame slot, or no object at all

basics

~20 s

That no reference to the object can be followed once the creating frame returns: it is never stored in longer-lived memory, returned, or published to another thread. Without that proof the object must go on the heap.

solid answer

~50 s

The compiler runs an **escape analysis** on the creation site. It follows every reference to the new value through assignments, field stores, array stores, arguments and returns, and asks one question: can any of those references still be reached after this frame is gone? If a reference is stored into a longer-lived structure, returned to the caller, captured by something that outlives the scope, or handed to a call whose body the compiler cannot examine, the value escapes and must be heap-allocated. If nothing escapes, the value's lifetime is exactly the frame's, so it can live in the frame — or be dissolved into individual registers and stack slots, which is scalar replacement and materialises no object at all. The analysis is conservative: anything it cannot prove counts as escaping, so a heap allocation is not evidence that the value really does escape.

code

pseudocode · 14 lines
pseudocode
procedure distance(rows):
    total = 0
    for each row in rows:
        p = make Point(row.x, row.y)     // created here
        total = total + p.x * p.x + p.y * p.y
    return total
    // every reference to p dies at the end of the iteration:
    // no store, no return, no publication -> does not escape

procedure remember(row):
    p = make Point(row.x, row.y)
    registry.last = p                     // a reference is stored
    return total_of(p)
    // registry outlives this frame, so p must outlive it too -> heap

go deeper

for a junior

Recall that where a value lives is decided by how long a reference to it can survive, not by how many bytes it occupies or which syntax created it.

for a middle

Explain the test itself: follow every reference from the creation site and ask whether one can be reached after the frame returns, then name the two rewards - a frame slot, or no object at all.

for a senior

Show that the analysis is conservative and has no source-level marker, so an ordinary-looking edit can move a value to the heap and change a service's per-operation allocation without any visible change in behaviour.

for a principal

Treat the proof as a dependency rather than a guarantee, and decide deliberately which paths in a codebase may rely on it and which must be shaped so the temporary is never created.

## The question the compiler is actually asking When code creates a value, two regions can hold it. A **stack frame** is built when a call begins and destroyed when it returns, so anything living in it has exactly the lifetime of that call and costs nothing to release — the release is the frame pop. The **heap** holds blocks whose lifetime is not tied to any frame, and something has to decide when each block dies: a tracing collector, a reference count, an explicit release, or a scope-exit rule. Every one of those mechanisms costs work that a frame slot does not. So the placement decision is not about how big the value is, and not about which syntax created it. It is about **lifetime**, and lifetime is settled by **reachability**: if any reference to the value can still be followed after the creating frame is gone, the value must outlive the frame, and only the heap can hold something that outlives the frame that made it. ## The escape test An **escape analysis** answers exactly that question. Starting at the creation site, it propagates the new reference through the code and classifies the outcome: 1. **Does not escape** — every reference to the value dies with the frame. The value may be laid out inside the frame, or broken apart into registers and stack slots (**scalar replacement**), so that no object is created anywhere. 2. **Escapes the frame** — some reference is stored into a longer-lived structure, returned to the caller, or captured by something that survives the scope. The value must be heap-allocated. 3. **Escapes the thread** — a reference becomes reachable from another thread. That is the strongest verdict of the same test, and it additionally withdraws any assumption that only one thread can observe the value. The analysis is a *may* analysis, not a *must* analysis: it reports whether a reference **can** get out, over all paths, not whether it does on the path that happens to run. ## What a proof of non-escape is worth | placement | cost to allocate | cost to release | what it gives up | |---|---|---|---| | heap block | find or carve space, write a header, publish the reference | work for whatever reclamation policy owns the region | nothing, but every block adds to that policy's workload | | frame slot | already paid inside the frame's fixed size | the frame pop, and nothing else | the value cannot outlive the call | | registers and stack slots after scalar replacement | none: no object is created | none | the value has no address and no identity | The third row is the interesting one. Once the fields live in registers, later passes can treat them as ordinary local numbers: fold constants through them, keep them in a loop, drop fields nothing reads. The value has stopped being an object and become a handful of variables. ## Why unproven has to mean escaping The analysis is **conservative by necessity**. If it guessed wrong in the direction of 'does not escape', a surviving reference would point at storage that the frame pop has already released, and the program would read memory that has been reused. A wrong guess in the other direction only costs a heap allocation. So the rule is asymmetric on purpose: **not proven non-escaping is treated as escaping.** This is why identical-looking code allocates differently in different places — one call site gives the analysis enough information, the other does not. ## Where the proof typically breaks - A reference is written into a field of something that already lives on the heap. - The value, or something reachable from it, is returned. - The value is handed to a call whose body the compiler cannot examine, so the callee must be assumed to retain it. - The value is captured by a closure or a deferred task that outlives the scope. - The value is published where another thread can find it. - The value's size is not known when the frame layout is decided, so no fixed slot can be reserved for it. ## What does not change A non-escape proof is invisible to a correct program. Field values, results and ordering are all unchanged; what changes is where the bytes live, who is obliged to release them, and how much work that obligation creates. Ecosystems differ in how far they take the reward — some place a whole object in the frame, others only ever scalar-replace and never lay an object out in a frame at all — but the proof obligation is the same everywhere: show that nothing which survives the frame can still reach the value.

  • If the analysis cannot examine the body of a call the value is passed to, what must it assume?
    That the callee may keep the reference, so the value escapes and goes on the heap. The assumption is forced: an unexamined callee could store the reference in a long-lived field or hand it to another thread, and a wrong verdict of 'does not escape' would leave a live reference pointing into a frame that has already been popped.
  • Does proving non-escape change what the program computes?
    No. It changes where the bytes live and who is obliged to release them, not the values themselves. Results, field contents and ordering are identical. What does change is memory-system behaviour: how much is allocated per operation, how much reclamation work follows, and how well the data sits in cache.
  • Why can the same construction escape at one call site and not at another?
    Because escape is a property of the surrounding code, not of the type. At one site every reference stays inside the frame; at another the same construction is stored, returned or handed to a call the compiler cannot see through. The analysis re-answers the question per site, with whatever information that site gives it.

saying these in an interview costs you the question

  • Says size decides the region: small values on the stack, large ones on the heap
  • Thinks the analysis inspects what the value points at rather than what points at it
  • Assumes a value the compiler could not prove non-escaping is known to escape
  • Believes placement follows the syntax used to create the value
  • Claims the decision is about how long the code runs rather than reachability
open as a page

What forces a short-lived local value onto the heap when its scope is a single function call?

level: middleimportance: must knowfreq 58%

basics

~20 s

Three independent causes. A reference escapes by being stored, returned or published; the size is not known when the frame layout is decided; or a closure captures the value and outlives the scope. Any one of them is enough.

open as a page

Why does a hot loop that built one small coordinate value per iteration without allocating start allocating heavily once each value is passed to a caller-supplied callback?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The value now escapes. A caller-supplied callback is a target the compiler cannot examine, so it must assume the reference is retained. The proof that kept the value in the frame is withdrawn and every iteration allocates a real block.

open as a page

Should a latency-critical component rely on the compiler proving its temporaries non-escaping, or be designed so those temporaries never exist?

level: principalimportance: should knowfreq 33%

basics

~20 s

Split the codebase. On paths with a hard latency budget, shape the interfaces so the temporary is never constructed, because the proof is an optimization with no source marker and an ordinary edit can withdraw it. Everywhere else, rely on it.

open as a page

What does it mean that a non-escaping value is scalar-replaced rather than allocated anywhere at all?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

The value is deleted as an object: each field becomes an ordinary local kept in a register or a stack slot. Nothing is allocated, nothing is released, and the value has no address and no identity.

open as a page