skip to content

A shrinker that deletes unreferenced code breaks a binder resolving members by name at run time — what can an analyser prove about generated binders instead?

level: seniorimportance: should knowfreq 52%

answer

  1. what can the analyser see
  2. reachability follows references from roots
  3. a string is not a reference
  4. emitted assignment is an edge; a name is not
  5. the keep list is unenforced truth

basics

~20 s

Generated binders name each member in ordinary code, so a reachability analysis follows that reference and keeps the member. A name resolved from a string while running is not a reference, so the analyser sees an unused member and deletes it.

solid answer

~40 s

A shrinker keeps what is reachable from a set of roots by following references in code. An emitted binder *is* such a reference: the assignment names the member, so the edge exists, the analysis keeps the member, and a build that must know everything up front sees the same graph. A binder that assembles a member name from a configuration key has no edge at all — the string is data, not a reference — so the member looks unused and goes, and the failure shows up at run time on whichever path binds it. The usual escape hatch is a hand-maintained list of names to keep. It works, but it is a second source of truth that drifts silently and grows with every dependency that binds by name.

code

pseudocode · 7 lines
pseudocode
// nothing here names the member; no edge exists
key = values.keys.first()                 // data, from a file
member = lookupMemberByName(typeOf(target), key)
setMember(target, member, values.get(key))

// emitted: an ordinary reference the analysis can follow
target.port = toInteger(values.get("port"))

go deeper

for a junior

Remember that tools which delete unused code decide what is used by following references written in code, and that a name pieced together from data is not one of those references.

for a middle

Explain the edge: an emitted assignment names the member, so the analysis keeps it; a lookup by name gives it nothing to follow, so the member is removed and the break lands at run time.

for a senior

Diagnose it out loud. Say how you would confirm the member was removed rather than mis-bound, and weigh emitting the binder against maintaining a keep list that nothing validates.

for a principal

Treat unenforced keep lists as a liability that grows with the dependency graph, and decide as a policy whether artifacts that must survive whole-program builds are allowed to bind by name at all.

## What the analyser is actually doing A dead-code shrinker, and a build that compiles a whole program ahead of time, both answer one question: **what can be reached?** They start from roots — the entry point, anything declared as an entry — and follow references in the code. A member that nothing references is, as far as the analysis can tell, not part of the program, so it is removed or never emitted. That is not a heuristic; it is the guarantee that makes the technique worth anything, because without it nothing could ever be deleted. The important consequence: **reachability is decided over references that exist in code, and a string is not a reference.** ## Why a name resolved while running is invisible A binder that looks a member up by name typically builds that name from data — a configuration key, a command-line argument, a file. To the analysis: - The lookup call is reachable, because the binder is called. - The **member** is not, because nothing in any compilation unit names it. - Nothing distinguishes a key that will be used from one that never appears, so the analysis cannot even be conservative in a useful way — every bound member of every bound type is equally invisible. So members are removed, or entire types are never included, and the program that survives is not the program that was tested. The failure surfaces at run time, on the path that binds, which is often a rarely exercised option. ## What emitted code gives the analysis | The question | Emitted binder | Binder resolving names while running | |---|---|---| | Is this member referenced? | yes — the assignment names it | no reference exists anywhere | | Which types can be bound? | the set the generator emitted for, enumerable in code | unknown; any type that reaches the binder | | Can this member be deleted? | no, the edge holds it | yes, and nothing warns you | | Where does a missing conversion surface? | the build, as a compile error | the run, when that member binds | | What holds the truth? | the emitted code, regenerated from the declaration | a list maintained beside the code | ## The escape hatch and what it costs When inspection must stay, the standard remedy is to tell the analyser which names to keep. The lifecycle of that list is predictable: 1. Someone lists the members the binder will reach for, so the shrunken artifact works. 2. The code changes — a member is added, renamed or moved — and the list does not, because nothing enforces the correspondence. 3. The drift is discovered when a user takes the path that binds the missing member, in a shipped artifact, long after the change. The list is a **second source of truth** with no compiler behind it. It also does not stay local: every dependency that binds by name contributes entries, so the list grows with the dependency graph rather than with your own code. Emitted binders avoid the list precisely because the generator re-derives the references from the declaration on every build, so the two cannot drift. ## Where generation is not automatically provable either Generation is not a magic word, and the honest version of the answer says so: - If the emitted binders are registered into a table and then selected by a run-time key, the analysis still needs the table's construction to be ordinary code naming each binder. Built that way, every edge exists; built by inspecting for binders at startup, the problem is back. - If the generator emits for a type nobody references, the analysis may delete the binder as unreachable — which is usually correct, and occasionally a surprise. - Whole-program analysis can only see the program it is given: code fetched after the build sits outside it either way. ## The second thing analysability buys Provability is not only about shrinkers. An engineer reading the code can follow an emitted assignment to the member it writes, and a tool that renames members updates emitted call sites on the next build because they are real references. A name living in a string is outside all of that: nothing points at it, so nothing maintains it, and a reader looking for who writes a member finds no call site at all. ## How it is asked Usually as a failure story: *it worked in development and the shrunken build broke.* The answer the interviewer wants is the mechanism — a reachability analysis follows references, a run-time name is not a reference, so the member was removed — followed by the two options and their real cost: emit the binder so the edge exists, or maintain a list of names and accept that it is unenforced.

  • Why does the failure from an over-aggressive shrink usually appear late rather than in a smoke test?
    Because binding is per member and per option. The members exercised by a smoke run are the common ones; a rarely supplied option binds a member nobody touched, so the missing member is only reached when a real user sets it. The artifact is wrong from the moment it is built, but nothing is there to say so.
  • Does emitting binders guarantee a whole-program build can see everything?
    No. It removes one class of invisibility — the member named by a string — but the analysis still only sees code it is given. A binder chosen from a table is provable only if the table itself is built by ordinary code naming each entry; if the table is filled by inspecting at startup, the edge is gone again.
  • What can the generator check that a run-time binder cannot check at all?
    Anything decidable from the declaration: a member whose declared type has no conversion, a duplicate key, a key naming no member. The generator can fail the build and point at the declaration. A run-time binder can only discover the same facts when it reaches that member, in whatever environment it happens to be running.

saying these in an interview costs you the question

  • Says the analyser should just keep everything that might be bound
  • Believes a member name built from a key counts as a reference to it
  • Thinks a keep list solves it, without noting nothing enforces the list
  • Claims generation makes every path statically provable, whatever the registry does
  • Blames the shrinker for a bug rather than naming the missing edge
  • Expects the break to show up in the build rather than in a shipped run