skip to content

How does the points-to static analysis compute the reachable set, and what starts it?

level: middleimportance: should knowfreq 40%

answer

  1. roots: main + registered reflection/JNI/serialization
  2. worklist to a fixed point
  3. types flow along edges resolve virtual calls
  4. unreachable = dead-code-eliminated
  5. hints add synthetic roots

basics

~20 s

It starts from entry points like main and registered roots, then follows every possible call and field access transitively until nothing new is discovered. Whatever it reaches is kept and compiled; the rest is removed.

solid answer

~50 s

The points-to analysis is a whole-program static analysis run by the native-image builder. It seeds a worklist with entry points: the application's main method plus any explicitly registered reflection, JNI, and serialization roots. It then propagates type information along call edges and field reads/writes, discovering which concrete types can flow to each call site and therefore which method implementations can actually be invoked. This continues to a fixed point — iterating until no new reachable types or methods appear. The resulting reachable set defines the closed world: those methods are AOT-compiled, those types get their fields and metadata included, and everything else is dead-code-eliminated. Because the analysis is conservative and static, dynamic dispatch through reflection or proxies is invisible unless you declare it, which is exactly why reachability metadata (Spring RuntimeHints, META-INF/native-image JSON) exists — it injects extra roots the analysis would otherwise miss.

go deeper

for a junior

Enough to say it starts at main and follows calls to decide what to keep.

for a middle

Should describe roots, fixed-point iteration, and dead-code elimination of the rest.

for a senior

Explain how type-flow sharpens virtual dispatch and why hints act as synthetic roots.

for a principal

Discuss conservativeness, image-size/build-time trade-offs, and designing code to stay analysis-friendly.

## What 'points-to analysis' means A **points-to analysis** is a form of static dataflow analysis that computes, for every variable/field/parameter, the set of object **types** it could possibly point to at runtime. GraalVM runs this over your **entire program** (application + libraries + JDK) — hence 'whole-program' analysis. ## Seeding: the roots The analysis needs starting points, called **roots** or **entry points**: - The application **`main` method** (and framework-provided entry points). - **Explicitly registered roots** from reachability metadata: reflectively accessed types, JNI-called methods, serialized types, and dynamic-proxy interfaces. In Spring these come from **`RuntimeHints`** contributed by `RuntimeHintsRegistrar`s and Spring's AOT engine. Without a root, a method is simply **not part of the world**. ## The propagation (fixed-point iteration) From the roots the analysis maintains a **worklist** and repeatedly: 1. Marks a method reachable and scans its body. 2. For each call, uses the **types that can flow to the receiver** to resolve which concrete overrides are callable (this is where points-to sharpens virtual dispatch — only implementations whose type can actually reach the call site are included). 3. For each `new`/allocation, marks that type **instantiated**, which can unlock more virtual targets elsewhere. 4. Propagates types through fields, parameters, and returns. Steps repeat until a **fixed point** — an iteration that discovers nothing new. Termination is guaranteed because the type universe is finite and the reachable set only grows. ## The output - **Reachable methods** → AOT-compiled to native code. - **Instantiated types** → included with the field layout and metadata they need. - **Unreachable everything** → removed. This is why native images are small and why an accidentally-unreferenced-but-reflected class disappears. ## Why dynamic features break it `Class.forName(userInput)` or `Proxy.newProxyInstance(...)` produce types/methods that **do not appear as static edges**. The analysis literally cannot see them, so the targets are absent from the closed world and you get a runtime failure. **Reachability metadata adds synthetic roots/edges** so those targets get pulled into the analysis. ## Spring's role Spring **AOT processing** (`ApplicationContextAotGenerator`, `BeanRegistrationAotProcessor`, `BeanFactoryInitializationAotProcessor`) runs before native compilation. It: - Emits generated Java that instantiates beans **directly** (no reflective scanning), giving the analysis clean static edges. - Registers `RuntimeHints` for the reflection/resources/proxies the remaining framework code still needs. The **tracing agent** (`-agentlib:native-image-agent=config-output-dir=...`) is the fallback: run the app on a JVM, exercise the dynamic paths, and it records the metadata to feed the analysis. ## Gotchas - The analysis is **conservative but not omniscient** about strings — computed reflection targets need explicit hints. - Adding hints **grows the image**; over-broad reflection config (e.g. registering a whole package) bloats size and slows the build. - Build time scales with reachable-set size; big apps have long native compiles.

  • Why does registering a type for reflection change what the analysis keeps?
    It adds a synthetic root/edge for that type's constructors/methods/fields, so the points-to analysis now treats them as reachable and keeps them in the image instead of eliminating them.
  • How does Spring AOT reduce reliance on the analysis seeing dynamic edges?
    It generates explicit Java code that instantiates and wires beans directly, turning reflective/scanning behavior into plain static call edges the analysis can follow naturally.

saying these in an interview costs you the question

  • Saying the analysis runs at application startup rather than build time
  • Claiming it can resolve Class.forName from a runtime string
  • Thinking every method on the classpath is always included
  • Confusing points-to reachability with simple 'is on the classpath'

context