How does the points-to static analysis compute the reachable set, and what starts it?
answer
- roots: main + registered reflection/JNI/serialization
- worklist to a fixed point
- types flow along edges resolve virtual calls
- unreachable = dead-code-eliminated
- hints add synthetic roots
basics
~20 sIt starts from entry points like main and registered roots, then follows every possible call and field access transitively until nothing new is discovered. Whatever it reaches is kept and compiled; the rest is removed.
solid answer
~50 sThe points-to analysis is a whole-program static analysis run by the native-image builder. It seeds a worklist with entry points: the application's main method plus any explicitly registered reflection, JNI, and serialization roots. It then propagates type information along call edges and field reads/writes, discovering which concrete types can flow to each call site and therefore which method implementations can actually be invoked. This continues to a fixed point — iterating until no new reachable types or methods appear. The resulting reachable set defines the closed world: those methods are AOT-compiled, those types get their fields and metadata included, and everything else is dead-code-eliminated. Because the analysis is conservative and static, dynamic dispatch through reflection or proxies is invisible unless you declare it, which is exactly why reachability metadata (Spring RuntimeHints, META-INF/native-image JSON) exists — it injects extra roots the analysis would otherwise miss.
go deeper
Enough to say it starts at main and follows calls to decide what to keep.
Should describe roots, fixed-point iteration, and dead-code elimination of the rest.
Explain how type-flow sharpens virtual dispatch and why hints act as synthetic roots.
Discuss conservativeness, image-size/build-time trade-offs, and designing code to stay analysis-friendly.
## What 'points-to analysis' means A **points-to analysis** is a form of static dataflow analysis that computes, for every variable/field/parameter, the set of object **types** it could possibly point to at runtime. GraalVM runs this over your **entire program** (application + libraries + JDK) — hence 'whole-program' analysis. ## Seeding: the roots The analysis needs starting points, called **roots** or **entry points**: - The application **`main` method** (and framework-provided entry points). - **Explicitly registered roots** from reachability metadata: reflectively accessed types, JNI-called methods, serialized types, and dynamic-proxy interfaces. In Spring these come from **`RuntimeHints`** contributed by `RuntimeHintsRegistrar`s and Spring's AOT engine. Without a root, a method is simply **not part of the world**. ## The propagation (fixed-point iteration) From the roots the analysis maintains a **worklist** and repeatedly: 1. Marks a method reachable and scans its body. 2. For each call, uses the **types that can flow to the receiver** to resolve which concrete overrides are callable (this is where points-to sharpens virtual dispatch — only implementations whose type can actually reach the call site are included). 3. For each `new`/allocation, marks that type **instantiated**, which can unlock more virtual targets elsewhere. 4. Propagates types through fields, parameters, and returns. Steps repeat until a **fixed point** — an iteration that discovers nothing new. Termination is guaranteed because the type universe is finite and the reachable set only grows. ## The output - **Reachable methods** → AOT-compiled to native code. - **Instantiated types** → included with the field layout and metadata they need. - **Unreachable everything** → removed. This is why native images are small and why an accidentally-unreferenced-but-reflected class disappears. ## Why dynamic features break it `Class.forName(userInput)` or `Proxy.newProxyInstance(...)` produce types/methods that **do not appear as static edges**. The analysis literally cannot see them, so the targets are absent from the closed world and you get a runtime failure. **Reachability metadata adds synthetic roots/edges** so those targets get pulled into the analysis. ## Spring's role Spring **AOT processing** (`ApplicationContextAotGenerator`, `BeanRegistrationAotProcessor`, `BeanFactoryInitializationAotProcessor`) runs before native compilation. It: - Emits generated Java that instantiates beans **directly** (no reflective scanning), giving the analysis clean static edges. - Registers `RuntimeHints` for the reflection/resources/proxies the remaining framework code still needs. The **tracing agent** (`-agentlib:native-image-agent=config-output-dir=...`) is the fallback: run the app on a JVM, exercise the dynamic paths, and it records the metadata to feed the analysis. ## Gotchas - The analysis is **conservative but not omniscient** about strings — computed reflection targets need explicit hints. - Adding hints **grows the image**; over-broad reflection config (e.g. registering a whole package) bloats size and slows the build. - Build time scales with reachable-set size; big apps have long native compiles.
- Why does registering a type for reflection change what the analysis keeps?It adds a synthetic root/edge for that type's constructors/methods/fields, so the points-to analysis now treats them as reachable and keeps them in the image instead of eliminating them.
- How does Spring AOT reduce reliance on the analysis seeing dynamic edges?It generates explicit Java code that instantiates and wires beans directly, turning reflective/scanning behavior into plain static call edges the analysis can follow naturally.
saying these in an interview costs you the question
- Saying the analysis runs at application startup rather than build time
- Claiming it can resolve Class.forName from a runtime string
- Thinking every method on the classpath is always included
- Confusing points-to reachability with simple 'is on the classpath'