From an architecture standpoint, how do you design plugin/redeploy systems to be resistant to classloader leaks?
answer
- Loader isolation: one disposable child loader per unit
- No strong refs from longer-lived to shorter-lived loaders
- Weak refs in shared caches + explicit deregistration
- Teardown: stop threads, clear ThreadLocals, deregister drivers/MBeans/hooks
- Prove collectibility: WeakReference to old loader + forced GC test
basics
~20 sGive each plugin/app its own classloader, never let long-lived infrastructure hold strong references to plugin objects (use weak references and explicit deregistration), and clean up threads, ThreadLocals, drivers, and listeners on undeploy. Test by redeploying many times and confirming the old classloader is collected.
solid answer
~50 sThe core principle is **loader isolation with clean teardown**: each deployable unit gets its own child classloader so its classes can be discarded as a unit, and nothing outside that loader holds a strong reference into it past undeploy. Concretely: shared/parent-loaded registries (caches, listener lists, service locators) should reference plugin objects **weakly** or expose explicit deregistration that the undeploy lifecycle calls; the container must stop threads the plugin started, clear ThreadLocals on pooled threads, deregister JDBC drivers and JMX MBeans, and remove shutdown hooks. Frameworks like OSGi formalize this with strict module boundaries and lifecycle callbacks. You verify resistance by **redeploying in a loop in a test** and asserting (via a heap dump or a weak reference to the old loader plus a forced GC) that the previous classloader is actually collected. Treat 'the old loader is collectible after undeploy' as a testable invariant, not a hope.
go deeper
Aware that plugins get their own classloaders and that cleanup matters; not expected to architect the system.
Can list common teardown steps (stop threads, remove ThreadLocals, deregister drivers) for a redeploy.
Designs weak-reference registries and explicit deregistration, and knows OSGi addresses this formally.
Treats loader collectibility as a tested system invariant, designs shared infra to never trap a loader, and chooses/encodes module boundaries with their trade-offs.
## The architectural problem Long-running JVMs that load and unload code at runtime — application servers, plugin hosts, scripting engines, hot-reload dev tools, multi-tenant platforms — depend on **classloaders being discardable**. A classloader leak defeats the entire model: you can't truly remove code without a JVM restart. So leak-resistance is a *design property*, not just a bug to chase later. ## Principle 1 — Isolate each unit in its own loader Give every plugin/app/tenant its **own classloader** (typically a child of a shared parent that holds the framework + shared libs). This makes each unit's classes a single collectible bundle and gives each unit its own namespace (the loader-as-namespace model: (name, loader) is the class identity). The shared parent loads only stable, long-lived code; volatile code lives in the disposable child. ## Principle 2 — No strong references from longer-lived to shorter-lived loaders The golden rule: **an object loaded by a longer-lived classloader must not strongly retain an object loaded by a shorter-lived one** past the latter's lifetime. Mechanisms: - **Weak references / weak collections** in shared caches and registries (`WeakHashMap`, `WeakReference`) so the GC can reclaim plugin objects. - **Explicit deregistration APIs** invoked by the undeploy lifecycle for anything that must be strongly held while active (listeners, callbacks, service entries). - **Avoid stashing plugin objects in JDK-global singletons** (`DriverManager`, `ImageIO`, `Logger`, security providers, `ThreadLocal`s on shared pools) — these are GC roots that outlive everything. ## Principle 3 — Own the teardown lifecycle Undeploy must actively unwind everything the unit touched: - **Stop threads** the unit started (a running thread is a GC root referencing its classes); interrupt + join, don't leak daemon threads. - **Clear ThreadLocals** on container-pooled threads (or design so the unit never sets them on pooled threads). - **Deregister JDBC drivers** (`DriverManager.deregisterDriver`), **JMX MBeans**, **shutdown hooks**, security providers, and any framework callbacks. - **Close resources** (connection pools, file handles, native handles) — often via `AutoCloseable`/try-with-resources and a `Cleaner`. - **Flush JVM/library caches** that may key on plugin classes (some reflection/introspection caches). ## Principle 4 — Make collectibility a tested invariant Don't assume it works — **prove it**: - In an integration test, deploy → use → undeploy, hold the old classloader through a `WeakReference`, drop all strong references, force GC (`System.gc()` a few times / use the test-friendly `Reference` APIs), and assert `weakRef.get() == null`. - Or run the redeploy loop N times under a capped `-XX:MaxMetaspaceSize` and assert no `OutOfMemoryError: Metaspace` and that `-Xlog:class+unload` shows old classes unloading. This converts a notoriously subtle, late-discovered failure into a fast deterministic check. ## Principle 5 — Prefer mature module systems where applicable **OSGi** and the **Java Platform Module System (JPMS)** (to a lesser degree, since JPMS doesn't do hot-unload) provide formal boundaries. OSGi in particular gives each bundle its own loader, explicit imports/exports, and start/stop lifecycle hooks — turning the ad-hoc rules above into enforced structure. For in-house plugin systems, deliberately mirror these guarantees. ## Trade-offs to reason about - **Weak references add complexity and surprise** (entries can vanish); use them where correctness tolerates it (caches), not where you need guaranteed retention. - **Strict isolation can duplicate classes** across loaders (memory cost) and complicate sharing of types across plugins — design a clear parent for shared API types so plugins interoperate without leaking. - **Aggressive teardown** can race with in-flight work; quiesce before unloading. ## The principal-level framing A staff/principal engineer treats 'every deployable unit's classloader is provably collectible after undeploy' as a **system invariant** with a test that guards it, designs shared infrastructure to never become the GC root that traps a loader, and chooses (or builds toward) a module system that encodes these boundaries rather than relying on every contributor to remember the rules.
- How would you write an automated test that a redeploy doesn't leak the classloader?After undeploy, keep only a WeakReference to the old classloader, drop all strong references, force GC, and assert the weak reference clears (weakRef.get() == null). Alternatively, loop the redeploy N times under a capped MaxMetaspaceSize and assert no Metaspace OOM while -Xlog:class+unload shows old classes unloading.
- Why does OSGi handle this better than an ad-hoc plugin loader?OSGi gives each bundle its own classloader with explicitly declared imports/exports and start/stop lifecycle callbacks, so module boundaries and teardown are enforced by the framework rather than left to each developer to remember — reducing accidental cross-loader strong references.
saying these in an interview costs you the question
- Stashing plugin objects in JDK-global singletons (DriverManager, ThreadLocal on shared pools)
- Assuming undeploy 'just works' without a test that the old loader is collected
- Using strong-referenced shared caches/listener lists with no deregistration
- Forgetting to stop plugin-started threads (each is a GC root)