skip to content

A Tomcat instance that is redeployed several times a day eventually dies with java.lang.OutOfMemoryError: Metaspace, and a full restart buys another few days. What happens on each redeploy to cause this, and how would you find and fix the cause?

level: seniorimportance: should knowfreq 40%

answer

  1. one loader per deployment
  2. one outside reference pins everything
  3. threads and thread-locals outlive the app
  4. Tomcat already logs the culprit

basics

~20 s

Each deployment gets a fresh web application class loader that should become garbage when the context stops. One reference from outside the application — an unstopped thread, a ThreadLocal, a registered JDBC driver — pins it, so every class it loaded stays in metaspace and each redeploy adds another full copy.

solid answer

~50 s

Tomcat creates a new class loader per deployment, so a redeploy is really "stop the old context, throw its loader away, build a new one". The loader can only be collected if nothing outside the application still points at it, and a single stray reference keeps it — and therefore every class it defined — alive; do that ten times and metaspace is full. The usual culprits are threads or executors the application started and never stopped, `ThreadLocal` entries left on Tomcat's request-processing threads, a JDBC driver loaded from `WEB-INF/lib` and registered with the JVM-wide `DriverManager`, shutdown hooks, and caches in shared libraries keyed by the thread context class loader. Tomcat tells you: at context stop it logs warnings naming the offending thread, ThreadLocal or driver, and the Manager application has a "Find leaks" diagnostic. The fix is to release those in `ServletContextListener.contextDestroyed`, move the driver to `$CATALINA_BASE/lib`, and — operationally — stop hot-redeploying long-lived JVMs at all.

code

java · 27 lines
java
import jakarta.servlet.ServletContextEvent;
import jakarta.servlet.ServletContextListener;
import jakarta.servlet.annotation.WebListener;
import java.sql.Driver;
import java.sql.DriverManager;
import java.sql.SQLException;
import java.util.Enumeration;

@WebListener
public class DriverCleanupListener implements ServletContextListener {

    @Override
    public void contextDestroyed(ServletContextEvent event) {
        ClassLoader webapp = Thread.currentThread().getContextClassLoader();
        Enumeration<Driver> drivers = DriverManager.getDrivers();
        while (drivers.hasMoreElements()) {
            Driver driver = drivers.nextElement();
            if (driver.getClass().getClassLoader() == webapp) {
                try {
                    DriverManager.deregisterDriver(driver);
                } catch (SQLException ignored) {
                    // nothing further can be done during shutdown
                }
            }
        }
    }
}

go deeper

for a junior

Know that each deployment gets its own class loader and that memory is not reclaimed unless the old one becomes garbage — this is why repeated hot redeploys eventually exhaust the JVM.

for a middle

Name the concrete pinning references — unstopped threads, ThreadLocals on container threads, a driver registered with DriverManager — and explain why one of them retains every class the application loaded.

for a senior

Drive the diagnosis: read Tomcat's stop-time leak warnings, use the Manager's leak check, take a heap dump and follow the path to GC roots, then fix the cleanup in contextDestroyed rather than raising a limit.

for a principal

Decide whether hot redeploy belongs in production at all. Restarting the JVM per release removes this class of failure entirely; if redeploy stays, own the standard for shutdown hygiene and how it is enforced across teams.

## Why redeploy needs a new class loader A deployed Tomcat context owns a **web application class loader** that defines every class in `WEB-INF/classes` and `WEB-INF/lib`. Java has no way to unload one class, so replacing an application means discarding the whole loader and building another. On redeploy Tomcat stops the context, drops its reference to the old loader, expands the new WAR and starts a new context with a new loader. A class loader is reachable *from* every class it defined, and every class is reachable from every instance of it. So the graph is all-or-nothing: if any live object anywhere in the JVM holds a reference into the old application, the loader survives, and with it all its classes' metadata in **metaspace** (the native region that replaced PermGen in Java 8), its static fields, and everything those retain on the heap. Repeat per release and metaspace grows monotonically until `OutOfMemoryError: Metaspace`. ## The classic pinning references - **Threads the application started.** A raw `Thread`, a `ScheduledExecutorService`, a library's background worker. A running thread is a GC root, and its `contextClassLoader` — plus its `Runnable`'s class — is the webapp loader. - **ThreadLocals on container threads.** Tomcat's request-processing threads outlive the application. A `ThreadLocal` whose *key type* or *value* is a webapp class, set during a request and never removed, keeps the loader alive through the thread's `ThreadLocalMap`. - **JDBC drivers.** `java.sql.DriverManager` is loaded by the JVM and lives forever. A driver in `WEB-INF/lib` registers itself there on first use and is never unregistered. - **Shutdown hooks** registered by the application with `Runtime.addShutdownHook`. - **Caches in shared libraries** keyed by the thread context class loader — logging frameworks, serialization caches, bean introspection caches, `ObjectStreamClass` caches. The library sits in `$CATALINA_BASE/lib`, so it outlives every application, while its cache entry points into one. - **JMX MBeans, RMI targets, custom `URLStreamHandler`s, `java.beans` introspection results** — anything registered in a JVM-global registry. ## Confirming it rather than guessing 1. **Read the logs at context stop.** Tomcat actively detects several of these and logs a warning naming the culprit: that the application appears to have started a thread and failed to stop it; that it created a `ThreadLocal` with a key of a given type and failed to remove it; that it registered a JDBC driver but failed to unregister it. These messages are the fastest possible diagnosis and are routinely ignored. 2. **Use the Manager application's "Find leaks" diagnostic.** It forces a full collection and lists contexts whose class loaders are still reachable after being stopped — a definitive yes/no per application. 3. **Take a heap dump after a few redeploys.** Count instances of Tomcat's web application class loader class: more than one live loader per deployed context is the leak. Then ask the analyser for the **path to GC roots** on the extra loaders, excluding weak references — that path names the exact field holding it. 4. **Watch metaspace over time.** Enable GC logging and native memory tracking, or read the metaspace gauge from JMX; a saw-tooth that never returns to its baseline after each deploy is the signature. ## Fixing it In the application, do the symmetric cleanup in a `ServletContextListener`: - shut down every executor and join every thread you started, with a bounded wait; - remove `ThreadLocal` values in a `finally` (or a filter that cleans up after the request); - deregister JDBC drivers whose class loader is the webapp's; - deregister JMX MBeans, cancel timers, close pools and clients. Structurally: move libraries the container itself uses — JDBC drivers above all — into `$CATALINA_BASE/lib` so they are loaded once and never bound to a webapp loader. Tomcat also offers **mitigations**, which are not fixes. `JreMemoryLeakPreventionListener` in `server.xml` pre-initialises JVM-level singletons that would otherwise be initialised by a webapp and pin its loader. Context attributes such as `clearReferencesStopThreads`, `clearReferencesStopTimerThreads` and `clearReferencesObjectStreamClassCaches` let Tomcat try to clean up on the application's behalf at stop time; `renewThreadsOnStartStop` on the Context recycles connector threads so their `ThreadLocal`s are dropped. They can hide a leak, sometimes violently — stopping threads at arbitrary points is not safe — so treat them as a stopgap while you fix the code. ## The operational answer The honest senior answer is that hot redeploy into a long-lived JVM is a development convenience. In production, deploy by starting a new JVM with the new artifact and shutting down the old one; then the class loader leak simply cannot accumulate, and the application's cleanup code — which you should still write, because it also governs graceful shutdown — is exercised once rather than continuously. Reserve `autoDeploy` and Manager redeploys for environments where a stale JVM costs nothing.

  • Why does a ThreadLocal set during a request outlive the application that set it?
    Because the thread does. Tomcat's request-processing threads belong to the connector's pool and are reused across contexts, so a value left in a thread's ThreadLocalMap stays reachable after the application stops. If either the key's type or the value's type came from the webapp loader, that loader is pinned. Remove the value in a finally block, or recycle the threads on stop.
  • How would you prove which reference is holding the old class loader?
    Take a heap dump a few redeploys in, find the instances of Tomcat's web application class loader for that context, and ask the analyser for the path to GC roots excluding weak and soft references. That path names the exact static field, thread or cache entry holding it. Tomcat's own stop-time warnings usually point at the same object first and cost nothing to read.
  • Tomcat's clearReferencesStopThreads can kill threads a stopped application left running. Why is that only a mitigation?
    Because it stops threads at an arbitrary point, which can leave locks held and data half-written — the same reason Thread.stop was deprecated. It buys time on an instance you cannot redeploy safely, but it treats the symptom: the application still has no shutdown path, so it will also misbehave on a real shutdown, in a container restart, or during graceful drain.

Undeploying is like moving out of a flat: you cannot hand back the keys while one of your appliances is still plugged into the building's shared power, and the landlord has to keep the whole flat reserved for it.

saying these in an interview costs you the question

  • Raising MaxMetaspaceSize and calling it fixed
  • Blaming the garbage collector for not collecting classes
  • Assuming Tomcat cleans up whatever the app leaves running
  • Ignoring the leak warnings Tomcat logs at context stop
  • Thinking heap dumps cannot show metaspace-class leaks

context