skip to content

A Java service on Tomcat behind a reverse proxy starts returning gateway timeouts to users, yet Tomcat's own logs look quiet, CPU is near idle and the host has plenty of memory. How do you confirm the connector's thread pool is exhausted, and what do you do once you have?

level: seniorimportance: must knowfreq 57%

answer

  1. idle CPU, rising latency
  2. count the exec threads in a dump
  3. one stack frame shared by all of them
  4. currentThreadsBusy pinned at maxThreads
  5. the access log is written at the end

basics

~20 s

Take a thread dump and count the connector's exec threads: if all maxThreads are alive and blocked in the same downstream call, the pool is exhausted. Confirm with the ThreadPool MBean's currentThreadsBusy, then bound the downstream call rather than enlarging the pool.

solid answer

~50 s

Idle CPU with rising latency is the signature of blocked threads, not of overload. Confirm it two ways. First, a thread dump: count threads named `http-nio-8080-exec-`; if there are exactly `maxThreads` of them and most share a stack frame parked in a downstream HTTP or JDBC call, the pool is the bottleneck. Second, JMX: the connector's `Catalina:type=ThreadPool` MBean exposes `currentThreadsBusy`, `currentThreadCount` and `connectionCount` — busy pinned at `maxThreads` while connections climb is the same story in a metric you can alert on. Tomcat's logs look quiet because the access log entry is written when the response completes, so in-flight requests are invisible there while the proxy has already given up on them. The fix is upstream of the pool: put a real timeout on the dependency that is blocking, add a circuit breaker or bulkhead so one slow call cannot take the whole pool, and only then reconsider sizing. Raising `maxThreads` alone just aims more concurrency at whatever is already failing.

code

bash · 11 lines
bash
# 1. how many request threads exist, and are they moving?
jcmd "$PID" Thread.print > /tmp/dump1.txt
grep -c 'http-nio-8080-exec-' /tmp/dump1.txt

# 2. what are they all waiting on?
grep -A 3 'http-nio-8080-exec-' /tmp/dump1.txt | grep -m 20 'at '

# 3. take a second dump and compare: stuck threads do not move
jcmd "$PID" Thread.print > /tmp/dump2.txt
diff <(grep 'http-nio-8080-exec-' /tmp/dump1.txt) \
     <(grep 'http-nio-8080-exec-' /tmp/dump2.txt)

go deeper

for a junior

Know that a thread dump exists and what to look for in it: threads named http-nio-<port>-exec-N, and whether they are running or waiting.

for a middle

Explain why saturation shows as latency with zero Tomcat errors, name currentThreadsBusy on the ThreadPool MBean, and connect the pool ceiling to maxThreads.

for a senior

Drive the whole diagnosis: correlate idle CPU with blocked threads, read the shared stack frame to name the dependency, and fix it with timeouts and isolation rather than a bigger pool.

for a principal

Own the standing defences — saturation metrics and dependency timeout budgets set deliberately per service, bulkheads for untrusted dependencies, and a retry policy at the edge that cannot amplify a backend stall.

## Read the symptom before touching config Three facts in the report already narrow the diagnosis. Users see gateway timeouts, which means the *proxy* gave up waiting — the failure is being reported by the hop in front, not generated by Tomcat. CPU is idle, so nothing is computing; threads are waiting. And Tomcat's log is quiet, which is not evidence of health: Tomcat writes an access-log entry when the response is complete, so requests still in flight — exactly the ones timing out — have not been logged yet. Every one of those points to the same shape: request threads alive but blocked. ## Confirming it with a thread dump A thread dump is the definitive evidence, and it costs almost nothing to take. ```bash jcmd <pid> Thread.print > /tmp/dump1.txt grep -c 'http-nio-8080-exec-' /tmp/dump1.txt ``` What you are looking for is two things at once: the **count** of `http-nio-8080-exec-` threads equalling the connector's `maxThreads`, and a **shared stack frame** near the top of most of them. If two hundred threads are all sitting in a socket read inside the same HTTP client call to the same downstream service, you have both the diagnosis and the culprit in one artefact. Take two or three dumps a few seconds apart: threads that are progressing move between dumps, threads that are stuck do not. The stack tells you which category of stall this is. Blocked in a socket read on a downstream call means a dependency is slow or has stopped answering. Blocked waiting for a JDBC connection means the *database pool*, not the connector pool, is the real ceiling and the connector pool is merely the queue behind it. Blocked on a `synchronized` monitor held by one thread means a lock convoy, and the holder's stack names the offending code. ## Confirming it continuously with JMX A dump is a snapshot; you also want this as a standing signal. The connector registers a `Catalina:type=ThreadPool` MBean whose `currentThreadsBusy` is the number of exec threads currently processing requests, alongside `currentThreadCount` (threads in the pool) and `connectionCount`. `currentThreadsBusy / maxThreads` is the single most valuable Tomcat saturation metric there is, and the sequence during an incident is characteristic: busy climbs to `maxThreads` and pins there, connections keep climbing behind it, throughput flattens, latency grows without bound, error rate at Tomcat stays zero. Alerting on sustained busy ratio catches this long before customers do. ## Why the errors appear at the proxy first Once all exec threads are occupied, Tomcat still accepts new connections — up to `maxConnections`, with the OS accept queue (`acceptCount`) behind that — and simply does not process them yet. So the connections succeed, the requests are never answered, and the proxy's read timeout fires. From the proxy's logs it looks like a backend that stopped responding; from the host it looks like an idle machine. This is precisely why a thread-exhausted Tomcat is repeatedly misfiled as a network problem, and why knowing the layering wins the interview. ## Fixing it in the right order 1. **Bound the blocking call.** The root cause is nearly always an unbounded wait. Connect and read timeouts on the HTTP client, a query timeout and a pool-acquisition timeout on JDBC. A request that cannot exceed two seconds cannot hold a thread for five minutes. This is the fix that actually works. 2. **Isolate.** If one endpoint depends on a flaky service, do not let it consume the pool the healthy endpoints share. A circuit breaker sheds load once the dependency is failing; a bulkhead — a separate bounded pool for that call, or a separate connector with its own `<Executor>` — caps the damage at a fraction of capacity. 3. **Shed load deliberately.** A shorter accept backlog and a client-side timeout shorter than the work's worst case turn a slow death into fast, retryable failures. Serving a response nobody is waiting for is pure waste. 4. **Only then, resize.** If threads are genuinely productive and merely numerous — lots of short waits, not one long one — a larger `maxThreads` is legitimate. Confirm first that the downstream can absorb the extra concurrency, because otherwise you have moved the queue, not removed it. ## What to change so it is caught next time Export the busy-thread ratio and the connection count as metrics, and log request duration so slow-but-successful requests are visible before they become timeouts. Have an operator-triggerable thread dump documented in the runbook. And write down the timeout for each downstream dependency as a deliberate number — the most common root cause here is not a missing metric but a client library that defaulted to waiting forever.

  • The dump shows exec threads waiting to acquire a JDBC connection rather than in an HTTP call. What changes about your diagnosis?
    The connector pool is a symptom, not the constraint — the database pool is the real ceiling and the exec threads are queued behind it. Look at the connection pool's size, acquisition timeout and leak detection, and at whether a slow query is holding connections. Growing maxThreads here only lengthens the queue for the same connections.
  • Why is a proxy read timeout shorter than the worst-case backend request actively useful here?
    It stops work that nobody will consume from being retried forever and turns an unbounded wait into a fast failure the caller can act on. The complement matters too: the proxy must not retry blindly onto other instances, or one slow dependency becomes a self-inflicted load multiplier across the whole pool.
  • What single Tomcat metric would you alert on to catch this earlier?
    The ratio of `currentThreadsBusy` to `maxThreads` on the connector's ThreadPool MBean, sustained over a few minutes. It rises well before error rate moves, because Tomcat reports no errors at all while it is saturated — the errors surface only at the proxy once its own timeout fires.

saying these in an interview costs you the question

  • Concludes it must be a network fault because Tomcat logs nothing
  • Raises maxThreads as the first and only action
  • Reads idle CPU as proof the server is healthy
  • Never takes a thread dump because the JVM 'looks fine'
  • Adds retries at the proxy while the backend is already saturated

context