skip to content

On a Tomcat service, one endpoint spends nearly all of each request blocked on a slow third-party API and its exec threads are saturated. How do you decide between raising maxThreads, converting the endpoint to an async servlet, and giving it its own <Executor> or connector?

level: principalimportance: should knowfreq 36%

answer

  1. three levers, three different questions
  2. slow versus failing is the first fork
  3. async removes a ceiling, not a wait
  4. the pool was also your backpressure
  5. isolation is measured on the other endpoints

basics

~20 s

Decide by what the threads are waiting on and who else shares the pool. Async frees the exec thread during the wait and lifts a throughput ceiling; a separate Executor buys isolation; a bigger pool only helps when the dependency can absorb more concurrency and the threads themselves are cheap.

solid answer

~60 s

Start from the constraint. If threads are merely parked waiting on I/O and the dependency is healthy but slow, the exec pool is an artificial ceiling and `AsyncContext` removes it: the servlet returns the container thread immediately and completes the response from a callback, so concurrency is limited by the downstream client's own capacity instead of by `maxThreads`. If the dependency is *failing*, none of the three options is the fix — a timeout and a circuit breaker are, because otherwise async simply lets you queue far more doomed work. If other endpoints are healthy and being starved by this one, the immediate win is isolation: a dedicated `<Executor>` or a separate connector so the blast radius is a bounded slice of capacity rather than everything. Raising `maxThreads` is the last option and only defensible when the downstream can genuinely absorb the extra concurrency and per-thread cost is affordable. Whatever you choose, set the async timeout deliberately — the connector's `asyncTimeout` defaults to 30 seconds — so freed threads do not simply become an unbounded backlog.

code

xml · 10 lines
xml
<Executor name="partnerPool"
          namePrefix="partner-exec-"
          maxThreads="30"
          minSpareThreads="5"/>

<Connector executor="partnerPool"
           port="8081"
           protocol="HTTP/1.1"
           asyncTimeout="3000"
           connectionTimeout="20000"/>

go deeper

for a junior

Know that a synchronous servlet holds one request thread for the whole call, so a slow dependency consumes capacity even while nothing is computing.

for a middle

Explain what startAsync changes mechanically — the container thread is returned and the response completed from a callback — and the requirement that filters and the servlet support async.

for a senior

Choose between the levers from evidence: read the thread dump, separate slow from failing, apply timeouts and a bulkhead, and justify why a bigger pool is or is not the answer here.

for a principal

Own the trade explicitly: async removes an implicit backpressure mechanism that must be replaced, isolation costs utilisation to buy blast-radius control, and any capacity increase is a decision about someone else's service too.

## Name the constraint before choosing the lever The three options answer three different questions, and picking one before you know which question you are in is the classic failure here. - **maxThreads** answers *how much concurrency may execute at once*. - **async** answers *must a container thread be held while we wait*. - **a separate Executor or connector** answers *who suffers when this endpoint misbehaves*. So the first job is to characterise the wait. A thread dump tells you whether the exec threads are blocked in a socket read (waiting on the third party), waiting for a connection from a client pool (the pool is the real ceiling), or running (a genuine CPU constraint, where none of these three options helps). The distinction between *slow but working* and *failing* matters as much: for a dependency that is timing out or erroring, the correct response is a bounded timeout and a circuit breaker, and every capacity lever below merely increases the volume of work you throw at something already broken. ## What async actually buys, and what it costs With `request.startAsync()`, the servlet registers a callback, returns, and the exec thread goes back to the pool while the response is still pending. The request is completed later from whatever thread the asynchronous client uses. ```java @WebServlet(urlPatterns = "/quote", asyncSupported = true) public class QuoteServlet extends HttpServlet { @Override protected void doGet(HttpServletRequest req, HttpServletResponse resp) { AsyncContext ctx = req.startAsync(); ctx.setTimeout(3000); upstreamClient.fetchAsync().whenComplete((body, err) -> { try { ctx.getResponse().getWriter().write(body); } catch (Exception e) { /* handle */ } finally { ctx.complete(); } }); } } ``` The gain is real and structural: 200 exec threads no longer cap you at 200 concurrent in-flight requests. But it comes with conditions people routinely miss. - The whole path must be non-blocking. If the async callback calls a *blocking* HTTP or JDBC client on some other pool, you have moved the queue, not removed it. - Every filter in the chain must support async, and the servlet must declare `asyncSupported`. - Debugging and tracing get harder: the stack that fails is no longer the stack that received the request. - Removing the thread ceiling removes a **natural backpressure mechanism**. A saturated pool used to reject new work implicitly by making it wait; without it, you can accept vastly more concurrent requests than the downstream can serve, and turn a slow endpoint into an outage for the third party too. Async needs an explicit concurrency bound to replace the implicit one, plus a timeout — the connector's `asyncTimeout` defaults to 30000 ms, and `AsyncContext.setTimeout` overrides it per request. ## What isolation buys If the honest problem is that a marginal endpoint is starving the ones that pay the bills, isolation is the highest-value change and often the easiest to justify. A dedicated `<Executor>` used by a second connector — or simply a bounded worker pool inside the application for that call — means the failure consumes a defined slice. ```xml <Executor name="partnerPool" namePrefix="partner-exec-" maxThreads="30" minSpareThreads="5"/> <Connector executor="partnerPool" port="8081" protocol="HTTP/1.1"/> ``` This is the bulkhead idea applied at the container: you accept that the flaky endpoint will fail under stress, and you buy the guarantee that it fails alone. The trade is utilisation — reserved capacity sits idle when that endpoint is quiet — and operational complexity, because traffic now has to be routed to the right port by whatever sits in front. ## When raising maxThreads is legitimate It is not always wrong. If each request blocks briefly, the dependency has plenty of headroom, and per-thread cost is comfortable at the new number, more threads convert directly into more throughput and cost you one line of configuration instead of a rewrite. Be explicit about the two ceilings you are moving toward: the memory and scheduling cost of the threads themselves, and the concurrency the downstream will tolerate. Doubling `maxThreads` against a dependency that is already at its limit does not add throughput; it adds queueing and pushes the failure onto someone else's service. ## How to decide, and how to know you were right A defensible order of reasoning: *bound the call* (timeout, circuit breaker) → *contain the damage* (isolate the pool) → *remove the artificial ceiling* (async, with an explicit concurrency limit) → *buy more capacity* (resize, having verified the downstream can take it). The first two are cheap and almost always correct; the third is a real engineering investment that should be justified by a measured ceiling; the fourth is a dial, not a design. Validate with numbers you agreed on beforehand: saturation as the ratio of busy exec threads to `maxThreads`, tail latency of the healthy endpoints during a slow-dependency episode, and the error rate of the third-party call. If tail latency on the unrelated endpoints stops moving when the third party degrades, the isolation worked — regardless of whether the slow endpoint got faster, which was never the goal.

  • Why can converting to async make an outage worse rather than better?
    Because the saturated pool was acting as implicit backpressure. Once threads are no longer the limit, the service happily accepts far more concurrent calls than the downstream can serve, amplifying load on something already struggling. Async must come with an explicit concurrency bound and a real timeout to replace the ceiling it removed.
  • What must be true of the code path for async to actually free capacity?
    It has to be non-blocking end to end. If the async callback calls a blocking client, or a filter in the chain does not support async, the work simply blocks a different pool and the ceiling reappears elsewhere. The servlet must declare asyncSupported and every filter in its chain must too.
  • How would you demonstrate afterwards that isolating the endpoint worked?
    By measuring the healthy endpoints, not the isolated one. During a slow-dependency episode their tail latency and error rate should be flat, and the busy-thread ratio of the main pool should stay well below saturation. The isolated endpoint is still expected to degrade — the objective was to make it degrade alone.

saying these in an interview costs you the question

  • Reaches for async before checking whether the dependency is simply broken
  • Assumes async makes individual requests faster
  • Sizes maxThreads without asking what the downstream can absorb
  • Ignores that async removes the pool's implicit backpressure
  • Shares one Executor across connectors and calls that isolation

context