skip to content

Setting `spring.threads.virtual.enabled=true` on a Spring Boot 3.2+ application running Java 21 changes how the embedded Tomcat handles requests. What changes, and which setting then bounds how many requests run at once?

level: seniorimportance: should knowfreq 42%

answer

  1. one property, no code change
  2. the pool stops being the ceiling
  3. the limit becomes connection-level
  4. the queue moves to whatever is scarce
  5. the old saturation metric goes flat

basics

~20 s

Spring Boot gives embedded Tomcat a virtual-thread-per-task executor, so each request gets a fresh virtual thread instead of one from the fixed pool. server.tomcat.threads.max stops bounding concurrency; server.tomcat.max-connections and accept-count become the real admission limits.

solid answer

~50 s

With the flag on, Spring Boot configures the embedded Tomcat's protocol handler to use a virtual-thread-per-task executor, so a request no longer borrows a platform thread from a pool of 200 — it gets its own virtual thread, and blocking on I/O parks that thread cheaply instead of pinning an OS thread. The consequence people miss is that `server.tomcat.threads.max` stops being the concurrency ceiling, because there is no bounded pool any more. Admission control moves to `server.tomcat.max-connections`, default 8192, plus `server.tomcat.accept-count`. That is a very different posture: the service will now happily admit thousands of concurrent requests and push the queue down onto whatever is actually scarce — usually the database connection pool or a downstream service. The thread pool used to be an accidental bulkhead, and turning this on removes it, so you need explicit limits somewhere. Blocking, thread-per-request code is fine; ThreadLocal-heavy caching and thread-pinned monitoring are what need review.

code

properties · 3 lines
properties
spring.threads.virtual.enabled=true
server.tomcat.max-connections=2000
server.tomcat.accept-count=100

go deeper

for a junior

Know that one property switches the embedded container to a new virtual thread per request instead of a fixed pool, and that the application code stays ordinary blocking code.

for a middle

Explain that the fixed pool disappears, so the maximum-threads property no longer limits concurrency and the connection-level limits take over.

for a senior

Show that the thread pool was acting as an unplanned bulkhead, and describe what must replace it: a realistic connection cap, bounded downstream pools, and outbound timeouts, with monitoring changed before the switch.

for a principal

Decide fleet-wide where the change pays: I/O-bound request serving benefits, CPU-bound work does not, and every adopting service needs an explicit admission-control policy so overload is refused at the edge rather than absorbed.

## What the flag actually does `spring.threads.virtual.enabled=true` is a single property, available from Spring Boot 3.2 on Java 21 or later. For the web tier it makes Boot hand the embedded Tomcat's protocol handler an executor that creates a new virtual thread per task, replacing the fixed pool of platform threads. Boot applies the same treatment to other places it owns thread creation, such as scheduled tasks and asynchronous execution, but the request path is where the operational change is felt. Before: a request is handed to one of at most 200 platform threads; if it blocks on a database call, that OS thread is unusable for the duration. After: a request runs on a virtual thread; if it blocks, the virtual thread parks and its carrier is released to run something else. The programming model does not change — the code is still ordinary blocking, thread-per-request servlet code — but the cost of blocking collapses. ## The bound moves This is the interview answer. `server.tomcat.threads.max` bounded concurrency because the pool was finite. With virtual threads there is no pool to exhaust, so that property no longer caps how many requests execute at once. The remaining limits are connection-level: `server.tomcat.max-connections`, default 8192, and behind it `server.tomcat.accept-count`. The practical shift is from roughly 200 concurrent requests to potentially thousands. Whether that is good depends entirely on what the requests do. If they are mostly waiting on a fast downstream, throughput improves and latency under burst gets much better. If they contend for a scarce resource, nothing improved — the queue simply moved. A service with a 20-connection datasource that now admits 4,000 concurrent requests has 3,980 of them waiting for a connection, holding memory and half-built object graphs, and probably timing out at the caller after work was already spent on them. ## The bulkhead you did not know you had The old thread pool was doing two jobs: multiplexing work onto OS threads, and *limiting admission*. Only the first job is obsolete. Turning on virtual threads removes the second silently, so it has to be replaced deliberately: - Lower `server.tomcat.max-connections` to a number the service can genuinely serve, so overload is refused at the edge rather than queued inside. - Bound each scarce downstream explicitly — connection-pool size with a short acquisition timeout, a semaphore or concurrency limit around a fragile dependency. - Keep read timeouts on outbound calls. Cheap threads make it tempting to skip them; a slow dependency now accumulates thousands of parked requests instead of 200, which is a larger pile of pending work, not a smaller problem. ## What to re-check in the application - **ThreadLocal caching.** Anything that caches an expensive object per thread on the assumption that there are 200 long-lived threads now allocates one per request. Correctness usually holds — request-scoped context propagation still works because each request has its own thread — but memory and hit rates change. - **Pooled resources tied to threads.** Any code that assumed thread identity is stable across requests, or that a bounded pool was throttling it, deserves a look. - **Pinning.** On Java 21, blocking inside a `synchronized` block pinned the virtual thread to its carrier, which could starve the small carrier pool under load; JDK 24 removed that limitation. If you run on 21 with heavily synchronized blocking code, this is worth measuring rather than assuming. - **Monitoring.** The busy-request-threads metric that used to be the saturation signal goes flat and stops telling you anything. Replace it with in-flight request count, connection count, and downstream pool wait time before you switch, not after. ## Deciding whether to turn it on The honest test is what the service does with its time. A request-serving API dominated by waiting on I/O is the ideal case. A CPU-bound service gains nothing: the work still needs cores, and virtual threads do not create any. A service already running comfortably at 200 concurrency with headroom gains little beyond burst tolerance. Roll it out on one service, watch tail latency and downstream saturation, and confirm the new admission limits are the ones you chose rather than the 8192 default nobody looked at.

  • Which kinds of Spring Boot services gain little or nothing from enabling virtual threads?
    CPU-bound ones. Virtual threads make blocking cheap; they do not create cores, so a service that spends its time computing sees no throughput gain and just more concurrency contending for the same CPU. Services already comfortable at 200 concurrency with idle downstreams gain mostly burst tolerance. The wins are concentrated in I/O-bound request serving with slow or bursty dependencies.
  • After enabling virtual threads, which saturation signal replaces the busy-thread metric?
    In-flight request count and connection count at the container, plus wait time to acquire whatever downstream resource is scarce — typically the database connection pool. The busy-thread gauge flattens because there is no bounded pool to saturate, so a dashboard built on it silently stops warning. Put the replacement metrics in place before the switch, not after the first incident.
  • Why can leaving server.tomcat.max-connections at its default be risky once virtual threads are enabled?
    Because it becomes the de-facto admission limit, and 8192 is far more concurrent work than most services can complete. Requests are admitted, queue on a small connection pool, consume memory, and time out at the caller after the service has already spent effort on them. Lower it to a number the service can genuinely serve so overload is refused at the edge.

saying these in an interview costs you the question

  • Says server.tomcat.threads.max still caps concurrency with virtual threads on
  • Claims virtual threads improve CPU-bound throughput
  • Believes blocking code must be rewritten to reactive style first
  • Assumes downstream pools scale automatically with request concurrency
  • Keeps the busy-thread dashboard as the saturation signal

context