A Spring Boot service on embedded Tomcat stops answering new requests under load while its CPU stays near idle. Which `server.tomcat.*` properties bound how many requests it can handle at once, and what happens to a request that arrives past each bound?
answer
- three limits, not one
- accepted is not the same as processed
- idle CPU means blocked, not busy
- the narrowest downstream resource sets the ceiling
- past the backlog the kernel refuses
basics
~20 sThree properties bound it: server.tomcat.max-connections (default 8192) caps accepted connections, server.tomcat.threads.max (default 200) caps requests processed at once, and server.tomcat.accept-count (default 100) sizes the OS backlog beyond that. Idle CPU means the 200 threads are blocked, not busy.
solid answer
~50 sConcurrency is bounded in three places. `server.tomcat.max-connections`, default 8192, is how many connections the connector will hold at once. `server.tomcat.threads.max`, default 200, is how many requests can be *processed* simultaneously — a connection above that is held but its request waits for a thread. `server.tomcat.accept-count`, default 100, is the OS accept queue depth once max-connections is reached; past that the kernel refuses the connection outright, which clients see as a connection reset or refused rather than a slow response. Idle CPU with no throughput is the diagnostic: the 200 threads are not computing, they are all blocked — typically on a downstream HTTP call or database connection with no timeout, or on a saturated connection pool. Raising `threads.max` then buys nothing and usually makes it worse, because more threads pile onto the same starved resource. The fix is a timeout on the blocking call and the right pool size behind it.
code
properties · 4 linesserver.tomcat.threads.max=200
server.tomcat.threads.min-spare=10
server.tomcat.max-connections=8192
server.tomcat.accept-count=100go deeper
Know that the embedded container processes requests on a bounded pool of threads whose default maximum is 200, and that the setting is a property rather than a code change.
Distinguish the three bounds — connections held, requests processed, kernel backlog — and describe exactly what a client experiences when each is exceeded.
Diagnose from the symptom: pinned busy threads with idle CPU means blocked I/O, and the correct response is timeouts and downstream pool sizing rather than a larger thread pool.
Own an admission-control policy across services: decide where load is shed versus queued, ensure thread pools are derived from the narrowest downstream resource, and make the saturation signal visible on every dashboard.
## The three bounds, in the order a request meets them A request arriving at an embedded Tomcat passes through three separate limits, and knowing which one it hit tells you what the client saw. **`server.tomcat.max-connections` (default 8192)** is the number of connections the connector will keep open and poll at one time. Tomcat's connector accepts connections on one thread and hands them to a poller; the connection existing does not mean a request is being processed. **`server.tomcat.threads.max` (default 200)** is the size of the pool of request-processing threads. This is the real concurrency ceiling for a blocking, thread-per-request application: at most 200 requests are executing controller code at any instant. Connection number 201 is accepted and its request simply waits for a thread to free up. `server.tomcat.threads.min-spare` (default 10) is how many are kept alive when idle. **`server.tomcat.accept-count` (default 100)** is the depth of the operating system's accept queue for the listening socket. It only matters once `max-connections` is reached and Tomcat stops accepting. New TCP connections then land in the kernel queue; when that queue is also full, the kernel refuses the connection. The client does not get a slow response or a 503 — it gets a refused connection or a reset, which is why this failure mode often shows up in a load balancer's logs as a connect error rather than in the application's logs at all. ```properties server.tomcat.threads.max=200 server.tomcat.threads.min-spare=10 server.tomcat.max-connections=8192 server.tomcat.accept-count=100 ``` ## Reading the symptom: idle CPU, no throughput A saturated thread pool with high CPU means the service is genuinely doing work and needs more capacity. A saturated thread pool with idle CPU means the opposite: every one of those 200 threads is parked in a blocking call and consuming nothing. The usual causes are a downstream HTTP call whose client has no read timeout, a database call waiting on a connection pool that is far smaller than 200, a lock held across I/O, or an external dependency that has become slow rather than failed. Slow is worse than down here — a dependency that returns errors in 5 ms frees the thread immediately, while one that takes 60 seconds to answer holds it for 60 seconds. This is also why the reflex to raise `threads.max` to 800 is usually wrong. If the constraint is a 20-connection database pool, 800 threads just means 780 of them queue on the pool instead of 180, memory per thread stack multiplies, and the latency distribution gets worse while throughput does not move. Worse, a bigger pool lets the service accept far more work than it can complete, so requests time out at the caller after the service has already spent effort on them. ## Which knob is the right one - **Timeouts first.** Every outbound call needs a connect and a read timeout shorter than the caller's patience. A bounded blocking time turns "the pool is gone forever" into "the pool churns and sheds load". - **Match the downstream pool.** Threads that will contend for a 20-connection datasource beyond a certain count do nothing but queue. Sizing should be derived from the narrowest downstream resource, not from a round number. - **Shed rather than queue.** A shallower `accept-count` fails fast at the edge, which lets a load balancer route around the instance sooner; a deep queue hides the problem and converts it into latency for everyone. - **Then raise threads.** Increase `threads.max` when threads are genuinely busy — CPU high, downstreams healthy — and the box has headroom. ## Property naming and versions The modern spellings are `server.tomcat.threads.max` and `server.tomcat.threads.min-spare`; older material and older configuration files use `server.tomcat.max-threads` and `server.tomcat.min-spare-threads`, which were renamed in Spring Boot 2.3. A stale property name is silently ignored — the application starts, the setting does nothing, and the default quietly applies. When someone insists a tuning value is in place, confirm it against the running configuration rather than the file. ## What to watch The useful signals are the count of busy request threads against the maximum, the queue or wait time before a request starts executing, and the connection count. When busy threads sit pinned at the maximum while CPU is flat, you have your answer before opening a profiler; a thread dump then shows exactly which downstream call all of them are parked in.
- Why is a very deep accept queue often worse than a shallow one for a service behind a load balancer?A deep queue absorbs far more work than the service can complete, so requests sit waiting and then time out at the caller after the service has already spent effort on them. A shallow queue refuses fast, the load balancer sees connection errors immediately and routes elsewhere, and recovery is quicker. Deep queues convert an overload into latency for everyone instead of failing a subset early.
- How would you confirm that the threads are blocked on a downstream call rather than genuinely busy?Take a thread dump and look at the request-processing threads: if most are parked in the same socket read or waiting on a connection pool, that is the answer. Corroborate with the busy-thread metric pinned at maximum, flat CPU, and the latency of that specific dependency. Genuinely busy threads show varied stacks and CPU that tracks throughput.
- A team sets server.tomcat.max-threads=400 in a recent Spring Boot application and nothing changes. Why?That property name was renamed to `server.tomcat.threads.max` in Spring Boot 2.3. An unrecognised property is not an error — it is simply ignored, so the default of 200 stays in force while the configuration file looks tuned. Always verify against the running configuration rather than the file.
saying these in an interview costs you the question
- Says raising the max thread count always increases throughput
- Treats accepted connections and in-flight requests as the same number
- Believes an overloaded Tomcat returns 503 rather than refusing connections
- Sizes the thread pool without considering the database connection pool
- Assumes idle CPU proves the service is under-loaded