skip to content

Explain `server.tomcat.threads.max` and `server.tomcat.accept-count`. How do they interact under load?

level: middleimportance: must knowfreq 70%

answer

  1. threads.max=200 → concurrency ceiling
  2. accept-count=100 → TCP backlog when threads busy
  3. backlog full → OS refuses connection
  4. min-spare=10 warm threads
  5. capped by somaxconn; don't outgrow DB pool

basics

~10 s

threads.max (default 200) is the max worker threads that process requests. accept-count (default 100) is the OS backlog queue of connections waiting when all threads are busy; beyond it, new connections are refused.

solid answer

~40 s

Tomcat processes each request on a worker thread from a pool capped by `server.tomcat.threads.max` (default 200). When every worker is busy, newly accepted connections wait in a backlog queue whose length is `server.tomcat.accept-count` (default 100) — this maps to the TCP listen backlog. If that queue is also full, the OS refuses further connections (client sees connection refused / reset). So the two work in sequence: threads do the work; accept-count absorbs a short burst beyond thread capacity. Raising `threads.max` increases concurrency but costs memory (each thread ~0.5–1 MB stack) and can overload downstream resources like a DB pool. A common tuning mistake is bumping threads.max huge without matching the DB connection pool, just moving the bottleneck. `threads.min-spare` (default 10) sets how many idle threads stay warm.

code

yaml · 9 lines
yaml
server:
  tomcat:
    threads:
      max: 200        # default; worker threads that run request handlers
      min-spare: 10   # default; idle threads kept warm
    accept-count: 100 # default; TCP backlog once all threads are busy
    max-connections: 8192 # connections accepted before backlog kicks in
# Rule of thumb: keep threads.max in line with downstream capacity
# (e.g. HikariCP maximum-pool-size) so you don't just move the bottleneck.

go deeper

for a junior

Knows threads.max is the worker pool (200) and accept-count is a queue (100).

for a middle

Explains the busy→backlog→refuse sequence and the memory cost of threads.

for a senior

Ties threads.max to downstream pool sizing and discusses fail-fast backpressure.

for a principal

Reasons about capacity planning end-to-end, somaxconn caps, async thread release, and load-shedding strategy.

## The request-processing pipeline in embedded Tomcat Tomcat (default connector: NIO) handles a request in stages, and these properties tune different stages. ### 1. Worker threads — `server.tomcat.threads.max` - The **thread pool** that actually executes your controller/servlet code. One thread is occupied for the full duration of a synchronous request. - **Default: 200.** - `server.tomcat.threads.min-spare` (default 10) = idle threads kept warm to absorb bursts without paying thread-creation cost. - **Concurrency ceiling:** with 200 threads, at most ~200 requests execute *simultaneously*. More arriving requests must wait. - **Cost:** each thread has a stack (~512 KB–1 MB) and context-switch overhead. Setting threads.max to thousands wastes memory and rarely helps for CPU-bound work; it can help for I/O-bound work that blocks on slow downstreams — but only up to what those downstreams (DB pool, remote API) can absorb. ### 2. The backlog — `server.tomcat.accept-count` - Once all worker threads (and the connection limit, see below) are saturated, the acceptor stops handing off new connections. Pending connections queue in the **OS-level TCP accept/listen backlog**. - `accept-count` sets that backlog length. **Default: 100.** - When the backlog is **full**, the OS **rejects** new TCP connections — the client gets `connection refused` (or a timeout/reset depending on OS). This is a hard shed of load, not queuing forever. ### How they interact under load (mental model) 1. Request arrives → connection accepted (subject to `max-connections`). 2. Assigned to a free worker thread if one exists (up to `threads.max`). 3. If all threads busy, the connection waits in the `accept-count` backlog. 4. If the backlog is full, the OS refuses the connection. So: **threads.max = concurrency**, **accept-count = short-burst buffer**, and beyond both you shed load fast (fail fast) rather than piling latency. ## Gotchas - **Moving the bottleneck:** raising `threads.max` to 500 while the HikariCP pool is 10 just means 490 threads block waiting for a DB connection. Tune the whole chain. - **accept-count is capped by the OS** (`net.core.somaxconn` on Linux); requesting 1000 may be silently clamped. - **Fail-fast vs. queue-forever:** a small accept-count makes the server reject early under overload, which is often *desirable* (backpressure) versus building an unbounded latency queue. - Property renamed over Boot versions: older `server.tomcat.max-threads` → now `server.tomcat.threads.max`; `server.tomcat.min-spare-threads` → `server.tomcat.threads.min-spare`. - Async requests (`DeferredResult`, `WebFlux`, `CompletableFuture` return) release the worker thread while waiting, so thread count is less of a ceiling for those. ## When to tune High-concurrency I/O-bound services may raise threads.max; latency-sensitive services often keep accept-count modest to shed load fast. Always load-test rather than guess.

  • You raise threads.max from 200 to 800 but throughput doesn't improve and latency gets worse. Why?
    The bottleneck is downstream (e.g. a DB pool of 20 or a CPU-bound path). Extra threads just block waiting or thrash the CPU with context switches and memory pressure — you moved, not removed, the constraint.
  • What happens when both the threads are saturated and the accept-count backlog is full?
    The OS stops accepting new TCP connections, so clients get connection refused / reset. This is deliberate load shedding (backpressure) rather than unbounded queuing.
  • Why might a small accept-count be a good thing?
    It makes the server fail fast under overload instead of building a huge latency queue, giving clients/load-balancers a fast signal to retry elsewhere.

saying these in an interview costs you the question

  • Claiming accept-count is a second thread pool
  • Saying accept-count queues requests indefinitely (it's a bounded OS backlog; full → refused)
  • Thinking bigger threads.max always increases throughput
  • Confusing accept-count (backlog) with max-connections (accepted-connection ceiling)

context