Why can nesting REQUIRES_NEW inside a REQUIRED transaction exhaust a connection pool and deadlock under load? Walk through the mechanics.
answer
- Two connections per thread
- N threads drain N-size pool on outer tx
- All then wait for connection #2 = circular wait
- connectionTimeout 30s then SQLTransientConnectionException
- Safe threads = poolSize/2
basics
~20 sEach thread holds two connections at once (the suspended outer plus the new inner). With a pool of N, once N threads each grab their first connection, every one needs a second that no one can release, so they all block and the pool deadlocks.
solid answer
~50 sREQUIRES_NEW inside REQUIRED makes one thread hold two connections simultaneously: the outer transaction's connection stays checked out (suspended, not released) while the inner borrows a second from the same pool. Say HikariCP has the default 10 connections and 10 concurrent requests arrive. Each thread enters the outer @Transactional and grabs one connection, so all 10 are checked out. Then each thread hits the REQUIRES_NEW call and asks for a second connection. The pool is empty. None of the threads will release their first connection because they are blocked inside the outer transaction waiting for the second. Circular wait: every thread holds one resource and waits for another that only a peer can free. Hikari blocks up to connectionTimeout (default 30s) and then throws SQLTransientConnectionException ('Connection is not available, request timed out'). Effectively a self-inflicted deadlock and pool exhaustion.
code
java · 28 lines// application.yml equivalent: spring.datasource.hikari.maximum-pool-size = 10
@Service
public class ReportService {
private final CheckpointService checkpointService;
// ...constructor...
@Transactional // REQUIRED -> borrows connection #1, holds it
public void generate(long batchId) {
for (Row r : repo.streamRows(batchId)) { // long-running outer tx
process(r);
// Inner independent tx to persist a checkpoint that must survive
// even if the batch later fails:
checkpointService.mark(r.getId()); // needs connection #2
}
}
}
@Service
class CheckpointService {
@Transactional(propagation = Propagation.REQUIRES_NEW) // second connection
public void mark(long rowId) {
checkpointRepo.save(new Checkpoint(rowId));
}
}
// Under >=6 concurrent generate() calls (pool=10), 10 threads pin 10 connections
// on the outer tx, then all block requesting the 11th -> 30s timeouts, pool dead.go deeper
Should grasp that two connections per thread + a fixed pool is the setup for exhaustion.
Must walk the circular-wait mechanics and name the poolSize/2 threshold and the timeout exception.
Connects it to Coffman conditions and to HikariCP metrics/thread-dump symptoms.
Generalizes to any 'two-resources-under-one-open-tx' anti-pattern and its capacity math.
This is a classic **resource-ordering deadlock** expressed through a connection pool. ### Setup - A **connection pool** (e.g. HikariCP, the Spring Boot default) hands out a fixed number of physical JDBC `Connection`s. `maximumPoolSize` defaults to **10**. When all are checked out, further `getConnection()` calls **block** up to `connectionTimeout` (default **30000 ms**), then throw. - An outer method is `@Transactional` (**REQUIRED**) — on entry it borrows **connection #1** and keeps it for the whole method. - Somewhere inside, it calls a method annotated `@Transactional(propagation = REQUIRES_NEW)`. Spring **suspends** the outer transaction (connection #1 is held aside, still checked out) and borrows **connection #2** for the inner transaction. ### The deadlock, step by step Assume `maximumPoolSize = 10` and 10 requests run concurrently (each on its own thread from the web container's thread pool, e.g. Tomcat's 200 threads — note the thread pool is far larger than the connection pool): 1. All 10 threads enter the outer transaction and each borrows one connection. **Pool is now fully drained (0 free).** 2. Each thread proceeds to the REQUIRES_NEW call and requests a **second** connection. 3. There are none. Every thread blocks in `HikariPool.getConnection()`. 4. A thread would only free its connection #1 by **finishing the outer transaction**, but it cannot finish because it is stuck waiting for connection #2. 5. **Circular wait**: all threads hold one and need one; no thread can make progress. This satisfies the classic Coffman deadlock conditions (mutual exclusion, hold-and-wait, no preemption, circular wait). 6. After `connectionTimeout` (30s) each blocked acquire throws `SQLTransientConnectionException: Connection is not available, request timed out after 30000ms`. Requests fail en masse; latency spikes to ~30s; the app appears hung. ### Why it is load-dependent (and hard to catch in tests) With fewer than `poolSize/2` concurrent nested requests, there are always spare connections, so it works fine in dev and low-traffic. The failure only appears once concurrent nested-REQUIRES_NEW requests approach **half the pool size**. That is why it slips through unit/integration tests and surfaces as a production incident under peak traffic. Precisely: the pool supports at most `floor(maximumPoolSize / 2)` such threads safely; the (poolSize/2)+1-th thread can trigger the stall. ### Symptoms in monitoring - HikariCP metrics: `hikaricp.connections.active` pinned at max, `hikaricp.connections.pending` climbing, `hikaricp.connections.usage`/`acquire` timing spiking. - Thread dumps: many threads parked in `HikariPool.getConnection` / `ConcurrentBag.borrow`. - Logs: bursts of `SQLTransientConnectionException` timeouts, often 30s apart. ### The subtlety people miss It is not that REQUIRES_NEW is 'slow' — it is that it **doubles the per-thread connection demand while the first connection is un-releasable**. Any pattern that holds two pool connections at once under an open transaction (nested REQUIRES_NEW, or opening a second `DataSource`-backed resource mid-transaction) has the same failure mode.
- At what concurrency does a pool of 10 start to deadlock with this pattern?It can stall once concurrent nested requests exceed floor(10/2) = 5. The 6th thread onward may find no free connection for its inner transaction while all outer connections are held, producing the circular wait.
- What exception and roughly what latency would you see when it triggers?HikariCP throws SQLTransientConnectionException ('Connection is not available, request timed out after 30000ms') after the default 30s connectionTimeout, so requests hang for ~30 seconds before failing.
saying these in an interview costs you the question
- Blaming slow queries rather than the two-connection hold
- Thinking a bigger web/thread pool helps (it makes it worse by allowing more concurrent outer holders)
- Believing each thread only ever uses one connection
- Assuming it would show up in low-traffic tests