You see production requests hanging ~30s then failing with 'Connection is not available'. You suspect nested REQUIRES_NEW. How do you confirm it and what are your fix options?
answer
- Thread dump parked in HikariPool.getConnection
- active=max, pending rising, leak-detection-threshold
- Fix = remove the double-hold, not raise pool
- after-commit event / outbox / async / dedicated pool
- Bigger pool = stopgap only
basics
~20 sConfirm via thread dumps (threads parked in HikariPool.getConnection) and Hikari metrics (active pinned at max, pending rising). Fix by removing the nesting: run the inner work in a separate transaction before/after the outer, publish it as an event/outbox, offload it, or as a stopgap raise the pool size.
solid answer
~50 sFirst confirm: take a thread dump and look for many threads blocked in HikariPool.getConnection / ConcurrentBag.borrow while inside an outer @Transactional; check hikaricp.connections.active pinned at max with pending climbing, and correlate the 30s timeouts with connectionTimeout. Enable leakDetectionThreshold to spot the long-held outer connection. Then fix by eliminating the simultaneous two-connection hold, not by masking it. Best options: restructure so the REQUIRES_NEW work happens outside the outer transaction (do it before it starts, or after it commits); replace it with a domain event/transactional-outbox consumed after commit; or move it to an async worker on a separate DataSource/pool. If you truly need immediate independent persistence, give that work its own dedicated small pool so it never competes with the outer pool. Raising maximumPoolSize only widens the window — use it as a stopgap, not the cure.
code
java · 25 lines// Fix via after-commit event: no two-connection overlap.
@Service
public class OrderService {
private final ApplicationEventPublisher events;
// ...
@Transactional // single connection held
public void placeOrder(Order order) {
orderRepo.save(order);
events.publishEvent(new OrderPlaced(order.getId())); // no DB call here
}
}
@Component
class AuditListener {
// Runs AFTER the outer tx commits and its connection is released,
// then this handler opens its own connection -> never held simultaneously.
@Async
@TransactionalEventListener(phase = TransactionPhase.AFTER_COMMIT)
@Transactional(propagation = Propagation.REQUIRES_NEW)
public void onOrderPlaced(OrderPlaced e) {
auditRepo.save(new AuditEntry("ORDER_PLACED", e.orderId()));
}
}
// Note: with @Async the listener runs on another thread after commit, so the
// outer connection is already back in the pool -> no circular wait possible.go deeper
Can name the symptom (timeouts) but not confirm or systematically fix it.
Diagnoses via metrics/thread dump and knows raising the pool is a stopgap.
Offers a ranked set of structural fixes (after-commit event, outbox, async, dedicated pool) and understands each trade-off.
Ties the fix choice to the real requirement (immediate vs independent vs in-scope) and the Modulith event/outbox architecture.
### Step 1 — Confirm the diagnosis - **Thread dump** (`jstack`, or the actuator `/threaddump`): the smoking gun is a cluster of worker threads `WAITING`/`TIMED_WAITING` in `com.zaxxer.hikari.pool.HikariPool.getConnection` → `ConcurrentBag.borrow`, each with a stack frame showing they are already inside a `@Transactional` outer method (a `TransactionInterceptor` frame). That proves they hold one connection and are waiting for a second. - **HikariCP metrics** (Micrometer): `hikaricp.connections.active` flatlined at `maximumPoolSize`, `hikaricp.connections.pending` rising, `hikaricp.connections.acquire` latency exploding, `hikaricp.connections.idle` = 0. - **Leak detection**: set `spring.datasource.hikari.leak-detection-threshold=20000`; Hikari logs a stack trace for any connection held longer than the threshold, pinpointing the long-lived outer connection. - **Logs**: repeated `SQLTransientConnectionException: ... request timed out after 30000ms`, roughly 30s cadence. - **Correlate with traffic**: it appears only above ~`poolSize/2` concurrent requests through the nested path. ### Step 2 — Fix by removing the double-hold The root cause is holding two pool connections under one open transaction. Ranked options: 1. **Move the independent work out of the outer transaction.** If the REQUIRES_NEW step doesn't need to run mid-transaction, do it **before** the outer transaction opens or **after** it commits. Then only one connection is ever held at a time. Use `TransactionSynchronizationManager.registerSynchronization(...)` `afterCommit`, or Spring's `@TransactionalEventListener(phase = AFTER_COMMIT)`. 2. **Event / transactional-outbox pattern.** Publish a domain event inside the outer transaction; a listener (after commit) or an outbox poller performs the side write on a fresh connection with no overlap. This is the idiomatic Spring Modulith approach and removes the nesting entirely. 3. **Asynchronous offload.** Hand the independent write to an `@Async` worker / message queue. It runs on its own thread and borrows its connection only after the outer has released — no simultaneous hold. (Ensure ordering/consistency requirements allow deferral.) 4. **Dedicated DataSource/pool for the inner work.** If the write genuinely must be immediate, independent, and inside the outer scope (rare), give it its **own** `DataSource` bound via a second `PlatformTransactionManager`. Then the inner never competes with the outer pool, so no circular wait. Costs a second pool and more DB connections total. 5. **Stopgap: raise `maximumPoolSize`.** Doubling the pool doubles the safe concurrency (`poolSize/2`), buying headroom, but the deadlock is still latent and returns at higher load. It also pushes more load onto the database. Use only to stabilize while you implement 1–4. ### What NOT to do - Don't just bump the web/thread pool — more concurrent outer holders make it worse. - Don't shorten `connectionTimeout` and call it fixed — you convert hangs into fast failures but still drop requests. - Don't 'fix' it by switching the inner to REQUIRED unless you actually want it to share the outer transaction and lose independence — that changes semantics (an inner failure now taints the outer). It does remove the deadlock, but only because it stops opening a second connection; make that a deliberate decision. ### Trade-off summary The cleanest fixes (1–3) preserve one-connection-per-thread and are the right long-term answer. Option 4 keeps immediacy at the cost of infrastructure. Option 5 is triage. Match the fix to whether the inner write really must be immediate, independent, and inside the outer scope — usually it does not need all three.
- Why is switching the inner method to REQUIRED a semantic change, not just a fix?REQUIRED makes the inner join the outer transaction on the same connection, so no second connection is borrowed (deadlock gone). But now the inner work is no longer independent: an inner failure marks the whole transaction rollback-only, and the inner write commits only when the outer commits, losing the 'must-survive-independently' guarantee REQUIRES_NEW gave.
- Which HikariCP setting helps you find who is holding a connection too long?leakDetectionThreshold (spring.datasource.hikari.leak-detection-threshold). When a connection is held longer than that many ms, Hikari logs a stack trace showing where it was acquired, exposing the long-lived outer connection.
saying these in an interview costs you the question
- Treating a bigger pool as the real fix
- Only lowering connectionTimeout to hide the hang
- Switching inner to REQUIRED without acknowledging the lost independence
- Ignoring that the inner work often doesn't need to run mid-transaction at all