A service intermittently fails with 'connection pool exhausted' under load but recovers when idle. How would you reason about and prevent this resource-leak class of bug?
answer
- borrow-but-never-return under load
- fails on victim, not culprit (non-local)
- active count climbs, never returns to baseline
- try-with-resources guarantees return on all paths
- leakDetectionThreshold + pool metrics/alerts
basics
~20 sSome code path borrows a connection from the pool and never returns it (no close on every path), so under load all connections get stuck and new requests block until they time out. Fix it by closing every borrowed resource with try-with-resources so it always returns to the pool.
solid answer
~50 sPool exhaustion under load with recovery when idle is the signature of a leak: a connection or statement is acquired but not always released, so each leaking request permanently removes one connection from the pool. Idle periods hide it because demand is below the leak rate; load exposes it. To reason about it: confirm acquisitions equal releases via pool metrics (active vs idle count climbing and never returning), find code paths that get a Connection/Statement/ResultSet without a guaranteed close — especially error paths and early returns — and check whether close lives in a finally that itself can be skipped. The fix is to wrap every pooled resource in try-with-resources so it is returned even on exception, close in last-acquired-first-released order, and never hold a connection across a slow external call. Add leak-detection (e.g. HikariCP leakDetectionThreshold) and pool metrics/alerts so the next regression is caught early rather than in production.
code
java · 11 lines// Leak-proof: every pooled resource returned on all paths.
try (Connection c = dataSource.getConnection();
PreparedStatement ps = c.prepareStatement("select * from users where id=?")) {
ps.setLong(1, id);
try (ResultSet rs = ps.executeQuery()) {
return map(rs);
}
} // c & ps returned to pool even if executeQuery or map throws
catch (SQLException e) {
throw new RepositoryException("load user " + id, e); // chain the cause
}go deeper
Understands that a borrowed connection must be closed/returned and that try-with-resources guarantees it, so the pool does not run out.
Connects the under-load/idle-recovery symptom to a leak, audits close-on-all-paths, and applies try-with-resources to fix it.
Drives diagnosis with pool metrics and leak detection, distinguishes true leaks from held-while-blocking, reasons about acquire-late/release-early and timeouts, and explains the non-local failure.
Establishes systemic safeguards: pool sizing models, leak-detection and saturation alerting across services, code standards mandating try-with-resources, and capacity/back-pressure design so bounded resources degrade gracefully.
## What a connection pool is Opening a database connection is expensive, so applications keep a fixed-size **pool** of pre-opened connections. Code **borrows** one, uses it, and **returns** it (by calling `close()` on the pooled wrapper, which does not really close the socket — it hands the connection back). The pool has a maximum size; when all connections are borrowed, the next request **waits** up to a timeout, then fails with an error like "pool exhausted" / "timeout acquiring connection". ## Why the symptom points to a leak A **leak** here means: borrowed but never returned. Each leaking request permanently consumes one slot. The tell-tale pattern: - **Under load**: many requests, several of them leak, the pool drains, new requests block then fail. - **When idle**: traffic is below the leak rate, the few leaked connections are tolerable or masked, and acquire timeouts stop happening — so it "recovers." Crucially it does not actually heal; it just stops being visible. This non-locality (failure appears later and elsewhere than the bug) is what makes leaks hard. The stack trace at failure points at the *victim* request that could not get a connection, **not** at the *culprit* that leaked one. ## How the leak happens in code The usual culprits: ```java Connection c = pool.getConnection(); PreparedStatement ps = c.prepareStatement(sql); ResultSet rs = ps.executeQuery(); // if this throws, c & ps never close process(rs); rs.close(); ps.close(); c.close(); // only runs on the happy path ``` Any exception, early `return`, or `break` before the manual `close()` leaks. A `finally` helps but is easy to get partially wrong (e.g. closing `rs` but the close throws and skips closing `c`). ## Diagnosis approach 1. **Metrics first.** Watch pool gauges: active connections, idle connections, pending acquires. A leak shows active climbing monotonically and never returning to baseline. 2. **Leak detection.** Pools like HikariCP have a `leakDetectionThreshold`: if a connection is held longer than N ms, it logs a stack trace **of where it was borrowed** — pointing straight at the culprit. 3. **Audit acquire/release symmetry.** Every `getConnection`/`prepareStatement`/`executeQuery` must have a guaranteed matching `close` on **all** paths (success, exception, early return). 4. **Check held-while-blocking.** Holding a connection across a slow HTTP call or lock multiplies effective demand and can exhaust the pool even without a true leak. ## The fix Use **try-with-resources** so every borrowed resource is returned on every path: ```java try (Connection c = pool.getConnection(); PreparedStatement ps = c.prepareStatement(sql); ResultSet rs = ps.executeQuery()) { process(rs); } // rs, ps, c closed in reverse order, even on exception ``` This guarantees return-to-pool, closes in last-acquired-first-released order, and turns a close failure into a suppressed exception rather than skipping the connection close. Additionally: - **Right-size and bound** the pool and the acquire timeout (fail fast rather than pile up). - **Don't hold connections across slow calls**; acquire late, release early. - **Add alerts** on pending-acquire count and on leak-detection logs so the next regression is caught in staging. ## Why this is the canonical resource-leak lesson File handles, sockets, threads, and locks all behave the same way: acquired-but-not-released under load exhausts a bounded supply, and the failure surfaces far from its cause. try-with-resources plus metrics/leak-detection is the general cure: make release automatic, and make leaks observable.
- Why does the stack trace at the point of failure usually not point at the leaking code?The failure happens in a later request that cannot acquire a connection (the victim); the request that leaked one already completed. That non-locality is why you need pool metrics and leak-detection (which captures the borrow-site trace) to find the culprit.
- Besides leaks, what else can exhaust a connection pool under load?Holding connections across slow operations (external HTTP calls, long transactions, lock waits), an undersized pool for the concurrency level, or a thundering herd of retries. These raise effective demand without an actual leak.
A pool of shared rental bikes: if some riders never dock their bike, the rack empties under busy hours and others wait. At night demand drops so the few missing bikes go unnoticed — the shortage 'recovers' without anyone returning a bike. try-with-resources is an auto-docking bike that returns itself when you let go.
saying these in an interview costs you the question
- Blaming the victim stack trace / the request that timed out instead of the leaking path.
- Just increasing pool size to 'fix' it — only delays exhaustion if there is a real leak.
- Closing only some resources in a manual finally and assuming the connection is returned.
- Holding the connection while doing slow I/O and calling it a pool-size problem.