skip to content

You add REQUIRES_NEW to an audit method called deep inside high-throughput request transactions. What production risks must you evaluate, and how would you mitigate them?

level: principalimportance: should knowfreq 34%

answer

  1. Two connections per request stack -> pool deadlock at saturation
  2. poolSize >= concurrency x (depth+1)
  3. Inner tx tiny; own table; no remote I/O
  4. Async @TransactionalEventListener / outbox alternative
  5. Inner blocks on outer's locks -> DB deadlock

basics

~20 s

Each request now holds two connections at once, so the pool can exhaust or deadlock under load. Also the inner tx can't see the outer's uncommitted rows/locks and may block. Mitigate by sizing the pool for nesting depth, keeping the inner tx tiny, or moving audit to an async/outbox path.

solid answer

~50 s

REQUIRES_NEW deep in a request means every in-flight request suspends its outer transaction while holding that connection and grabs a second connection for the inner tx. Peak connection demand doubles for that call stack, so a pool sized for one-per-request can deadlock: all connections held by outer transactions, all waiting for inner connections that never free. I'd size the pool above the maximum simultaneous REQUIRES_NEW nesting and cap request concurrency accordingly, keep the inner transaction extremely short (single insert, no remote calls) to return its connection fast, and consider whether the audit really needs synchronous independence — an async listener or transactional-outbox pattern removes the second-connection-in-line entirely. I'd also watch lock visibility: the inner tx on a separate connection can't see the outer's uncommitted writes and can block on rows the outer locked, risking DB deadlock. Load-test with realistic concurrency and monitor pool wait time / timeout counts.

code

java · 19 lines
java
// Alternative that avoids holding a 2nd connection in-line:
// write the audit AFTER the business tx settles, on a listener thread.
@Component
class AuditListener {
    private final AuditRepository repo;
    AuditListener(AuditRepository repo) { this.repo = repo; }

    // Fires after the outer tx completes; its own tx, own connection,
    // but NOT nested inside the still-open request transaction.
    @Async
    @TransactionalEventListener(phase = TransactionPhase.AFTER_COMPLETION)
    @Transactional(propagation = Propagation.REQUIRES_NEW)
    public void onBusinessEvent(BusinessAudited e) {
        repo.save(new AuditRow(e.action(), e.userId(), e.committed()));
    }
}

// Business code just publishes — no second connection held during the request:
// applicationEventPublisher.publishEvent(new BusinessAudited(...));

go deeper

for a junior

Not expected to evaluate production pool/lock risks.

for a middle

Should at least flag doubled connection usage and keeping the inner tx short.

for a senior

Should quantify pool sizing and suggest async/outbox alternatives.

for a principal

Should present a full decision framework: pool math, lock/visibility deadlock, failure semantics, observability, and when NOT to use REQUIRES_NEW.

## The core risk: doubled connection demand + pool deadlock With REQUIRES_NEW nested inside an outer transaction, the outer connection is **suspended but still checked out** while a **second** connection is acquired for the inner tx. So a single request stack can hold **two** connections simultaneously. **Deadlock scenario:** pool size = N. N concurrent requests each begin their outer transaction and hold one connection. Each then reaches the REQUIRES_NEW audit call and requests a second connection — but the pool is empty, and no request will release its outer connection until *after* its inner call returns. Every thread blocks forever (until `connectionTimeout` fires, turning it into a storm of failures). This is a classic, well-known Hikari/pool exhaustion pattern. ### Mitigations for the pool 1. **Size the pool for peak nesting.** Effective need ≈ (max simultaneous requests in this path) × (max REQUIRES_NEW depth + 1). Or cap request concurrency (thread pool / semaphore) below `poolSize / (depth+1)`. 2. **Keep the inner transaction minimal.** A single `INSERT` with no network I/O returns its connection in milliseconds, shrinking the window where two are held. 3. **Reconsider synchronicity.** If the audit doesn't need to be written *before the request continues*, publish a Spring `ApplicationEvent` and handle it with `@TransactionalEventListener(phase = AFTER_COMPLETION)` or an async worker — the audit write then happens on its own thread/connection outside the request's critical section. Better still, a **transactional outbox**: write the audit row in the *same* outer transaction into an outbox table, and a separate poller ships it — no second connection, and the write is atomic with the business data (though it then shares the outer's commit/rollback, which may or may not be what you want). 4. **Separate DataSource/pool** for audit writes so audit connection demand can't starve the business pool (adds complexity). ## Second risk: lock & visibility surprises Because the inner tx runs on a **different connection**, it: - **Cannot see** the outer transaction's uncommitted rows (different connection, normal isolation), so audit logic that reads 'current' business rows may see stale/absent data. - **Can block on locks** the outer tx holds. If the outer transaction has locked a row that the inner audit tries to update, the inner waits — and since the outer won't proceed until the inner returns, you get a **real database deadlock** (distinct from pool deadlock). Keep the inner tx touching *disjoint* tables (a dedicated audit table) to avoid this. ## Third risk: masked failures / semantics - The whole point (survives outer rollback) also means an audit row can exist for a business action that **never committed** — downstream consumers of the audit log must tolerate 'attempted but rolled back' entries. - Exceptions from the inner tx (e.g. its own constraint violation) will propagate into the outer method unless caught; an audit failure could then roll back the business transaction — usually the opposite of intent. Wrap the audit call so its failure is logged but swallowed, if audit is best-effort. ## Observability & validation - Monitor pool metrics: `hikaricp_connections_pending`, acquire/wait time, `connectionTimeout` counts, active vs idle. - Load-test at realistic concurrency *and* nesting depth — the deadlock only appears near saturation. - Add tracing to confirm two connections are actually taken and how long the inner tx holds one. ## Decision framework - Need durability independent of the caller **and** synchronous **and** low volume → REQUIRES_NEW is fine; size the pool. - High throughput / hot path → prefer async event listener or outbox to avoid the second-connection-in-line. - Must be atomic *with* the business tx → don't use REQUIRES_NEW at all; write in the same transaction (or outbox).

  • Give the rough formula for a safe pool size when REQUIRES_NEW nests one level deep.
    You need up to two connections per concurrent request in that path, so poolSize should exceed maxConcurrentRequests x 2 (generally (depth+1) x concurrency). Alternatively cap concurrency below poolSize/(depth+1) so a second connection is always available.
  • Why might an @TransactionalEventListener with AFTER_COMMIT be a better fit than an in-line REQUIRES_NEW for auditing?
    It runs after the business transaction settles, so it doesn't hold a second connection while the outer tx is still open — eliminating the nested pool-deadlock risk. With AFTER_COMMIT it only records genuinely committed actions; combined with @Async it also moves the write off the request thread.
  • How can an in-line REQUIRES_NEW audit cause a database (not pool) deadlock?
    The inner tx is on a separate connection and may try to touch a row the still-open outer tx has locked. The inner blocks on that lock, but the outer can't release it until the inner returns — a genuine circular DB wait. Keeping the inner tx on a disjoint audit table avoids it.

saying these in an interview costs you the question

  • Assuming pool usage is unchanged because 'it's the same request'
  • Ignoring that the outer connection stays checked out during suspension
  • Believing the inner tx sees the outer's uncommitted data
  • Not realizing an audit-write failure can roll back the business tx if uncaught

context