As an architect, how would you size connection pools and design the transaction topology of a service so that patterns like nested REQUIRES_NEW cannot deadlock it, and how do you enforce that at scale?
answer
- Bound connections-held-per-request (aim 1)
- safeConcurrency = poolSize / k
- Budget total vs Postgres max_connections / PgBouncer
- Prefer after-commit events / outbox (k=1)
- Dedicated pool for real k=2; enforce via ArchUnit + load tests
basics
~20 sCap the maximum connections any single request can hold at one time (ideally one), and size the pool to peak concurrent requests times that per-request hold, staying within the database's connection ceiling. Ban mid-transaction nesting via review/architecture tests, prefer after-commit events/outbox, and give any unavoidable second-connection work a dedicated pool.
solid answer
~50 sI design so that no request holds more than one pool connection at a time; then pool sizing is just 'peak concurrent requests capped by the DB's max_connections,' and nested-REQUIRES_NEW deadlock is structurally impossible. Where an operation genuinely needs an independent immediate write, I isolate it on its own DataSource/pool so it can never contend with the primary pool. I push most 'independent side effect' cases to after-commit domain events or a transactional outbox, which removes overlap entirely and fits Spring Modulith. For sizing I follow HikariCP's guidance — pools should be small; connections are scarce DB resources — and I budget total connections across all pods against Postgres max_connections (or a PgBouncer layer). I enforce the ban on mid-transaction nesting with code review plus ArchUnit/Modulith tests and load tests that drive concurrency past poolSize/2 so the failure surfaces in CI, not production.
go deeper
Not expected to design pool topology.
Can size a single pool but may miss the cross-pod DB ceiling and enforcement.
Chooses after-commit/outbox and dedicated pools and sizes against the DB ceiling.
Frames it as a per-request connection-hold invariant, budgets all pools against max_connections/PgBouncer, and encodes the rule in ArchUnit/Modulith + load tests so it cannot regress.
### The governing invariant The deadlock exists only because a single request can hold **more than one** pool connection at once. So the architectural rule is: **bound the maximum connections held simultaneously per request** — ideally to 1. If `maxHeldPerRequest = 1`, then `safeConcurrency = maximumPoolSize` and no circular wait can form. If you allow `k` held at once (nested REQUIRES_NEW gives k=2), then `safeConcurrency = floor(maximumPoolSize / k)` and you must size accordingly. ### Sizing the pool - **Small is correct.** HikariCP's own guidance is that connection pools should be **small**; a database serves throughput best with a modest number of busy connections, not hundreds of mostly-idle ones. A common heuristic is roughly `connections = ((core_count * 2) + effective_spindle_count)` as a starting point, then tune with load tests. - **Respect the database ceiling.** Postgres `max_connections` (often ~100–200) is shared across **all** application instances/pods. Total demand = `pods * (sum of maximumPoolSize across all DataSources per pod)`. If each pod has a 10-connection primary pool plus a 5-connection dedicated pool, 20 pods already demand 300 connections — over a default ceiling. Introduce **PgBouncer** (transaction pooling) to multiplex if you need many app instances. - **Account for every pool.** A dedicated pool for independent writes (the clean way to keep a legitimate second-connection use) adds to the total budget — size and cap it explicitly. ### Transaction topology design 1. **Default to one connection per request.** Keep transactions short; do external calls and non-DB work outside the transaction. 2. **Independent side effects → after commit.** Use `@TransactionalEventListener(AFTER_COMMIT)` or a **transactional outbox** so the side write happens on a fresh connection after the primary releases. This is the Modulith-friendly pattern already established in this codebase (domain events). 3. **Unavoidable in-scope independent write → dedicated DataSource.** Bind a second `PlatformTransactionManager` to a separate small pool so the two never share a pool; the two-connection hold no longer forms a cycle within one pool. 4. **Never nest REQUIRES_NEW inside a long-running / high-fan-out loop** — that multiplies hold time and demand. ### Enforcement at scale (make it un-reintroducible) - **Review + lint:** flag `@Transactional(propagation = REQUIRES_NEW)` called from within another `@Transactional` scope. - **Architecture tests:** ArchUnit / Spring Modulith rules to constrain where REQUIRES_NEW may appear, or to require side effects go through the events module. - **Load/soak tests in CI:** deliberately drive concurrency **past `poolSize/2`** through nested paths so the deadlock manifests in the pipeline, not at 2am. Assert on p99 latency and zero `SQLTransientConnectionException`. - **Observability guardrails:** alert on `hikaricp.connections.pending > 0` sustained and on acquire-timeout counts; set `leakDetectionThreshold` in all environments. - **Capacity budget doc:** a living table of every pool, its max size, and the total against `max_connections` / PgBouncer limits, reviewed when adding pods or pools. ### The principal-level framing This is fundamentally a **capacity and invariants** problem, not a Spring-annotation trivia problem. You (a) choose an invariant (≤ k connections per request), (b) size all pools against the shared DB ceiling with that invariant, (c) prefer topologies (after-commit events/outbox) that keep k=1, (d) isolate the rare legitimate k=2 case onto its own pool, and (e) encode the rule in tests and alerts so it cannot silently regress. Then nested-REQUIRES_NEW pool deadlock stops being a lurking incident and becomes a design constraint the system provably satisfies.
- Ten pods each run a 10-connection primary pool and a 5-connection dedicated pool. Postgres max_connections is 150. Is this safe?Demand is 10 pods x (10 + 5) = 150 connections, exactly the ceiling with zero headroom for superuser/maintenance connections and other clients, so it is unsafe. You would reduce pool sizes, cut pods, or front the DB with PgBouncer transaction pooling to multiplex.
- Why is a dedicated pool a valid fix when raising the single shared pool is only a stopgap?A dedicated pool physically separates the outer and inner connection sources, so the two-connection hold can never form a circular wait within one pool. Raising the shared pool keeps both connections in the same pool, so the deadlock is merely deferred to higher load.
saying these in an interview costs you the question
- Sizing pools huge 'to be safe' (ignores DB max_connections and hurts DB throughput)
- Forgetting that pool budget is shared across all pods/instances
- Treating it as an annotation quirk rather than a capacity invariant
- Not enforcing the rule in tests, so it silently regresses