skip to content

An Object Pool with a fixed maximum size receives an acquire request while every instance is already checked out. What policies can the pool apply, and how do you choose between them?

level: middleimportance: must knowfreq 66%

answer

  1. Block-with-timeout / fail-fast / overflow / block-forever
  2. Timeout shorter than the caller's deadline
  3. Bound the wait queue too
  4. Little's Law sizes the pool, not intuition
  5. Exhaustion = leak or slow downstream, usually

basics

~20 s

The pool can make the caller wait until someone returns an object (usually with a timeout), fail fast with an error, or — in a soft-limited pool — create a temporary extra instance. Waiting forever is the dangerous option: it can hang the whole system.

solid answer

~50 s

Exhaustion policies: (1) block with a timeout — the caller parks on a wait queue and either gets an instance or receives a timeout error; this is the default for most connection pools. (2) Fail fast — reject immediately, shedding load and preserving latency. (3) Grow beyond the max (overflow/soft limit) — trades the safety cap for availability. (4) Block indefinitely — almost always wrong, because one stuck borrower deadlocks every caller and the failure surfaces as an unexplained hang. Choose by asking what the caller can do with the answer: an interactive request with a client-side deadline should fail fast or use a short timeout well inside that deadline; a batch job can afford to wait. Add a bounded wait queue so pending acquirers don't grow without limit, and always emit metrics — wait time, wait-queue depth, timeout rate — because exhaustion is a capacity signal, not just an error. Remember exhaustion is often a symptom of leaks or slow downstream calls, not of an undersized pool.

code

pseudocode · 7 lines
pseudocode
// deadline-aware acquisition
remaining = request.deadline - now()
if (remaining < MIN_USEFUL) return SHED_LOAD

conn = pool.acquire(timeout = min(remaining * 0.5, 250ms))
  ?: return error("pool exhausted; inUse=" + pool.inUse()
                  + " pending=" + pool.pending())

go deeper

for a junior

Name the options: wait (with a timeout), or fail with an error. Say why waiting forever is dangerous.

for a middle

Add bounded wait queues, timeout chosen relative to the caller's deadline, and that exhaustion is often caused by leaks or slow downstream calls.

for a senior

Bring in Little's Law sizing, FIFO vs barging handoff, nested-acquire deadlock, and the metrics you'd watch (pending count, wait-time p99, timeout rate).

for a principal

Position the pool as an admission-control point in a system-wide backpressure design: deadline propagation, load shedding above the pool, separate pools for critical paths, and the failure mode you deliberately choose under overload.

## What "exhaustion" means A pool is exhausted when **every instance it is allowed to create is currently checked out** and a new caller asks for one. It is not an error in itself — it is the pool doing its job as a concurrency limiter. The design question is what the pool does with the waiting caller. ## The policy menu ### 1. Block with timeout (most common default) The caller is parked on a wait queue. If an instance is released within the timeout, it is handed to a waiter; otherwise the acquire call throws/returns a timeout error. - **Pros**: absorbs short bursts without failing; smooth under normal jitter. - **Cons**: converts a capacity problem into latency. If the timeout is longer than the caller's own deadline, you burn resources on work nobody will use. - **Handoff order matters**: FIFO waiter wakeup gives fairness and bounded worst-case wait; LIFO/barging gives better throughput but can starve a waiter indefinitely. ### 2. Fail fast Return an error immediately when nothing is idle. - **Pros**: instant, honest backpressure; latency stays predictable; upstream can retry elsewhere, serve degraded content, or shed the request. - **Cons**: no tolerance for micro-bursts — a pool that would have freed an instance in 3 ms still rejects. - Common in latency-critical paths and in systems with a circuit breaker or load shedder above the pool. ### 3. Overflow / soft limit A "core size" of pooled instances plus permission to create temporary extras up to a hard ceiling, destroyed on return. - **Pros**: rides out spikes. - **Cons**: the extras cost full construction price and can overwhelm the downstream resource — the very thing the cap protected. Only safe when the downstream has real headroom. ### 4. Block forever - Almost always a bug. One borrower that never returns (a leak, an infinite loop, a socket read with no timeout) will eventually block **every** caller. The symptom is a totally silent system: no errors, no logs, just threads parked in `acquire`. If you must support it, make it opt-in and never the default. ### 5. Caller-supplied instance / bypass Some designs let a caller construct an unpooled instance when the pool is dry. This defeats the cap and is rarely appropriate for connections; occasionally acceptable for buffers, where the fallback is only a memory allocation. ## Choosing: work backwards from the caller's deadline The key question is **"what does the caller do with a slow answer?"** - User-facing request with a 300 ms budget → acquire timeout well under that (e.g. 100–250 ms), then fail with a clear error. Waiting 30 s for a connection to answer a request the browser abandoned at 1 s is pure waste. - Async worker or batch job → longer wait is fine; throughput matters more than tail latency. - Health checks and admin endpoints → should ideally use a separate small pool so they still work when the main pool is saturated. A useful discipline: the acquire timeout should be **strictly shorter** than the caller's total deadline, so a timeout is reported as a pool problem rather than as a generic upstream timeout. ## Bounding the wait queue Blocking without a bounded queue just moves the unboundedness one layer up: thousands of parked threads or futures consume memory and, in thread-per-request systems, threads. Bound the queue and reject beyond it. This is the same reasoning as bounding a thread pool's task queue. ## Sizing: why bigger is not better People reflexively raise `maxPoolSize` when they see exhaustion. Often that makes things worse: - The downstream database has its own connection limit and per-connection memory cost; blowing past it causes server-side thrashing or refused connections. - Past the point where the resource is saturated, more concurrency adds queueing everywhere without adding throughput. **Little's Law** (`concurrency = throughput × latency`) gives the honest target: to serve 500 requests/s each holding a connection for 10 ms, you need about 5 concurrent connections, not 200. - Small pools give better latency *and* clearer backpressure. HikariCP's well-known guidance — a pool of ~10 often beats a pool of 100 — comes directly from this. ## Exhaustion is a symptom — diagnose before tuning Before resizing, check: 1. **Leaks** — objects acquired and never released. Look for a monotonically rising in-use count that never drops at idle traffic. 2. **Slow downstream** — a query that went from 5 ms to 500 ms multiplies required concurrency 100×. Fix the query, don't grow the pool. 3. **Holding across I/O** — code that acquires a connection and then makes an HTTP call before releasing holds the resource for the whole remote round trip. 4. **Nested acquires** — a task that holds one instance while waiting for a second can self-deadlock the pool. If pool size is N and each task needs 2, any N simultaneous tasks each holding 1 deadlock forever. Fix by acquiring both up front, using a separate pool, or sizing so `size > tasks × (needed − 1)`. ## Observability that makes exhaustion debuggable Minimum metrics: idle count, in-use count, total created, pending-acquire count, acquire wait-time distribution (p50/p99), acquire timeout rate, and creation failures. Pending count > 0 sustained is the leading indicator; timeouts are the lagging one. Log timeouts with the current in-use count and, if the pool supports it, the stack traces of long-held instances.

  • A service under load shows rising acquire timeouts. Why might doubling the maximum pool size make things worse?
    If the bottleneck is the downstream resource, more concurrent users of it increase contention and per-operation latency, so each instance is held longer and the pool saturates again at a higher cost — plus you may exceed the server's own connection limit. Little's Law says required concurrency = throughput × hold time; the fix is usually to shrink hold time.
  • How can a bounded pool deadlock even with no leaks and no bugs in the pool itself?
    If a task acquires a second instance while still holding the first and pool size is N, N tasks can each hold one and wait forever for another. Acquire all needed instances up front, use separate pools, or size the pool above tasks × (needed − 1).

A restaurant with a fixed number of tables: you can seat waiting guests in a queue with a quoted wait time (block with timeout), turn them away at the door (fail fast), or drag extra tables from the storeroom (overflow) — but telling guests to wait indefinitely with no update is how a queue turns into a riot.

saying these in an interview costs you the question

  • "Just make the pool bigger" as the first response to exhaustion
  • Blocking indefinitely on acquire with no timeout
  • Setting an acquire timeout longer than the caller's own request deadline
  • Blocking with an unbounded wait queue, moving the resource exhaustion one layer up
  • Treating exhaustion purely as an error rather than as a capacity/backpressure signal
  • Assuming an overflow policy is free — the extras still hit the downstream resource

context