skip to content

A fixed-cell object pool that grew during a traffic spike still holds those cells hours later. What design choices control that retention?

level: seniorimportance: nice to knowfreq 27%

answer

  1. sized to the worst minute
  2. nothing ever votes to shrink
  3. on the free set, so not a leak
  4. cap, evict, shrink or refuse to grow
  5. window longer than the spike interval

basics

~20 s

A pool that grows on demand and never trims is sized to its all-time peak, not to steady state, so a spike becomes a permanent footprint. The controls are a hard cap, idle eviction above a floor, shrinking on a low-water observation, or refusing to grow at all.

solid answer

~60 s

Growth in a pool is a decision that is easy to make implicitly: acquire finds the free set empty, creates a cell, and that cell is now reachable and reusable forever. Nothing in the design ever concludes that a cell is surplus, so the pool ends up sized to the highest concurrency it has ever seen. This is retention rather than a leak — the cells are on the free set and will be reused — but the process footprint is a high-water mark, not a steady-state number. Four controls exist, and they are not exclusive: a **hard cap** that turns a spike into waiting or a clean failure instead of growth; **idle eviction** that discards cells unused for a period, above a minimum floor; **shrinking** toward an observed low-water mark over a window; and a **fixed pool** with backpressure, which refuses to grow at all. Each trades footprint against the cost of recreating cells when load returns, so the eviction window must be longer than the spike period or the pool thrashes.

go deeper

for a junior

Remember that a pool which is allowed to create new cells will keep them afterwards unless something is written to throw them away.

for a middle

Explain why the held count tracks the all-time peak, and distinguish cells sitting idle on the free set from cells that were never returned.

for a senior

Pick a control and defend its parameters: a floor under eviction, a window longer than the spike interval, and a cap that covers every shard rather than each one.

for a principal

Decide what the service should do at the cap — wait or fail — and argue why a bounded, visible degradation is a better fleet-wide trade than footprint that ratchets upward.

## Growth is a decision, not an accident A pool that may grow does so at exactly one moment: `acquire` is called, every cell is out on loan, and the code chooses to create another rather than wait or fail. The new cell is indistinguishable from the others afterwards. When load falls, the cells come back to the free set and stay there — reachable, reusable, and never reconsidered, because nothing in a plain pool ever forms the opinion that it has too many. The result is a pool **sized to its all-time peak concurrency**, and a process footprint that is a record of the worst minute the service has had rather than a description of what it is doing now. In a fleet this shows as memory that only ever ratchets upward and resets on restart, which is easy to misread as a leak. ## Why it is retention and not a leak The distinction matters because it changes the investigation. A leaked cell is one nobody returned: it is not on the free set, no borrower will give it back, and the pool drains until acquirers wait or fail. A retained cell is on the free set, available immediately, and costs only its bytes. The tell is the free set's own size — if it is large while throughput is low, the pool is retaining, not leaking. ## The four controls - **Hard cap.** The pool never exceeds N cells; when all N are lent out, acquirers wait or fail. This converts an unbounded memory response to a spike into a bounded, visible latency or error response — which is usually the trade you want, because the spike was going to exceed some limit either way and this one you chose. - **Idle eviction.** A cell unused for longer than a chosen period is discarded, subject to a **minimum floor** so the steady-state path never pays creation cost. The period must exceed the natural spike interval, or the pool discards cells shortly before it needs them again. - **Low-water shrink.** Over a window, record the smallest number of cells simultaneously in use; anything above that plus a margin is surplus and can be dropped at the end of the window. This adapts to a workload whose steady state changes across a day, and it is gentler than per-cell timers. - **No growth at all.** A fixed pool created at startup with backpressure for everyone else makes the footprint a constant and moves the spike's cost entirely into the queue. The failure mode is explicit, which for capacity-limited resources is often preferable to memory that grows quietly. | Control | Gives back | Costs | |---|---|---| | Hard cap | Bounded worst case | Waiting or failures during a spike | | Idle eviction | Footprint after the spike | Recreation cost when load returns | | Low-water shrink | Adapts to a changing baseline | A window's delay before it reacts | | Fixed size | A constant, predictable footprint | No headroom for an unusual burst | ## Two multipliers people miss 1. **Sharding.** Pools are often split per worker or per shard to avoid contention on the free set. A pool sharded 8 ways, each shard grown to 512 cells, holds 4096 cells even if total concurrency never exceeded 600 — because each shard grew to its own local peak and no shard can lend to another. Sharded pools need a global cap, not only a per-shard one. 2. **Cell size.** Retention is counted in bytes, not cells. A pool that grew by 200 cells of 64 KB is holding roughly 12.8 MB; the same 200 cells at 4 MB each is 800 MB. A cap expressed in cells hides this, so express it in bytes when cell size varies between deployments. ## What to measure before choosing Three series answer the question without guessing: cells in use over time, cells held over time, and the gap between them. A large, flat gap after a spike is pure retention and argues for eviction or a low-water shrink. A gap that closes and reopens repeatedly argues for a floor and a longer window, because trimming into that pattern pays recreation cost over and over. And a cells-in-use line that flattens exactly at the pool's size is the cap doing its job — at which point the interesting metric is the wait time at acquire, not memory at all.

  • How do you tell this apart from cells that are never being returned?
    Look at the free set. Retention means many cells are held and idle while throughput is low, and acquire never waits. A missed return means the opposite: cells held is high but free is near zero, acquirers start waiting or failing, and the shortfall only ever grows because no cell comes back.
  • What goes wrong if the idle-eviction period is set too short?
    The pool thrashes: it discards cells during the quiet stretch between bursts and recreates them when the next burst arrives, paying creation cost repeatedly and adding latency exactly at the moment load is rising. The eviction period has to be longer than the natural gap between spikes, and a minimum floor should keep the steady-state path warm.

saying these in an interview costs you the question

  • Reports pool growth after a spike as a memory leak without checking the free set.
  • Adds idle eviction with no minimum floor, so the steady path recreates cells constantly.
  • Caps a sharded pool per shard only, multiplying the true peak by the shard count.
  • Expresses the cap in cells when cell size varies widely between deployments.
  • Lets the pool grow without limit so a spike becomes a process-wide memory event.