skip to content

A service runs three kinds of work on one shared worker pool: fast in-memory lookups, slow calls to a third-party HTTP API, and periodic report generation. Make the case for splitting them into separate pools, and explain how you would allocate capacity across the pools.

level: principalimportance: should knowfreq 45%

answer

  1. different wait/service ratios = no single correct size
  2. head-of-line blocking: short tasks inherit long tails
  3. bulkhead: a slow dependency drains a shared pool
  4. CPU is the real shared budget; workers are permits
  5. splitting loses statistical multiplexing

basics

~20 s

One pool cannot have one correct size: the three classes have wildly different wait/service ratios, and slow tasks occupy workers so fast ones queue behind them. Split into pools sized per class, budget CPU across them, cap each class's concurrency, and measure utilization and queue wait per pool.

solid answer

~60 s

Three arguments. First, sizing: the correct worker count follows the wait-to-service ratio, which differs by orders of magnitude across these classes, so no single number is right - it is either far too small for the HTTP class or wastefully oversubscribed for the lookups. Second, head-of-line blocking: a shared pool lets long tasks occupy workers while short ones queue behind them, so the fast path inherits the slow path's latency, and the bursty report job is the worst offender. Third, isolation: if the third-party API slows down, its tasks accumulate and consume the shared pool, taking the healthy lookup path down with it. Separate pools act as bulkheads. For allocation, treat CPU as the shared budget: keep the CPU-bound classes near the core count in total, size the blocking class from its wait ratio but cap it at what the dependency can absorb, and give background work a small, deliberately deprioritized pool with a deep queue. Then run per-pool metrics - utilization, queue depth, queue wait, rejection rate - and tune with explicit per-class limits rather than one global number.

go deeper

for a junior

Say that fast work should not sit behind slow work and that pools with different characteristics need different sizes.

for a middle

Explain head-of-line blocking and size each pool from its own wait/service ratio.

for a senior

Lead with failure isolation - a stalled dependency draining a shared pool - and add per-class limits, bounded queues and per-pool telemetry.

for a principal

Treat the partition as an isolation architecture with an explicit CPU budget, name the multiplexing cost you are paying, and define the metrics and limits that let the decision be revisited with evidence.

## Why one pool cannot be sized Pool size follows from the ratio of waiting to computing. In-memory lookups are almost pure service time, so their pool wants roughly the core count. Third-party HTTP calls may wait a hundred times longer than they compute, so their pool wants many times the core count. Report generation is CPU-heavy, long-running, and bursty. A single pool must pick one number, and every choice is wrong for two of the three classes: small enough for the lookups starves the HTTP work; large enough for the HTTP work oversubscribes the CPU whenever reports and lookups are active. ## Head-of-line blocking and the tail With one queue and one set of workers, task ordering couples the classes. A burst of report tasks or slow API calls occupies the workers; the next lookup, which needs a millisecond, waits behind them. The lookup's latency distribution becomes a mixture dominated by the slowest class - a textbook tail-latency amplifier. Fairness within a single pool does not fix it, because the unit of occupancy is a whole task: once a long task holds a worker it keeps it until it finishes. ## Isolation and correlated failure The strongest argument is failure containment. When the third-party dependency degrades from 50 ms to 5 s, its tasks stop leaving the pool. The arrival rate is unchanged and residence time is up a hundredfold, so by Little's law the number resident rises a hundredfold and consumes every worker. The lookup path - which needs nothing from that dependency - now fails too. Separate pools are bulkheads: the blast radius is one class. This is also why a per-class concurrency limit belongs on the calling side even when the executor is shared. ## Allocating capacity CPU is the genuinely shared resource; workers are just permits. 1. **Budget the CPU.** Sum the CPU demand of the CPU-bound classes and keep the total near the core count. Two pools of eight on eight cores means sixteen runnable threads competing - the isolation is real but the CPU is still shared, so protect it with priority or with smaller pools. 2. **Size the blocking class from its ratio, then cap it by the dependency.** The wait-ratio rule may say 200; if the third party degrades past 40 concurrent calls, 40 is the answer and the rest becomes visible queueing or fast rejection at your edge. 3. **Give background work a floor, not a share.** Reports get a small pool, lower priority, and a bounded queue; they are graded on completion, not latency, so running them at high utilization is correct. 4. **Decide the overflow policy per class.** Interactive classes want a shallow queue and fast rejection so callers can retry or degrade; background classes can queue deeply. A single global queue cannot express both. 5. **Consider elasticity.** For blocking classes a pool that grows to a hard ceiling under queueing pressure and shrinks when idle often beats a fixed number, provided the ceiling is set by the dependency. ## The cost of splitting Partitioning loses statistical multiplexing: three pools each sized for their own peak need more total capacity than one pool sized for the aggregate peak, because idle workers in one pool cannot help another. Fragmentation also multiplies the tuning surface and the number of places a change must be made, and every extra pool adds threads competing for the same cores. That is the honest counterargument, and the answer is not 'never split' but 'split along failure and latency boundaries, not along every function' - typically one pool per dependency or per latency class, not one per endpoint. ## How you would know it is right Per-pool telemetry: utilization, queue depth, queue wait, task duration percentiles, rejections. Two signals justify a split after the fact - a latency class whose percentiles track another class's activity, and a dependency incident whose impact extended beyond the tasks that used it. Two signals argue against over-splitting: pools persistently idle while another rejects work, and a total thread count far above what the cores can use. ## The framing to leave with Pool boundaries are an isolation decision first and a sizing decision second. The sizing math tells you how big each pool should be; the partition itself is chosen so that one workload's bad day cannot become another's.

  • What is the strongest argument against splitting, and how do you weigh it?
    Losing statistical multiplexing: separate pools each sized for their own peak need more total capacity than one pool sized for the aggregate peak, since idle capacity cannot be shared, and each extra pool adds threads contending for the same cores plus another thing to tune. The weighing rule is to split along isolation boundaries that matter - one pool per external dependency and per latency class - rather than per feature, and to accept the extra capacity as the price of bounding the blast radius.
  • How would you decide the queue and rejection policy for each pool?
    From what the class is graded on. Interactive classes get a shallow bounded queue so waiting stays inside the client's timeout and excess load is rejected quickly, letting callers shed or degrade rather than pile up. Background classes get deep queues and no urgency, because completion matters more than latency. The bound follows from the target wait: by Little's law, allowed depth is roughly target wait times throughput.
  • Would you give the report pool lower operating-system priority?
    It can help, but priority alone is blunt: it does not bound how many cores the background work occupies, and aggressive priority differences risk priority inversion if the classes share a lock. A small fixed pool plus a bounded queue limits the background footprint more predictably, with lower priority as a secondary refinement.

Supermarket checkouts: the express lane exists not because small baskets matter more, but because putting them behind full trolleys makes their wait unpredictable, and one price check would otherwise stall everybody.

saying these in an interview costs you the question

  • Answering only 'use a bigger pool' without addressing the workload mixture
  • Claiming separate pools cost nothing, ignoring lost multiplexing and extra threads on the same cores
  • Splitting per endpoint or per feature rather than per dependency and latency class
  • Forgetting that separate pools still contend for one CPU budget
  • Assuming priorities on a shared queue remove head-of-line blocking, when task occupancy is the blocking unit

context