skip to content

questions

5

When a worker pool's bounded work queue is full and a new task is submitted, several policies are possible: refuse with an error, drop the newest item, drop the oldest item, block the submitter until space appears, or execute the task on the submitting thread. Walk through the tradeoffs and when each is the right choice.

level: middleimportance: must knowfreq 55%

answer

  1. full queue = forced decision, every option loses something
  2. refuse: fastest, most honest, needs backoff
  3. drop-newest: optional work only, must count it
  4. drop-oldest: latest-value semantics, freshness over fairness
  5. block/caller-runs: real backpressure, can stall the intake loop

basics

~20 s

Refuse = fastest, most honest signal; needs a caller that can handle it. Drop-newest sheds load and keeps old work; drop-oldest keeps the freshest data. Block propagates backpressure but can stall the producer. Run-inline throttles the producer by making it do the work.

solid answer

~60 s

There is no default; the policy encodes what your system values under overload. - **Refuse (fail fast)**: the submitter gets an immediate error. Cheapest, most observable, and it lets the caller retry, reroute, or degrade. Right when work is externally retryable or optional. - **Drop newest**: silently discard the arriving task. Right for optional, fungible work such as metrics or cache warm-ups — but silent loss must be counted, or you will not know it happened. - **Drop oldest**: discard the head to make room. Right when only the freshest item matters — a latest-value sensor reading, a UI refresh, a heartbeat. Wrong for anything with per-item semantics. - **Block the submitter**: propagates backpressure straight to the producer. Correct for closed pipelines with a bounded producer; dangerous when the submitter is an event loop, accept loop, or a thread whose stall stops intake. - **Run on the submitting thread**: an implicit throttle — the producer is busy for the duration and cannot submit more. Strong backpressure with real hazards. Always instrument the policy: rejections and drops are your saturation metric.

code

text · 15 lines
text
submit(task):
  if queue.offer(task):        return ACCEPTED

  switch policy:
    REFUSE:      metrics.rejected++;  return ERROR_OVERLOADED
    DROP_NEW:    metrics.dropped++;   return ACCEPTED   # silent - must count
    DROP_OLD:    victim = queue.removeHead(); metrics.evicted++
                 queue.offer(task);   return ACCEPTED
    BLOCK:       queue.put(task, timeout)  # backpressure; may stall producer
    CALLER_RUNS: task.run()          # producer throttled while it works

# complementary to any policy:
worker loop:
  task = queue.take()
  if now() > task.deadline: metrics.expired++; continue   # skip stale work

go deeper

for a junior

Name the options and one sentence each: refuse now, drop the new item, drop the oldest item, make the submitter wait, or make the submitter run it. Say that the choice depends on whether the work is optional.

for a middle

Give the tradeoff per policy and match each to a workload — telemetry versus payments versus latest-value sensor data — and stress that silent drops must be counted.

for a senior

Add the operational failure modes: retry storms after fail-fast, stalled accept or event loops after blocking, self-deadlock when a pool worker blocks on its own queue, and per-item deadlines so workers skip stale tasks.

for a principal

Treat it as the system's overload contract: which classes of work are shed first, which are protected by separate pools, how the signal reaches producers and autoscaling, and how the policy is expressed in the API so callers know what acceptance means.

## The decision point Bounding a queue creates a moment: the queue is full and a task arrives. Something must happen, and every option loses something. The policy is a statement about which loss you prefer. Making it explicit is the whole value of the bound. ## Refuse (fail fast / abort) The submission fails immediately with an error the caller can see. - **Gains**: constant-time, no memory growth, no producer stall, a clean saturation metric, and the decision moves to the caller — who often has better options (retry with backoff, try a different node, return a cached or degraded result, tell the user). - **Loses**: work that could have completed a moment later, and it converts overload into caller-visible errors. - **Use when** the work has an owner who can respond: a request-serving pool, a client with retry logic, an upstream that can reroute. - **Pitfall**: naive callers that retry immediately turn refusal into a retry storm. Refusal must be paired with backoff, jitter, and ideally a circuit breaker. ## Drop newest (discard the arrival) The task is silently discarded and the submission reports success. - **Gains**: never blocks, never grows, keeps the pool working on already-accepted items. - **Loses**: the work, silently — the most dangerous property in the list. Silent loss with no counter is how systems develop mysterious gaps. - **Use when** items are optional and fungible: telemetry samples, cache pre-warms, opportunistic prefetch, best-effort notifications. - **Never use for** anything with a durability or exactly-once expectation. ## Drop oldest (discard the head) Evict the item that has waited longest to make room for the new one. - **Gains**: the queue holds the *freshest* items, so under overload you serve current state instead of history. It also bounds staleness rather than bounding only count. - **Loses**: the oldest item, which in a FIFO pipeline is usually the one someone has been waiting for the longest — badly unfair if items represent user requests. - **Use when** the queue carries *latest-value* semantics: sensor readings, position updates, UI repaints, health snapshots. In those cases stale items are actively worse than nothing. ## Block the submitter The submitting thread waits until space appears (optionally with a timeout). - **Gains**: true backpressure — the producer's rate is forced down to the consumer's rate, with no loss and no error. In a closed pipeline (stage A feeds stage B) this is often exactly right and requires no policy elsewhere. - **Loses**: liveness properties of the submitter. If the submitter is a shared event loop, an accept loop, a request thread with its own deadline, or itself a pool worker, blocking there propagates the stall upward — connections stop being accepted, unrelated work stops, and latency degrades globally. Blocking a pool's own worker on the same pool's queue is a classic self-deadlock. - **Use when** the producer is a dedicated thread whose only job is feeding this stage, and prefer *blocking with a timeout* so the stall is bounded and reportable. ## Run on the submitting thread (caller-runs) The submitter executes the task itself instead of queuing it. - **Gains**: no loss, no error, and an automatic throttle — while the producer is executing, it is not producing. It also degrades gracefully instead of cliff-edging. - **Loses**: the submitter's latency and, more importantly, its role. It is really "block" with useful work attached, and it inherits the same hazards. ## Cross-cutting rules 1. **Count everything.** Rejections, drops, block time. A silent policy is an unobservable outage. These counters are also the natural autoscaling and alerting signal. 2. **Add deadlines, not just depth limits.** Depth bounds the count; a per-item deadline bounds *staleness*. Workers should discard items whose deadline has already passed rather than execute work no one is waiting for. 3. **Consider LIFO under overload.** When you are already behind, serving newest-first keeps *some* requests fast instead of making all of them slow — the same intuition that motivates drop-oldest. 4. **The policy is part of the API contract.** Callers must know whether a submission means "accepted," "maybe accepted," or "will run, eventually, possibly on your thread." 5. **Do not mix silently.** Choosing drop for a pool that also carries must-run work is how important tasks vanish; separate pools per criticality class instead.

  • Which policy would you pick for a pool that writes metrics samples, and which for a pool that processes customer payments?
    Metrics: drop newest (or drop oldest), because samples are optional and fungible and losing a few is far better than stalling the application that emits them — but count the drops so gaps are visible. Payments: never drop and never silently accept. Refuse fast so the caller sees an explicit error and can retry idempotently, or push the work to durable storage rather than an in-memory queue at all.
  • Refusing seems safest, yet it can still bring a system down. How?
    Through retry amplification. If callers retry immediately on refusal, each rejected request returns as one or more new arrivals, so the arrival rate rises exactly when the system is already over capacity — congestion collapse. Refusal only works when paired with client-side exponential backoff with jitter, retry budgets or circuit breakers, and ideally a response that tells the caller how long to wait.
  • Why add per-item deadlines when the queue is already bounded?
    A depth bound limits how many items wait; it does not limit how long any one waits when the pool slows. Recording each item's deadline lets workers skip items whose caller has already given up, so capacity goes to work whose result will actually be used. It also converts an invisible latency problem into an explicit "expired tasks" counter.

A full emergency room. You can turn ambulances away so they go elsewhere (refuse), stop admitting new arrivals (drop newest), discharge the longest-waiting patient (drop oldest), make the ambulance idle at the door (block), or hand the driver a stethoscope (caller-runs). Each is defensible for a different kind of patient, and none is free.

saying these in an interview costs you the question

  • Treating one policy as universally correct rather than as a statement about what the system values under overload.
  • Choosing a silent drop policy without a counter, so loss is invisible.
  • Blocking the submitter without noticing the submitter is an accept loop, event loop, or a worker of the same pool.
  • Assuming drop-oldest is fair — it penalizes the longest-waiting item, which is usually the worst choice for user requests.
  • Recommending fail-fast without client-side backoff, producing a retry storm.

context

open as a page

A worker pool is fed by an in-memory work queue that has no size limit, so every submission is accepted. What are the failure modes of that design, and what does putting a bound on the queue actually buy you?

level: middleimportance: must knowfreq 62%

basics

~20 s

An unbounded queue never rejects, so overload becomes unbounded memory growth and unbounded latency instead of a visible failure. Bounding it turns overload into an immediate, observable signal you can shed, retry, or push back on.

open as a page

One way to handle a full work queue is to execute the rejected task on the thread that submitted it instead of queuing it. Explain why that creates backpressure, and what can go wrong with it in a real system.

level: seniorimportance: should knowfreq 38%

basics

~20 s

While the submitter runs the task itself it cannot submit anything else, so the intake rate is throttled to the pool's drain rate automatically. The risks: whatever else that thread was responsible for stalls — accepting connections, an event loop, ordering guarantees — and if the task needs the same pool, it can deadlock.

open as a page

Contrast a worker pool whose intake is a synchronous handoff — a submission succeeds only if a worker is free to take it immediately — with one that buffers submissions in a queue of some depth. How does the choice change the pool's behavior under load, including when the pool is allowed to grow?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Direct handoff has zero buffer: a submission either starts now or fails now, so pressure is felt instantly and latency is not hidden. Buffering absorbs bursts but delays that signal — and in pools that only add workers when intake fails, a deep buffer means the pool never grows.

open as a page

You own a service whose worker pool is periodically overwhelmed. Design its behavior at the submission boundary — how the system should decide what to accept, what to shed, and how that decision reaches the producers — and justify the choices you would make.

level: principalimportance: should knowfreq 34%

basics

~20 s

Decide the overload contract explicitly: bound the queue from a latency budget, separate work by criticality into its own pools, shed the lowest-value class first, attach deadlines so stale work is skipped, and make the refusal a signal producers act on with backoff and capacity planning.

open as a page