skip to content

Blocking a Pipeline

A single blocking call inside an asynchronous pipeline can stall every other request that shares its worker. Interviewers want the symptom, how you detect it, and the fix.

on this pageshow

questions

4

Why does one pipeline stage calling a synchronous routine that waits slow down unrelated requests sharing its worker pool?

level: middleimportance: must knowfreq 64%

answer

  1. shared workers, not one per request
  2. occupied is not the same as busy
  3. pool sized to the processor count
  4. a few waits park every worker
  5. unrelated routes degrade together

basics

~20 s

A parked worker is unavailable to every subscription that shares it. Asynchronous pipelines multiplex many requests over a pool sized to the processor count, so a few simultaneous waits occupy every worker and unrelated routes queue behind them.

solid answer

~50 s

An asynchronous pipeline does not give each request a worker of its own. Every subscription's stages are dispatched onto the same small pool, which is sized roughly to the processor count because non-blocking stages are expected to transform a value and return in microseconds. A stage that calls a synchronous routine holds its worker for the whole wait: the worker is *occupied* while consuming no processor at all, and an ordinary synchronous call offers no point at which the runtime could take it back. With four or eight workers, four or eight concurrent calls to that one endpoint park the pool, and every other subscription's signals — values, completions, even cancellations — sit queued behind them. That is why the symptom is system-wide: routes with no code in common slow down and time out together, correlated with traffic to the single feature that blocks.

code

pseudocode · 10 lines
pseudocode
// one small pool of workers runs the stages of every subscription

pipeline =
    sourceOf(requestId)
      .transform(id -> legacyLookup(id))   // synchronous: returns only when the reply arrives
      .transform(row -> format(row))

// for the whole of legacyLookup the worker is parked:
// it cannot deliver a value, a completion or a cancellation
// to any other subscription that shares the pool

go deeper

for a junior

Recall that stages do not each get their own worker. A small shared pool runs them all, so a stage that waits is holding something other requests need.

for a middle

Explain why the pool is sized to the processor count, what it means for a worker to be occupied without consuming processor time, and how few concurrent waits it takes to hold all of them.

for a senior

Show how you would confirm it live — latency rising together on unrelated routes, near-idle processors, workers parked in repeated stack snapshots — and then isolate the offending stage onto a worker set sized for waiting.

for a principal

Weigh whether a synchronous dependency belongs inside this service at all, and what isolation and bounds you require before any team is allowed to call one from a shared pipeline.

## How the work is actually dispatched An asynchronous pipeline is a chain of **stages** assembled once and then run for every subscription. Nothing in that chain owns a thread. When a **signal** arrives — a value, a completion or a cancellation — the runtime hands it to a free **worker** from a shared pool, the worker runs the stage that was waiting for it, and the worker goes straight back to the pool. Thousands of in-flight subscriptions are multiplexed over a handful of workers, and that multiplexing is the whole reason the arrangement scales: the stack memory and the switching cost of a worker are paid a few times, not once per request. The design rests on one assumption: - every stage **returns promptly**, having done nothing but transform a value and hand it downstream; - anything that takes real time is expressed as *another source in the chain*, so the waiting is done by the transport, which parks nothing; - therefore a pool sized to the processor count is sufficient, because a worker is never idle-but-held — it is either computing or back in the pool. ## Occupied is not the same as busy A synchronous call breaks that assumption invisibly. The worker enters a stack frame and does not leave it until the reply arrives. It consumes no processor while it waits, so nothing in a usage graph flags it, but it is unavailable: it cannot deliver a value for any other subscription, cannot run a completion, cannot even propagate a cancellation. **Occupied and busy are different states**, and only one of them shows up as load. There is no reclamation mechanism to rescue it. A stage that returns a deferred result gives the runtime a hand-back point; a stage that returns a finished value has, by definition, already waited on the worker that ran it. ## The arithmetic, and why a small pool has no slack Average occupancy follows **Little's Law**: the average number of workers held is the arrival rate of blocking calls multiplied by how long each one waits. Let `n` be the average number of workers occupied. | Calls per second into the blocking stage | Wait per call | `n` workers occupied on average | |---|---|---| | 5 | 200 ms | 1 | | 20 | 200 ms | 4 | | 200 | 20 ms | 4 | | 40 | 500 ms | 20 | On a four-worker pool, the second and third rows are already total occupancy — and the third row is the dangerous one, because a 20 ms wait looks harmless in a single trace. The pool was deliberately sized with **no slack**: extra workers buy nothing when every stage returns in microseconds, and they cost stack memory and context switches. That efficiency is exactly what leaves no headroom to absorb waiting. ## What it looks like from outside 1. Latency climbs on routes that share no dependency, no data store and no code, because the only thing they share is the pool. 2. Processor usage stays low while latency and timeouts rise — the signature of waiting rather than computing. 3. The delay appears **before** the stage runs: the time from accepting a request to starting its first stage grows, while the stage's own duration is unchanged. 4. The degradation correlates with traffic to one feature, which is the thread to pull. ## The fix, and the fix that isn't - Move that one stage onto a **separate worker set sized for waiting** — many more workers than processors, because they spend their lives parked — and leave the rest of the chain on the small non-blocking pool. - Give the isolated region a **bound**, so a dependency that slows down cannot grow it without limit, and a **timeout**, so a call that never returns eventually releases its worker. - Enlarging the small pool is not a fix. It raises how many concurrent waits are needed to park it, and traffic supplies them; you also pay the switching cost the small pool existed to avoid. - Retrying into the same parked pool makes it worse: each retry adds occupancy to the resource that is already the constraint. ## Why this bites late The stage is usually correct and usually fast. At low traffic, occupancy stays under the worker count and nothing is visible; the first symptom often arrives with a traffic peak, a dependency that got slower, or a feature flag switched on for everyone. Because the damage lands on *other* routes, the investigation typically starts in the wrong place — on the endpoint that got slow, not the endpoint that made everything slow.

  • What is the fix once you have identified the stage that waits?
    Run that one stage on a separate worker set sized for waiting — far more workers than processors, since they sit parked — and leave the rest of the chain on the small non-blocking pool. Bound the isolated region and give the call a timeout, so a dependency that slows cannot grow occupancy without limit.
  • Why does enlarging the small non-blocking pool not solve it?
    It moves the threshold rather than removing the wait. With twice the workers you need twice the concurrent calls to park them, and traffic supplies that. You also pay the cost the small pool was chosen to avoid: more stacks in memory and more context switching for stages that never needed to wait.
  • If the call waits only 20 milliseconds, is it still a hazard?
    Yes, once the arrival rate is high enough. By Little's Law the average occupancy is rate times wait, so 200 calls a second at 20 ms holds about four workers — a whole small pool. Short waits are more dangerous in one way: they surface as a latency floor on every route rather than as obvious timeouts.

Four tellers serve every queue in a bank. If one teller sits on hold with a supplier for ten minutes, he is occupied without serving anyone, and every queue slows — not just the one that needed the supplier.

saying these in an interview costs you the question

  • Thinks only the endpoint that blocks gets slower.
  • Assumes the runtime detects the wait and reuses the worker elsewhere.
  • Reaches for a bigger pool instead of moving the blocking stage.
  • Reads near-idle processor usage as proof nothing is stuck.
  • Believes a stage cannot block because it sits inside a pipeline.
  • Adds retries, which pile more occupancy onto the parked pool.
open as a page

A pipeline has no obvious network waits, yet its workers still park; where does the blocking hide?

level: middleimportance: should knowfreq 46%

basics

~20 s

Blocking hides in anything that returns a finished value: a synchronous data-access driver, a contended lock, first-use setup such as opening a connection or resolving a name, a file or device read, and a synchronous logging or metrics write.

open as a page

Every route on an asynchronous service is slow; what evidence separates parked workers from one slow dependency?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Ask which routes degraded. A slow dependency hurts only its callers; parked workers hurt routes with nothing in common, alongside near-idle processors, growing time before a stage starts, and stack samples finding the pool waiting.

open as a page

Several pipeline stages in your service offload blocking calls to one waiting-sized pool; when do you split that pool per dependency?

level: principalimportance: should knowfreq 38%

basics

~20 s

Split when one dependency's slowdown must not stall the others. One shared offload pool reproduces the original failure a level down; separate pools contain it but reserve capacity. A permit per dependency over one pool buys most of that isolation.

open as a page