skip to content

Schedulers & Threading

Which worker runs each stage, how to move work between them, and why one blocking call on a small pool stalls everything. Misplaced switching is the most common reactive performance bug.

on this pageshow

explore

questions

23

Why does one pipeline stage calling a synchronous routine that waits slow down unrelated requests sharing its worker pool?

level: middleimportance: must knowfreq 64%

answer

  1. shared workers, not one per request
  2. occupied is not the same as busy
  3. pool sized to the processor count
  4. a few waits park every worker
  5. unrelated routes degrade together

basics

~20 s

A parked worker is unavailable to every subscription that shares it. Asynchronous pipelines multiplex many requests over a pool sized to the processor count, so a few simultaneous waits occupy every worker and unrelated routes queue behind them.

solid answer

~50 s

An asynchronous pipeline does not give each request a worker of its own. Every subscription's stages are dispatched onto the same small pool, which is sized roughly to the processor count because non-blocking stages are expected to transform a value and return in microseconds. A stage that calls a synchronous routine holds its worker for the whole wait: the worker is *occupied* while consuming no processor at all, and an ordinary synchronous call offers no point at which the runtime could take it back. With four or eight workers, four or eight concurrent calls to that one endpoint park the pool, and every other subscription's signals — values, completions, even cancellations — sit queued behind them. That is why the symptom is system-wide: routes with no code in common slow down and time out together, correlated with traffic to the single feature that blocks.

code

pseudocode · 10 lines
pseudocode
// one small pool of workers runs the stages of every subscription

pipeline =
    sourceOf(requestId)
      .transform(id -> legacyLookup(id))   // synchronous: returns only when the reply arrives
      .transform(row -> format(row))

// for the whole of legacyLookup the worker is parked:
// it cannot deliver a value, a completion or a cancellation
// to any other subscription that shares the pool

go deeper

for a junior

Recall that stages do not each get their own worker. A small shared pool runs them all, so a stage that waits is holding something other requests need.

for a middle

Explain why the pool is sized to the processor count, what it means for a worker to be occupied without consuming processor time, and how few concurrent waits it takes to hold all of them.

for a senior

Show how you would confirm it live — latency rising together on unrelated routes, near-idle processors, workers parked in repeated stack snapshots — and then isolate the offending stage onto a worker set sized for waiting.

for a principal

Weigh whether a synchronous dependency belongs inside this service at all, and what isolation and bounds you require before any team is allowed to call one from a shared pipeline.

## How the work is actually dispatched An asynchronous pipeline is a chain of **stages** assembled once and then run for every subscription. Nothing in that chain owns a thread. When a **signal** arrives — a value, a completion or a cancellation — the runtime hands it to a free **worker** from a shared pool, the worker runs the stage that was waiting for it, and the worker goes straight back to the pool. Thousands of in-flight subscriptions are multiplexed over a handful of workers, and that multiplexing is the whole reason the arrangement scales: the stack memory and the switching cost of a worker are paid a few times, not once per request. The design rests on one assumption: - every stage **returns promptly**, having done nothing but transform a value and hand it downstream; - anything that takes real time is expressed as *another source in the chain*, so the waiting is done by the transport, which parks nothing; - therefore a pool sized to the processor count is sufficient, because a worker is never idle-but-held — it is either computing or back in the pool. ## Occupied is not the same as busy A synchronous call breaks that assumption invisibly. The worker enters a stack frame and does not leave it until the reply arrives. It consumes no processor while it waits, so nothing in a usage graph flags it, but it is unavailable: it cannot deliver a value for any other subscription, cannot run a completion, cannot even propagate a cancellation. **Occupied and busy are different states**, and only one of them shows up as load. There is no reclamation mechanism to rescue it. A stage that returns a deferred result gives the runtime a hand-back point; a stage that returns a finished value has, by definition, already waited on the worker that ran it. ## The arithmetic, and why a small pool has no slack Average occupancy follows **Little's Law**: the average number of workers held is the arrival rate of blocking calls multiplied by how long each one waits. Let `n` be the average number of workers occupied. | Calls per second into the blocking stage | Wait per call | `n` workers occupied on average | |---|---|---| | 5 | 200 ms | 1 | | 20 | 200 ms | 4 | | 200 | 20 ms | 4 | | 40 | 500 ms | 20 | On a four-worker pool, the second and third rows are already total occupancy — and the third row is the dangerous one, because a 20 ms wait looks harmless in a single trace. The pool was deliberately sized with **no slack**: extra workers buy nothing when every stage returns in microseconds, and they cost stack memory and context switches. That efficiency is exactly what leaves no headroom to absorb waiting. ## What it looks like from outside 1. Latency climbs on routes that share no dependency, no data store and no code, because the only thing they share is the pool. 2. Processor usage stays low while latency and timeouts rise — the signature of waiting rather than computing. 3. The delay appears **before** the stage runs: the time from accepting a request to starting its first stage grows, while the stage's own duration is unchanged. 4. The degradation correlates with traffic to one feature, which is the thread to pull. ## The fix, and the fix that isn't - Move that one stage onto a **separate worker set sized for waiting** — many more workers than processors, because they spend their lives parked — and leave the rest of the chain on the small non-blocking pool. - Give the isolated region a **bound**, so a dependency that slows down cannot grow it without limit, and a **timeout**, so a call that never returns eventually releases its worker. - Enlarging the small pool is not a fix. It raises how many concurrent waits are needed to park it, and traffic supplies them; you also pay the switching cost the small pool existed to avoid. - Retrying into the same parked pool makes it worse: each retry adds occupancy to the resource that is already the constraint. ## Why this bites late The stage is usually correct and usually fast. At low traffic, occupancy stays under the worker count and nothing is visible; the first symptom often arrives with a traffic peak, a dependency that got slower, or a feature flag switched on for everyone. Because the damage lands on *other* routes, the investigation typically starts in the wrong place — on the endpoint that got slow, not the endpoint that made everything slow.

  • What is the fix once you have identified the stage that waits?
    Run that one stage on a separate worker set sized for waiting — far more workers than processors, since they sit parked — and leave the rest of the chain on the small non-blocking pool. Bound the isolated region and give the call a timeout, so a dependency that slows cannot grow occupancy without limit.
  • Why does enlarging the small non-blocking pool not solve it?
    It moves the threshold rather than removing the wait. With twice the workers you need twice the concurrent calls to park them, and traffic supplies that. You also pay the cost the small pool was chosen to avoid: more stacks in memory and more context switching for stages that never needed to wait.
  • If the call waits only 20 milliseconds, is it still a hazard?
    Yes, once the arrival rate is high enough. By Little's Law the average occupancy is rate times wait, so 200 calls a second at 20 ms holds about four workers — a whole small pool. Short waits are more dangerous in one way: they surface as a latency floor on every route rather than as obvious timeouts.

Four tellers serve every queue in a bank. If one teller sits on hold with a supplier for ten minutes, he is occupied without serving anyone, and every queue slows — not just the one that needed the supplier.

saying these in an interview costs you the question

  • Thinks only the endpoint that blocks gets slower.
  • Assumes the runtime detects the wait and reuses the worker elsewhere.
  • Reaches for a bigger pool instead of moving the blocking stage.
  • Reads near-idle processor usage as proof nothing is stuck.
  • Believes a stage cannot block because it sits inside a pipeline.
  • Adds retries, which pile more occupancy onto the parked pool.
open as a page

Why does a tenant id stashed in per-worker ambient storage come back empty after a pipeline stage hops workers?

level: middleimportance: must knowfreq 64%

basics

~20 s

Ambient storage is attached to the worker, not to the request: a read resolves against whatever worker is executing right now. Once a stage runs on a different worker, nothing ever wrote that worker's slot.

open as a page

Why would you deliberately run a pipeline stage on the calling thread, with no execution-context hop at all?

level: middleimportance: must knowfreq 52%

basics

~20 s

A hop is not free: it enqueues the value, wakes another worker and adds latency for every element. For a short non-blocking stage, running inline on whichever worker delivered the value costs less than moving the work somewhere else.

open as a page

In a reactive media pipeline, how do you choose the execution context for a decode stage versus a stage that waits on remote storage?

level: middleimportance: must knowfreq 64%

basics

~20 s

Match the worker to where a stage spends its time. Compute-bound decoding belongs on a small fixed pool sized near the processor count; a stage that spends its wall-clock time waiting on remote storage belongs on an elastic pool that can grow.

open as a page

Why does assigning a stream pipeline to a pool of many workers still leave its elements processed one at a time?

level: middleimportance: must knowfreq 58%

basics

~20 s

A reactive sequence is serial by contract: values reach each stage one at a time, in order, however many workers the execution context owns. That choice decides which worker runs a stage, never how many elements run at once.

open as a page

In a stream chain, why does the switch that relocates the producing work apply no matter where it is placed?

level: middleimportance: must knowfreq 62%

basics

~20 s

Subscription travels upstream from the subscriber to the source, so that switch is reached wherever it sits and hands the rest of the upward walk to its worker. The source therefore starts on that worker.

open as a page

After fanning a scoring pipeline out across workers, why do results arrive out of order, and how do you restore it?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Results converge in completion order, not source order, because the work for each record takes a different amount of time. Restore order by carrying each record's position through the fan-out and re-emitting by position, or by making the write-back address records by key.

open as a page

A desktop window freezes while a chain reads a large file; adding a switch that moves only the stages after it changed nothing — why?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The freeze is the read itself, and the read runs on whatever worker the subscription reached the source from — the window worker here. A switch that moves only later stages never touches it; it just adds a hop.

open as a page

A pipeline has no obvious network waits, yet its workers still park; where does the blocking hide?

level: middleimportance: should knowfreq 46%

basics

~20 s

Blocking hides in anything that returns a finished value: a synchronous data-access driver, a contended lock, first-use setup such as opening a connection or resolving a name, a file or device read, and a synchronous logging or metrics write.

open as a page

In a context carried with a subscription, where must a value be written for a given stage to read it?

level: middleimportance: should knowfreq 46%

basics

~20 s

A subscription context is assembled while the subscription travels from the subscriber back toward the source, so a write serves the stages it passes on that journey: those between it and the source, not those after it in the written chain.

open as a page

A nightly scoring batch must use every core, so what are the two ways a serial reactive sequence gets real parallelism?

level: middleimportance: should knowfreq 45%

basics

~20 s

Two routes exist. Turn each element into an inner operation carrying its own work and keep several subscribed at once under a concurrency bound; or split the sequence into a fixed number of worker-pinned tracks and rejoin them afterwards.

open as a page

Every route on an asynchronous service is slow; what evidence separates parked workers from one slow dependency?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Ask which routes degraded. A slow dependency hurts only its callers; parked workers hurt routes with nothing in common, alongside near-idle processors, growing time before a stage starts, and stack samples finding the pool waiting.

open as a page

Why does a stage that reads a missing tenant id from the subscription context and scopes nothing usually pass its tests?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Absence is handled as a default rather than a failure: the scoping clause is simply left out, the query succeeds, and the result is bigger rather than broken. Tests that always populate the context never execute that branch.

open as a page

An elastic execution context's worker ceiling is set arbitrarily high; why does that look like a thread leak in production?

level: seniorimportance: should knowfreq 44%

basics

~20 s

An elastic context opens a worker whenever every existing one is busy, and a waiting call holds its worker for the whole wait. When downstream latency rises, the worker count climbs with it, producing the same monotonic graph a leak produces.

open as a page

Which work in an ingest pipeline justifies a single-worker execution context, and what does serialising it cost?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Work that must happen one at a time earns it: assigning sequence numbers, or driving a resource only one thread may touch. A single serial worker gives exclusion and queue order without a lock, at the cost of a one-worker ceiling and head-of-line delay.

open as a page

What breaks when a pipeline fans every element out into its own inner operation with no limit on how many run at once?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Unbounded fan-out subscribes to an inner operation per element, so in-flight work scales with the input. Memory for outstanding work, open connections and the load the dependency behind the stage sees all grow with the record count, and the failure surfaces far from the pipeline.

open as a page

How would you arrange switches so a chain decodes a large file off the window worker but paints progress on it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Relocate the producing work once, so the read and decode run in the background, then place one forward-only boundary just before the painting stage so values arrive on the window worker. Each boundary opens a new segment.

open as a page

Several pipeline stages in your service offload blocking calls to one waiting-sized pool; when do you split that pool per dependency?

level: principalimportance: should knowfreq 38%

basics

~20 s

Split when one dependency's slowdown must not stall the others. One shared offload pool reproduces the original failure a level down; separate pools contain it but reserve capacity. A permit per dependency over one pool buys most of that isolation.

open as a page

When you set one standard across many reactive services, what criteria decide whether a value may ride in the subscription context instead of an explicit stage parameter?

level: principalimportance: should knowfreq 34%

basics

~10 s

A value qualifies when it is ambient, constant for the whole subscription, small and immutable, and read by stages the value is not about. Anything a stage operates on directly stays an explicit parameter.

open as a page

A platform runs many ingest pipelines in one process; how do you decide how many execution contexts it exposes and which workloads may use each?

level: principalimportance: should knowfreq 33%

basics

~20 s

Decide by blast radius rather than convenience. A small shared inventory keeps the process's total worker count predictable and understandable; a dedicated context per workload buys isolation, paid for in threads, idle memory and harder global reasoning.

open as a page

Your team proposes parallelising every batch pipeline by default, so how do you decide whether that pays?

level: principalimportance: should knowfreq 34%

basics

~20 s

Decide from measurement, not policy. Find where the wall clock actually goes, bound the achievable gain by the share of work that stays serial, then weigh that ceiling against permanent costs: a bound to tune, order to restore, and failures that no longer reproduce.

open as a page

Should a published stream chain pin the worker its producing work runs on, or leave that choice to each caller?

level: principalimportance: should knowfreq 36%

basics

~20 s

Pin only where the work is unsuited to any caller's worker, because a pin nearest the source silently overrides every caller's own switch. Otherwise leave the choice out and publish which worker values are delivered on.

open as a page

If one chain carries two switches that each relocate the producing work, which one decides where the source runs?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

The one the subscription reaches last, which is the one nearest the source. The others only relocate the remaining part of the upward walk; they never change where the source produces or where values are delivered.

open as a page