skip to content

Why would you deliberately run a pipeline stage on the calling thread, with no execution-context hop at all?

level: middleimportance: must knowfreq 52%

answer

  1. the fourth option is no worker
  2. a hop is a handoff with a price
  3. short non-blocking stages need no move
  4. you inherit the delivering worker
  5. per-element enqueue and wake-up cost

basics

~20 s

A hop is not free: it enqueues the value, wakes another worker and adds latency for every element. For a short non-blocking stage, running inline on whichever worker delivered the value costs less than moving the work somewhere else.

solid answer

~50 s

Caller-thread execution is a real fourth choice alongside the fixed, elastic and serial contexts, and it is also what happens when nothing asks for a move: the stage runs inline on whichever worker delivered the value. You choose it when the stage is short and never waits, because a handoff has a fixed per-element cost — enqueue the item, wake a worker, and run the code on a processor whose caches know nothing about this data. A chain that hops three times for stages of a few microseconds each pays that cost three times per value, and the scheduling can easily exceed the work. What the choice makes you responsible for is the other side of the same coin: the worker that delivered the value now runs your code, so the stage must stay short and must not wait on anything.

code

pseudocode · 10 lines
pseudocode
// a few microseconds of arithmetic per element
inline_chain = source()
    .map(item -> item.byte_count * 8)
    .subscribe(report)

// identical work, now preceded by a handoff paid for every element
hopped_chain = source()
    .on_context(WAITING)
    .map(item -> item.byte_count * 8)
    .subscribe(report)

go deeper

for a junior

Remember that a stage runs on whichever worker delivered the value unless something moves it, and that not every stage needs a context of its own.

for a middle

Explain the per-element cost of a handoff — enqueue, wake-up, cold caches — and why a short non-blocking stage is better left where it is.

for a senior

Judge from measurement: weigh a stage's own duration against the handoff it would add, and hunt for chains that hop repeatedly for no stated reason.

for a principal

Set the house default — whether teams place stages explicitly or only where a workload demands it — and account for what that costs in review effort and latency budget.

## What caller-thread execution means Every stage runs somewhere. If nothing in the chain asks for a different worker, the stage runs **inline on whichever worker delivered the value to it** — the worker that produced the item, or the one that signalled it from a timer or a network completion. That is caller-thread execution, and it is both the default for a stage nobody moved and a deliberate option worth naming: *this work stays here; do not hand it over*. It is worth being precise about a common confusion. Caller-thread execution does not mean the stage runs on the thread that assembled the chain, and it does not mean the thread that subscribed. Assembling a chain does no work; the worker is decided at delivery time, and for a source that emits later it may be a worker the caller has never seen. ## What a hop costs A hop is a handoff between two workers, and a handoff is machinery: - the value is placed into the target context's queue, which is itself concurrent state; - a worker on that context is woken, or an already-running one picks the item up on its next turn; - the work resumes on a processor whose caches hold none of the data the previous stage just touched; - latency is added per element, not per pipeline — a chain that hops pays it for every value that flows. None of those costs is large. All of them are fixed, and a stage of a few microseconds is smaller than the fixed cost of moving it. | | Stage kept inline | Stage moved to another context | |---|---|---| | Cost per element | the stage's own work only | the stage plus one enqueue and one wake-up | | Cache behaviour | continues on warm data | resumes cold on another processor | | Who is occupied | the delivering worker | a worker of the target context | | Suited to | short, non-blocking steps | compute that deserves its own pool, or work that waits | ## When the no-hop choice is the right one 1. **The stage is cheap and pure.** Mapping a field, computing a size, filtering on a predicate. The work is far smaller than the handoff. 2. **The delivering worker is already the right kind.** If values arrive on a compute pool and the stage is compute, moving it to another compute context buys a handoff and nothing else. 3. **Latency matters more than isolation.** Each avoided hop removes a queueing point from the per-element path. ## What staying inline makes you responsible for The worker that delivered the value is now running your code, and it cannot do anything else — including deliver the next value — until your stage returns. That is fine for arithmetic and fatal for a stage that waits on a remote call or takes a contended lock, because the worker it is occupying belongs to whatever context produced the value, chosen by somebody else's placement decision. The rule that falls out of this is simple to state and easy to check in review: **stay inline only where the stage cannot wait and cannot run long.** There is a second, quieter responsibility. Because the delivering worker varies with the source, code that stays inline must not assume anything about which worker it is on. A stage that would only be correct on one particular worker is not a candidate for inline execution; it wants a context that guarantees the worker. ## Recognising hops nobody needed The production signature of needless hops is latency that scales with the number of elements rather than with the amount of work: each stage is individually fast, the end-to-end time is not, and a large share of processor time is spent in scheduling and queue handling rather than in the pipeline's own functions. Chains assembled by several people over time collect these, because each author adds the move that made their own stage safe and nobody removes the one that has become redundant. The corrective is to place stages on purpose rather than by habit: a move should have a reason you can name — this work computes and wants the compute pool, this work waits and wants the elastic one, this work must be serialised. Where no such reason exists, the cheapest and most honest placement is none at all.

  • Does caller-thread execution mean the stage always runs on the thread that subscribed?
    No. It means no move was requested, so the stage runs on whichever worker delivers the value. For a source that emits immediately on subscription that may indeed be the subscribing thread, but for a source that emits later it will be a timer worker, a network worker, or a worker of the context an earlier stage was placed on. Caller-thread execution names the absence of a hop, not a particular thread.
  • How do several needless hops in one chain show up in production?
    As latency that grows with element count rather than with work. Every hop adds an enqueue, a wake-up and a cold restart per value, so three redundant moves pay that three times for each element. The signature is a pipeline whose individual stages are fast while end-to-end latency and processor time are dominated by scheduling and queue handling.

saying these in an interview costs you the question

  • Every stage should be placed on some named context
  • Moving work to another worker is free
  • Caller-thread execution always means the subscriber's own thread
  • A hop makes a slow stage faster by itself
  • Short stages should hop so they cannot block anything