A consumer's outstanding demand for a ledger stream is 4 — may the producer deliver those four values concurrently?
answer
- how many, not how many at once
- one handover at a time
- serial, but possibly different workers
- parallelism lives behind the handover
- the subscription is the unit of parallelism
basics
~20 sNo. Demand is permission to send a number of values, never permission to send them at once. Delivery to one subscriber is serialised: each value is handed over only after the previous handover has returned.
solid answer
~40 sA count bounds **how many**, not **how many at a time**. Signals to a single subscriber are serialised, so four outstanding demand means four handovers one after another, each starting only once the previous one returns. That is what lets the consumer's handler keep ordinary, unguarded state and rely on the order of the ledger rows it is committing. Serial is not the same as single-threaded, though: consecutive handovers may happen on different workers, and the protocol only promises they do not overlap and that each one sees the effects of the one before it. Concurrency stays perfectly legal *behind* the handover — fetching pages, decoding, computing — and the real unit of parallelism is the subscription, not the request count.
code
pseudocode · 14 lines// WRONG: reads demand of 4 as four deliveries at once
on requestSignal(4):
for each row in nextFourRows():
spawnWorker: emit(row) // four handovers race
// RIGHT: four deliveries, strictly one after another
on requestSignal(4):
outstanding = outstanding + 4
if not draining: // single owner of the loop
draining = true
while outstanding > 0 and hasNextRow():
emit(nextRow()) // returns before the next one starts
outstanding = outstanding - 1
draining = falsego deeper
Remember the two words that differ: a count says how many values, not how many at the same time. Values arrive one after another, in order, for a given subscription.
Explain the serialisation rule and what it buys the consumer: no locking around handover-local state, preserved ordering, and a happens-before relationship that carries state across a worker change.
Demonstrate the diagnosis: throughput that does not move when the request count is raised, a consumer that spawns work inside the handover and loses its own bound, and a fan-out done properly with separate subscriptions.
Treat it as an interface guarantee you are buying for every operator author in the codebase. Concurrent delivery under one count would push thread-safety into every consumer; keeping handover serial concentrates that cost in one drain loop.
## What a demand of four buys, and what it does not A demand-driven stream gives the consumer one lever: a count. When it says *I can take four more*, it has said something precise and something narrow. Precise: the producer may hand over at most four values before it must stop. Narrow: it has said nothing about **when**, on **which worker**, or **how many at a time**. The rule the protocol fixes is that signals to a single subscriber are **serialised**. Each value is handed over in its own call, and the next handover does not begin until the previous one has returned. Four outstanding demand is four handovers in a row, not four handovers at once. Return to the ledger export. The report writer asks for four rows because it can validate and commit four. If the producer answered by invoking the writer's handler from four workers simultaneously, the writer would have to guard every field it touches, its commit batch could interleave with itself, and the row ordering that makes the export a ledger at all would be gone. The count would have bounded memory and bought nothing else. ## Serial does not mean single-threaded The common over-correction is to read *serialised* as *always the same worker*. It does not say that. Consecutive handovers may well happen on different workers — a stage that shifts delivery onto another pool does exactly that — and the guarantee is only that they never overlap and that each handover sees the effects of the one before it. For the consumer's handler that means, concretely: - it does **not** need a lock around state it touches only inside the handover; - that state **is** safely visible from one value to the next, because the non-overlap comes with an ordering guarantee, not merely with mutual exclusion; - it must **not** assume worker-affine storage, or anything else keyed to the identity of the executing worker, survives from one value to the next. ## Where concurrency is still legal | Concern | Bounded by the request count? | Serialised? | |---|---|---| | values delivered to this subscriber | yes | — | | handovers to this subscriber | yes | yes | | work the producer does to create a value | no | no | | deliveries to a second, separate subscription | counted separately | independent of the first | | work the consumer starts and does not wait for | no | no | The producer is free to fetch three ledger pages in parallel, decode them on a pool, and do anything it likes to prepare values. What it may not do is let two resulting handovers overlap. Implementations achieve this with a drain loop that has a single owner at a time: whichever worker finds the loop free takes ownership, emits while demand lasts, and releases it; any other worker merely adds its contribution and returns. The mirror image matters on the consumer's side. If the consumer hands each value off to its own worker inside the handover and returns immediately, the handover no longer paces it — and the count it requested no longer reflects the work actually in flight. That is the subtle trap: **demand bounds deliveries, not the outstanding work a consumer has spawned behind them**. A consumer that fans out internally has to bound that fan-out itself, and should only replenish demand when the spawned work finishes rather than when the handover returns. ## How you actually get parallelism Not by requesting more. Raising the count from four to four hundred changes how much the producer may send ahead, not how many values are processed simultaneously; the handovers still happen one after another. Parallelism comes from splitting the work across **separate subscriptions**: several pipelines, each with its own outstanding count and its own serialised handover, with their results merged back through one more serialised handover downstream. The unit of parallelism is the subscription; the request count is only the unit of flow control within one. ## Why the protocol is drawn this way It is a deliberate division of labour. Serial handover makes the consumer's job writable by an ordinary engineer: no locks, no re-entrancy analysis, order preserved. The count makes the producer's job bounded: it always knows exactly how far ahead it may run. Had the protocol allowed concurrent delivery under one count, every consumer — including every operator in the middle of a pipeline — would have had to be written thread-safe against itself, and the ordering guarantees that most operators depend on could not be stated at all. The practical tell in an interview is the follow-through: a candidate who says *demand of four means four in parallel* will usually also propose raising demand to fix a slow stage, and will be surprised when throughput does not move.
- Where may the producer still use concurrency without breaking the guarantee?Everywhere behind the handover. It can fetch pages in parallel, decode on a pool, and prepare values however it likes, provided every finished value passes through one drain loop that emits them one at a time. Only the handover itself is serialised, not the work that produced it.
- Two consumers subscribe to the same source — must deliveries to them be serialised with each other?No. Serialisation is per subscription. Each subscription carries its own outstanding count and its own ordered handover, and the two may progress at completely different speeds and on different workers without any relationship between them.
- If a consumer hands each value to a worker and returns, is it still protected by its own demand?Not really. The handover returns immediately, so the count now measures deliveries rather than work in flight, and the consumer can accumulate unbounded spawned work while looking well-behaved. It must bound the fan-out itself and replenish demand when the spawned work completes.
saying these in an interview costs you the question
- Reads a demand of N as permission for N parallel deliveries
- Believes serial delivery means every value arrives on the same worker
- Thinks raising the request count is how you add parallelism to a slow stage
- Assumes serial handover forces the producer to prepare values sequentially too
- Expects worker-affine state to survive between two consecutive values
- Thinks demand still bounds work after the consumer spawns it and returns