Compare three ways for a consumer to receive a sequence of items from a producer: a blocking pull loop where the consumer asks for the next item, unrestrained push where the producer invokes a consumer callback, and demand-driven push where the consumer sends batched requests upstream. What does each cost, and when does the third win?
answer
- Pull: free flow control, one parked thread each
- Push: thread-cheap, no rate control, unbounded queue
- Demand push: permits batched, bounded, nobody parked
- Suspending pull: loop shape, no parked thread
- Local in-memory sequence: just use a loop
basics
~20 sBlocking pull gives free rate control but parks a thread per stream. Unrestrained push parks nothing but hands rate control to the producer, so surplus items pile up. Demand-driven push combines them: batched permits bound memory, push delivery keeps threads free. It wins with many concurrent, remote, rate-mismatched sources.
solid answer
~60 s**Blocking pull (iterator style):** the consumer asks for the next item and waits. Rate control is inherent — nothing arrives unless you ask — but each active sequence occupies a thread while it waits. With thousands of slow remote sources that is thousands of parked threads: memory for stacks and scheduler pressure. **Unrestrained push (callback style):** no thread waits, and items arrive with minimum latency, but the consumer cannot say slow down. Any speed mismatch becomes an unbounded queue, a blocked thread, or dropped data. Composition is nesting and the completion and error conventions are ad hoc. **Demand-driven push:** the consumer sends batched permits upstream, the producer delivers push-style but never exceeds outstanding demand. You get pull's bounded memory with push's thread economy, amortized because permits are batched rather than per item. It wins when there are many concurrent, remote or asynchronous sources with uneven rates and a hard memory bound. It loses on a single local in-memory sequence, where a plain loop is simpler and faster, and on a single value, where a future is the right shape. Suspending pull is a fourth option worth naming: it reads like a loop but does not park a thread.
code
text · 10 lines# 1. blocking pull - thread waits inside next()
while (item = source.next()) != END: handle(item)
# 2. unrestrained push - nothing waits, nothing throttles
source.onItem(item -> handle(item)) # surplus -> queue / block / drop
# 3. demand-driven push
sub = source.subscribe(consumer)
sub.request(64) # permits, batched
# on each 32 handled: sub.request(32) # replenish, never exceed windowgo deeper
Name the three models and the headline trade: pull waits and controls the rate, push does not wait and does not control it, demand-driven push tries to get both.
Explain the parked-thread cost of pull and the unbounded-queue cost of push, and that batching permits is what makes demand cheap.
Give the workload conditions that justify the third model — high concurrency, remote sources, rate mismatch, bounded memory — and be equally clear about where a plain loop wins.
Bring in cheap user-mode threads and suspending iteration as alternatives that change the arithmetic, and judge the complexity against team debuggability and ecosystem support.
## The three transfer models, on their merits ### 1. Blocking pull The consumer drives: *give me the next item*, then wait until it exists. - **Rate control is free.** Nothing is produced or delivered until the consumer asks, so there is no surplus, ever, and no buffering policy to design. - **The code is the easiest to read.** A loop with a body. The stack tells you where you are, exceptions propagate normally, and a debugger steps through it. - **The cost is a parked thread per concurrent sequence.** Waiting on a remote source means an idle thread holding a stack. Fine for a handful of sequences; ruinous for tens of thousands of slow connections, which is precisely the workload async models were built for. - **A non-blocking variant exists** — pull returning a future per item — but that makes one signal round trip per item, which is chatty and latency-bound over a network. ### 2. Unrestrained push The producer drives: whenever an item exists, invoke the consumer's callback. - **No thread waits**, so one thread can serve many sources. This is the model of event-driven I/O. - **Minimum latency:** the item is handed over the instant it exists, with no request round trip. - **The consumer has no rate control.** If the consumer is slower, the surplus goes into a queue (unbounded means eventual memory exhaustion), into a blocked thread (which forfeits the model's whole benefit), or on the floor (data loss). - **Weak protocol.** Completion and failure signalling are per-API conventions, composition nests, and nothing guarantees a callback is invoked exactly once. ### 3. Demand-driven push The consumer sends permits upstream (*request n*); the producer delivers push-style but must never exceed the outstanding permits. - **Bounded in flight by construction.** The surplus is never created, so no buffering policy is forced on you at the transport level. - **No parked threads.** When demand runs out, the producer simply stops emitting; nobody waits. - **Amortized signalling.** Permits are batched — request a window, replenish when part of it is consumed — so you do not pay a round trip per item as chatty async pull does. - **Costs.** A more complex protocol with real rules (serial delivery, at most one terminal signal, cumulative demand), pipelines that are harder to debug because the causal chain is not the stack, and correctness burdens on operator authors. Requesting an unbounded amount silently opts out of the whole mechanism. ### 4. The hybrid worth naming: suspending pull Coroutine-style asynchronous iteration keeps pull semantics — the consumer asks for the next item — but suspends instead of blocking, so no thread is parked. It reads like the loop of model 1 with the thread economics of model 3, and back-pressure is implicit because the source only advances when the consumer asks. Its weakness is that each item is still a request-and-resume, so a very high-rate source pays per-item coordination that batched demand amortizes; and it has no standard cross-library protocol for composing independently written stages. ## When demand-driven push wins - **Many concurrent sequences**, each mostly idle — thousands of connections, subscriptions or feeds. Thread-per-sequence pull cannot get there. - **Remote or otherwise asynchronous producers**, where a per-item request round trip is a latency tax. - **Uneven or unpredictable rates** across pipeline stages, where you need the slow stage to throttle the source automatically instead of by hand-tuned queue sizes. - **A hard requirement for bounded memory under overload** — the system must degrade in latency, not fall over. - **Multi-stage pipelines assembled from independently written stages**, where a shared protocol is what makes composition safe. ## When it loses - **A single local in-memory sequence.** A plain loop is faster and enormously easier to debug; the protocol buys nothing when there is no I/O and no rate mismatch. - **A single result.** That is a future. - **Sources whose rate you do not control** — sensors, user input, external feeds. They cannot honour demand, so you still have to choose what happens to the excess; the demand protocol alone does not answer that. - **Teams and codebases where readability and debuggability dominate**, and where sequential suspending code delivers most of the benefit at a fraction of the cognitive cost. ## How to answer Structure it as two axes — who delivers, who controls volume — then place all three models on it and say the third exists to take rate control from the producer without giving anyone a parked thread. Finish with the honest boundary: it is worth its complexity where concurrency is high, sources are remote and memory must stay bounded, and not otherwise.
- If blocking pull already gives perfect flow control, why not just use more threads?Because the cost is per concurrent sequence, not per unit of work. Ten thousand mostly idle remote sources means ten thousand stacks in memory plus scheduler and context-switch overhead, all to wait. Cheap user-mode threads change the arithmetic and make the pull loop viable at much higher concurrency, but where the runtime does not provide them, demand-driven push gets the same flow control with a handful of carrier threads.
- Where does an event loop sit in this comparison?An event loop is the execution substrate rather than a transfer model: it is how push-style delivery is dispatched without a thread per source. Demand-driven push typically runs on top of one, adding the volume-control channel that a bare event loop does not have. The pairing matters operationally, because a slow consumer callback occupies the loop thread and delays every other source it serves.
saying these in an interview costs you the question
- Presenting demand-driven push as universally superior, including for a local in-memory sequence
- Claiming blocking pull has no flow control, when consumer-driven pull is the strongest form of it
- Ignoring that pull's real cost is a parked thread per concurrent sequence rather than raw throughput
- Assuming demand signalling can throttle any source, including ones with an external clock
- Requesting one item at a time and concluding that demand-driven push is inherently chatty