Where a producing operator pushes each record downstream as it is made, what replaces the bucket a consumer would otherwise collect?
answer
- routing unchanged, storage removed
- a send buffer, not a file
- received, not requested
- nothing kept means nothing re-read
basics
~20 sNothing durable. The routing rule still picks one destination per record, but the bucket is only a send buffer flushed over a channel to a consumer that is already running. There is no stored share to request, and none to request twice.
solid answer
~50 sThe routing half of the exchange is unchanged: each record's key still names exactly one destination, so equal keys still meet. What changes is the other half. Instead of a share written to local disk and requested afterwards, the producer holds a small per-destination buffer and flushes it across a channel to a consumer that is running at the same moment. The consumer does not *collect*; it **receives**. Three consequences follow. Both steps occupy capacity simultaneously, for the life of the exchange, rather than the producing side finishing and being released. Nothing can be read a second time, so a consumer that dies cannot re-gather its input and the job has to return to a recorded point instead. And the two sides are coupled: if the consumer cannot keep up, its buffers fill and the resistance travels back to the producer.
go deeper
Recall that some runtimes never store the exchange at all: the record goes straight to the worker that will handle its key, so there is no share sitting anywhere to be collected.
Explain what stays and what goes. The routing rule and the per-destination separation stay; the written share, the request and the possibility of a second read all go.
Show you can name the regime behind any claim you make, and say what each implies for capacity held, for re-reading and for how a failure is repaired.
The trade is architectural: a pushed hand-off buys latency and spends coupling plus simultaneous capacity, and it makes a replayable source non-negotiable — a commitment that reaches well past the one job.
## The half that does not change This regime is often described as if it were a different exchange. It is not. The record's key is still put through the **routing rule** — the rule that names exactly one destination per key — and equal keys still converge on one consumer, which is the entire point of a redistribution. Buckets still exist as a *concept*: the producer still separates its outgoing records by destination. What differs is where a bucket lives and who moves it. ## What a bucket is when nothing is stored - It is a **small send buffer per destination**, held in the producing operator's memory. - It is flushed when it fills, or after a short interval so a trickle of records does not sit waiting forever. - It is addressed to a **specific running instance** of the consuming operator, over a channel that stays open for the life of the exchange. Designs vary in whether each producer-consumer pair gets its own network connection or many logical channels are multiplexed over fewer physical ones. - It is **gone once acknowledged onward**. There is no offset to come back to and no file to reopen. The fan is the same arithmetic as before — one channel per producer-consumer pair, so the product of the two widths — but it is paid as concurrent channels and buffers rather than as transfers requested later. ## What the consuming side does instead of collecting It receives. It does not decide when its input arrives, it cannot ask for a particular producer's contribution, and it cannot ask twice. That produces four properties an interviewer is listening for: 1. **Both sides are alive at once.** The producing and consuming steps hold capacity simultaneously for as long as the exchange lasts, rather than the producing side completing and its workers being handed back. 2. **No second read.** A consumer that fails cannot re-gather what it was sent. The job returns to a recorded point and re-derives the input from the source — a recovery mechanism of its own, and a different subject from the exchange. 3. **A lost machine takes something different.** There is no written share to lose. What is lost is the records in flight and whatever the consumer had accumulated, which is why recovery here is about restoring a consistent picture rather than re-making a bucket. 4. **The sides are coupled.** A consumer that cannot keep up stops draining its channel, the producer's buffers for that destination fill, and the resistance propagates upstream — backpressure, the mechanism by which a slow step throttles the ones feeding it. Reading and acting on that signal is an operating subject in its own right; what matters here is that a materialised exchange does not have it, because a slow collector simply asks later. ## Three regimes side by side | | shares written down first | pushed as produced | run as repeated small finite jobs | |---|---|---|---| | where a bucket lives | the producing machine's local disks | a send buffer in the producer's memory | local disks, once per little job | | how the consumer gets it | it requests its numbered share | it is sent, on an open channel | it requests, within that little job | | can a share be read twice | while those bytes survive | no | within that little job only | | both steps alive at once | not required | required | required within each little job | | what a machine loss takes | the written shares on it | in-flight records and accumulated consumer state | that little job's shares | The third column is the one people forget, and it is a large part of this market: a continuous computation can be executed as a rapid succession of small finite jobs, each of which performs the written-down version at small scale. It therefore inherits materialised exchange mechanics *and* the ability to re-read a share within a little job, while looking from the outside like a continuous pipeline. ## How to answer without asserting one engine's model The failure mode in an interview is to state one regime as *the* way exchanges work. Both statements below are true only of their own regime: - *a redistribution writes to local disk and is collected afterwards* — the batch lineage's design; - *records cross the network as they are produced, with nothing materialised* — the record-at-a-time lineage's design. Name the regime with the claim: **where the shares are written down**, collection can outlive the producer and a share can be read again; **where records are pushed as they are produced**, there is nothing to re-collect, both sides must be alive, and the price of that is a coupling between them that has to be operated.
- If nothing is stored, how does such a job recover a consumer that dies?Not by re-collecting, because there is no share to collect. It goes back to a recorded point — a previously captured picture of the job's progress and accumulated state — and re-derives what came after it from the source. That makes a replayable source a hard requirement here, whereas a materialised exchange can often just be re-read.
- Why does a slow consumer affect the producer in this regime but not in a materialised one?Because the producer's only outlet is a buffer bound to that consumer. When the consumer stops draining, the buffer fills and the producer blocks, and the slowdown travels upstream. Where shares are written down, the producer's outlet is local disk, so a slow collector delays itself and leaves the producer alone.
- Does pushing records as they are made remove the all-to-all fan?No. Every consumer can still receive from every producer, so the pair count is still the product of the two widths. It is paid differently: as channels and buffers held open concurrently rather than as transfers requested afterwards, which turns a transfer-count problem into a connection-and-memory problem.
saying these in an interview costs you the question
- Thinks pushing records removes the need for a routing rule
- Says every engine materialises a redistribution before anything is collected
- Believes a consumer can re-request records it has already received
- Assumes the producing side can finish and release its capacity
- Treats a slow consumer as a purely downstream problem here