skip to content

Two workers in a hand-written stream source push values to one subscriber concurrently — why does the contract forbid that?

level: middleimportance: must knowfreq 58%

answer

  1. one at a time, not one worker
  2. overlap is the forbidden thing
  3. the next signal sees the last
  4. downstream stages hold unsynchronised state
  5. one gate at the emission boundary

basics

~20 s

Signals must be delivered one at a time, with each one seeing what the previous one wrote. Every stage downstream is written on that promise and keeps its accumulated state unsynchronised, so overlapping deliveries corrupt it.

solid answer

~50 s

The contract says signals reach a subscriber **serially**: at most one in flight, and an ordering edge between consecutive ones so the second sees what the first wrote. That promise is what lets every stage between the source and the final consumer keep counters, buffers and folds in plain unsynchronised fields. Two workers pushing at once turn all of that into a data race — lost or duplicated elements, a corrupted accumulator, values arriving out of order, and a terminal signal racing a value so something lands after the ending. Note what the rule does *not* forbid: the source may use as many workers as it likes, and consecutive signals may come from different ones. What is forbidden is **overlap**. The fix is a single gate at the emission boundary that every producing worker passes through.

code

pseudocode · 10 lines
pseudocode
// forbidden: two workers inside the consumer at once
worker A:  consumer.value(a1)
worker B:  consumer.value(b1)

// conforming: one gate, delivery inside it, worker may vary
function emit(signal):
    lock(gate):
        if terminated: return
        if signal is terminal: terminated = true
        deliver(signal) to consumer

go deeper

for a junior

Remember the shape of the promise: a subscriber is handed one signal at a time. Parallel work inside the source is fine; two simultaneous handovers are not.

for a middle

Explain both halves — non-overlap and the ordering edge — and why the promise is what lets intermediate stages keep counters and buffers without synchronisation.

for a senior

Diagnose from symptoms: duplicated or lost elements under load, a corrupted running total, a value landing after the ending, all clean in single-element tests. Then point at the missing gate.

for a principal

Treat it as where the cost of correctness is placed: one gate paid once per signal at the source, versus synchronisation spread across every stage that any team ever writes.

## What serial delivery actually promises The rule has two halves, and candidates usually recall only the first. 1. **No overlap.** At any instant at most one signal is being delivered to a given subscriber. A second worker that arrives while the first is inside the consumer must wait. 2. **An ordering edge between consecutive signals.** The worker delivering signal *n+1* must see everything the worker delivering signal *n* wrote. Non-overlap in wall-clock time is not enough on its own: without a handoff that establishes that edge — a lock, a queue, or an equivalent ordering guarantee — the second worker may observe stale state. What the rule does **not** say is equally important. It does not say the work runs on one worker, and it does not say consecutive signals come from the same one. A source may fan its work across a pool and still conform, provided deliveries are funnelled so they never overlap. ## Why the contract puts the obligation on the source Serialisation has to happen somewhere. The contract chooses the source because that is the cheapest and safest place for it. | where synchronisation could live | what it costs | |---|---| | in every stage between source and consumer | every stage pays on every element, and one stage that forgets reintroduces the race for the whole chain | | in the final consumer only | the stages in the middle still see overlap, so their own state is still racing | | at the source's emission boundary | paid once per signal, in one place, and the rest of the chain stays lock-free | The third row is the whole design. Because delivery is serial, an author writing a stage that counts elements, batches them, or folds them into a running total can use ordinary fields with no synchronisation at all, and be correct. That is not an optimisation detail; it is why small composable stages are writable by ordinary engineers. ## What breaks when two workers push at once The failures are the classic data-race set, made worse by being invisible in light testing: - **Lost or duplicated elements** — two workers read-modify-write the same buffer index or counter. - **Corrupted accumulated state** — a partially built batch is emitted, or a fold reads a half-updated value. - **Reordering** — two values overtake one another, so a stage that assumed monotonic ordering draws the wrong conclusion. - **A signal after the ending** — a value delivery overlaps a terminal delivery, and the value lands after the sequence has closed, which breaks the grammar rule as a side effect of breaking this one. - **Load-dependent symptoms** — a single-element test never overlaps, so the defect ships and appears first under production concurrency. ## Where the gate goes in a hand-written source One emission boundary, and every producing worker goes through it. The same gate is the natural home for the terminated flag, which means one small critical section enforces both rules at once: it rejects anything after the terminal and guarantees the non-overlap. The critical section must cover the delivery itself, not merely the decision to deliver — releasing the gate and *then* calling the consumer reintroduces exactly the overlap you were preventing. An equivalent shape replaces the lock with a queue: producing workers enqueue, and one drain loop delivers. The queue handoff supplies the ordering edge, and a drain loop that is guaranteed to run one at a time supplies the non-overlap. Which shape is better depends on whether you would rather block a producing worker or buffer for it — and the buffering choice belongs to the flow-control subject, not to this one. ## Distinguishing this from two things it is often confused with - **It is not a rule about where work runs.** Fetching, decoding and computing may all happen in parallel. Only the act of handing a signal to the subscriber is serialised. - **It is not satisfied by a thread-safe final consumer.** The source cannot know what sits between it and that consumer, and the intervening stages were written on the serial promise. A source that delivers concurrently because *its* consumer happens to cope is a source that breaks the moment anyone puts a stage in front of it. In review, the question to ask is narrow and answerable: *can two code paths in this source be inside the consumer at the same moment?* If you cannot answer no by pointing at a single gate, the answer is yes.

  • If two deliveries never overlap in time, is any synchronisation still needed between them?
    Yes. Something must establish that the later delivery sees what the earlier one wrote — a lock, a queue handoff, or an equivalent ordering guarantee. Non-overlap alone says nothing about visibility, so a consumer's unsynchronised counter can still be read stale by the next worker to deliver.
  • Does serial delivery mean the source cannot use a pool of workers?
    No. Producing work may be as parallel as you like; only the handing over of signals is serialised. Consecutive signals may even be delivered by different workers, provided they pass through one gate so they never overlap and the ordering edge is preserved.

saying these in an interview costs you the question

  • Thinks concurrent delivery is fine if the final consumer is thread-safe
  • Believes every signal must come from one fixed worker
  • Adds locks in every stage instead of serialising at the source
  • Confuses serial delivery with doing all the work single-threaded
  • Dismisses overlapping delivery as a rare race not worth fixing
  • Releases the gate before calling the consumer, keeping the overlap