skip to content

A stream-based import validates ten thousand address rows and row twelve signals a failure — what happens to rows thirteen onward?

level: middleimportance: must knowfreq 72%

answer

  1. an ending, not a skipped element
  2. one ending per sequence
  3. the remainder is never requested
  4. subscription released, source stops
  5. loop catch resumes, stream does not

basics

~10 s

Nothing processes them. A failure signal is terminal: it ends the sequence at row twelve and releases the subscription, so the source is never asked for rows thirteen onward and nothing downstream sees them.

solid answer

~40 s

A stream ends exactly once, and a failure is one of the ways it ends. It is not a per-element event that the pipeline steps over. When validation of row twelve raises a failure, that failure travels downstream as the sequence's ending: the subscriber gets it instead of the remaining values, the subscription is released, and the source stops being asked for rows. Rows thirteen to ten thousand are never read. This is the opposite of a loop with a per-iteration catch, where handling a bad row resumes the same iteration and the remaining rows still run. Rows already written before row twelve stay written — a terminal failure stops the flow, it does not roll anything back.

code

pseudocode · 8 lines
pseudocode
accepted = 0
for each row in rows:
    try:
        validate(row)            // raises a rejection on row 12
        accepted = accepted + 1
    catch rejection:
        record(rejection)        // the loop moves on to row 13
return accepted                  // all 10000 rows were visited

go deeper

for a junior

Remember the one-line fact: a failure ends the whole sequence rather than skipping the bad element, so the values that would have come after never arrive.

for a middle

Explain the mechanics: the failure travels downstream as the sequence's ending, the subscription is released, and the source is never asked for the remaining input. Contrast it with a loop whose per-iteration catch resumes.

for a senior

Show that you have debugged the consequence: a nightly job that imports an unpredictable fraction of a file, stops at whatever the first bad record happens to be, and leaves the destination partially applied because terminality is not rollback.

for a principal

The judgment is which conditions deserve to end a run at all. Terminality is a claim that the remainder has no value; treat it as a design decision with a cost, not as the default you inherited.

## A failure is an ending, not a skip A stream delivers values over time and ends **exactly once**. There are two kinds of ending: a normal completion, meaning the source had nothing more to give, and a failure, meaning the sequence stopped because something went wrong. The load-bearing word in both is *ending*. After either one, that subscriber receives nothing further from that sequence, and the subscription connecting it to the source is released. That is the whole of the surprise in the scenario. The failure raised while validating row twelve did not stay local to row twelve. It **became the ending of the sequence**. Rows thirteen through ten thousand are not skipped, not deferred, not queued for later. They are never requested from the source at all. ## What the ending actually costs Four consequences follow, and a weak answer names only the first: - **Nothing downstream sees another value.** Each stage between the failing stage and the subscriber passes the ending along and stops handling values. - **The subscription is released.** The source is no longer asked for anything, so whatever it held open for this subscriber — a reader over the file, a connection, a timer — is dropped rather than left producing into a dead path. - **Stages keyed to a *normal* ending never fire.** An accumulator that writes its batch when the sequence completes normally receives a failure ending instead, so it writes nothing. - **Work already performed is not undone.** The eleven rows written before the failure are still written. Terminality stops the flow; it is not a transaction boundary. ## Why the loop intuition misleads Most engineers arrive with the per-iteration catch as their mental model of "handling a bad row", and that model is exactly wrong here. | | loop with a per-iteration catch | stream with a failure signal | |---|---|---| | scope of the failure | the single iteration it was raised in | the whole sequence | | after it is handled | the next iteration runs | nothing further is delivered | | remaining input | still visited | never requested from the source | | default outcome | the run finishes with a bad row recorded | the run stops at the bad row | | where handling lives | inside the body, per element | on the pipeline, as the ending's destination | The row that raised the failure is also *not* the interesting part. The interesting part is the **remainder**: the loop's remainder survives, the stream's remainder is cancelled. ## The defect this produces at scale The shape is stable under testing and fails in production for a statistical reason. A fixture of twenty clean rows never raises anything, so the pipeline finishes and the shape looks correct. A real file of ten thousand address records almost certainly contains at least one row that validation rejects; once it does, the run stops there, and each nightly attempt stops at the first bad row again. The observable symptom is a job that reports a failure and imports a fraction of the file, with no obvious relationship between the fraction and the fault — it is simply wherever the first bad row happened to sit. The fix is not to make failures gentler. It is to decide which outcomes were ever failures. A malformed postcode is an **expected per-record outcome** and belongs in the sequence as data, so the run keeps flowing; a destination that refuses every write is a condition under which continuing has no value, and terminality is the right answer there. Terminality is a claim about the *remainder*: it says the rest of this run is not worth producing. ## What an interviewer is listening for - That you say **ending**, not "the element is dropped". - That you mention the **remainder** — the rows never requested — rather than only the failing row. - That you do not describe the failure as recoverable by default; continuing past it is something you arrange deliberately, and if you arrange nothing, the sequence is over. - That you separate **terminality from rollback**: already-emitted work stands, which is why a partially applied import is a real state a rerun has to cope with. Said compactly: in a stream, failure is the ending you did not plan; in a loop, failure is an element you did.

  • The import ended with a failure at row twelve — what happened to the eleven rows that were already written?
    They stay written. Each value was delivered and handled as it passed, and the terminal failure does not reverse that. The run leaves the destination in a partially applied state, which is why a rerun has to be safe against rows that are already there.
  • Does it matter whether the failure came from the source itself or from a validation stage in the middle?
    Not for what the subscriber sees: either way the sequence ends and no further value arrives. It matters for diagnosis and for what the remainder would have contained — a failing source means there was nothing more to read anyway, while a mid-pipeline failure abandoned input that was perfectly readable.
  • Can a sequence deliver a failure for one element and still end normally later?
    No. A sequence has one ending. The failure is that ending, so there is no later normal completion to wait for. Any stage that only acts on a normal ending simply never acts on this run.

A failure signal is the phone call dropping, not one word being misheard. Everything you still meant to say is never delivered, and the line is gone.

saying these in an interview costs you the question

  • Thinks the failing row is skipped and the sequence continues
  • Believes the failure ends only the stage that raised it
  • Expects a normal completion to arrive after the failure
  • Assumes rows already written are rolled back by the failure
  • Says the source keeps producing into the released subscription
  • Describes streams as recovering from element failures by default