A per-row validator rejects an address for a malformed postcode — why carry that rejection as a value rather than a failure signal?
answer
- many values, one ending
- expected outcomes are data
- the ending condemns the remainder
- widen the element to carry an outcome
- rejections still need a destination
basics
~20 sA rejected row is an expected outcome, not a broken stream. Signalled as a failure it ends the run at that row; carried as a value in the sequence it flows on, so the remaining rows are still processed and both counts are reported.
solid answer
~40 sA stream has two channels with different meanings. Values are the things the run is about, and there are many of them. The failure channel is used once and it is the ending, so it can only say "this run is over" — it cannot say "this record is bad, the next one is fine". A malformed postcode is a normal, anticipated outcome of validating untrusted input, so expressing it as the ending is a category error: one bad record in ten thousand kills the run. Make the element type carry the outcome instead — an accepted row or a tagged rejection with a reason — and the sequence keeps flowing, rejections can be routed and counted, and the failure channel is left for conditions under which continuing genuinely has no value.
code
pseudocode · 18 linesACCEPTED = "accepted"
REJECTED = "rejected"
function validateRow(row):
if postcodeWellFormed(row.postcode):
return { kind: ACCEPTED, row: row }
return { kind: REJECTED, id: row.id, why: "postcode format" }
function handleOutcome(outcome):
if outcome.kind == ACCEPTED:
writeToDestination(outcome.row)
else:
writeToRejects(outcome)
subscribe(
map(source(rows), validateRow), // never raises for a bad row
handleOutcome,
handleFailure) // left for conditions that stop the rungo deeper
Learn the split: things that happen to one record are data the sequence carries, and the failure channel is for the run being over. Confusing the two is what kills long imports.
Explain why the failure channel cannot express a per-record outcome: it fires once and is the ending, so using it condemns every record after the bad one. Then describe widening the element to carry an accepted or rejected outcome.
Show the operational half: rejections need a real destination, a count on the run report and a rate threshold, or you have traded a noisy abort for silent data loss that nobody notices for weeks.
The interesting call is where the threshold sits and who owns the rejects. Decide what rejection rate means the input contract changed, and make that the condition that genuinely stops a run.
## Two channels, two different meanings A stream carries **many values** and exactly **one ending**. That asymmetry is the whole argument. - The value channel is per element, repeated, and ordinary. It is where the run's subject matter lives. - The failure channel fires once and *is* the ending. Using it makes a claim about the **remainder**: that the rest of this run is not worth producing. So the question is never "is a malformed postcode an error?" — of course it is, in plain English. The question is: **does one malformed postcode mean the remaining 9,999 rows should not be processed?** Almost always no. If the answer is no, the failure channel is the wrong channel, because the failure channel cannot express anything narrower than "stop". ## What "expected" means here A useful test is whether the outcome is part of the job's normal operating range: 1. **Would you write a count for it on the run report?** Nine hundred rejected postcodes is a number an operator wants; it is data, not an incident. 2. **Does it depend on one record, or on the environment?** Per-record properties of untrusted input are expected; a refused destination or exhausted credentials is not. 3. **Would the next record probably succeed?** If yes, the run still has value and terminality is a lie about the remainder. ## The shape that carries an outcome Instead of raising, the validation stage **returns** something for every row: either the accepted row or a tagged rejection carrying an identifier and a reason. The element type widens from "a row" to "a per-row outcome", and every downstream stage matches on which it got. Nothing about the sequence's ending is involved, so the run reaches its normal ending and reports both counts. | | rejection signalled as a failure | rejection carried as a value | |---|---|---| | effect on the run | ends at the first bad record | run continues to its normal ending | | records processed | those before the bad one | all of them | | what the report can say | one abort, one reason | accepted count, rejected count, reasons | | how many can be reported | one — the ending happens once | as many as occur | | where the handling lives | on the pipeline's ending | in the per-element handling | | what the failure channel is left for | nothing — it is spent | genuine stop conditions | ## What it costs Carrying outcomes as values is not free, and an interviewer will push on the cost: - **Every downstream stage must handle both cases.** A stage written as if it only receives accepted rows will write rejections into the destination or silently drop them. - **A rejection with no destination is data loss.** If rejections ride along and nobody routes them anywhere, the run reports a clean ending while records disappear. That is worse than the abort it replaced, because the abort was at least loud. - **Volume changes meaning.** A handful of rejections is normal; forty per cent rejected usually means the source changed shape or you are importing the wrong file, and the run should stop. Terminality does not disappear — it moves from "any one bad record" to a deliberate threshold. ## Where the failure channel still belongs Reserve it for conditions where the remainder genuinely has no value: the destination refusing every write, credentials rejected, the source truncated so that nothing further can be read. These have the property a terminal signal asserts — continuing produces nothing useful. Keeping the channel for exactly those cases also restores its diagnostic value: when a run ends with a failure, that now means something specific rather than "row 9,998 had a typo". ## How to say this in an interview Name the asymmetry (many values, one ending), then the consequence (the failure channel cannot describe a single record without condemning the rest), then the shape (widen the element to carry the outcome), then the cost (both branches must be handled and rejections need a real destination and a threshold). That sequence is the complete answer, and it is short.
- What must every stage downstream of the validator do once rejections travel as values?Match on which outcome it received. A stage written as if only accepted rows arrive will either push rejections into the destination or quietly drop them, turning a loud abort into silent corruption. The widened element type is what forces each stage to declare its intent.
- Which outcomes in this import should still use the failure channel?The ones where the remainder has no value: the destination refusing every write, credentials rejected, or the source truncated mid-file. Each condemns the rest of the run, which is exactly the claim a terminal signal makes.
- How does carrying rejections change what the run can report?The sequence reaches its normal ending, so the report can carry an accepted count, a rejected count and the reasons alongside each other. Signalled as a failure, the run can only report the one that ended it, because the ending happens once.
saying these in an interview costs you the question
- Treats every invalid input record as an exceptional condition
- Thinks the failure channel can describe one record without ending the run
- Assumes a rejected record is skipped by the pipeline automatically
- Carries rejections as values but routes them nowhere
- Believes carrying outcomes removes any need for a failure channel
- Ignores rejection volume, so a wrong source file imports silently