Your batch import reports only the first bad meter row; what must change for the traversal to report every bad row?
answer
- one report or forty round trips
- stop short-circuiting, run every step
- failures need a defined merge
- only valid when rows are independent
- successes still dropped on any failure
basics
~20 sTwo things must change: the failure channel needs a way to merge two failures into one, typically a failure holding a list, and the traversal must run every row's step instead of stopping at the first. Collecting needs independent rows.
solid answer
~50 sReporting only the first bad row is the signature of **fail-fast** sequencing: each step is taken only when the previous one succeeded, so the traversal returns at the first failure. To report all of them you change two things. First, the failure side must be able to hold more than one failure, with a defined, associative way to merge two - a failure carrying a list of failures is the usual shape. Second, the traversal must run the step on **every** row and combine the results rather than returning at the first bad one. That is only meaningful when the rows are independent: if row `k`'s step consumes row `k-1`'s output, there is nothing to run after a failure. The cost is that you always pay every step call, and the successful values are still discarded when any row fails.
code
pseudocode · 10 linesfunction traverseCollecting(rows, step):
values = empty list
failures = empty list
for each row in rows: // every row is visited
r = step(row)
if r is failure: failures.append(failure of r)
else: values.append(value of r)
if failures is not empty:
return failure(failures) // merged; values dropped
return success(values)go deeper
Recall that there are two behaviours - stop at the first bad row, or check them all - and that the second needs somewhere to put more than one failure.
Name both changes precisely: a failure type with an associative merge, and a traversal that runs every step instead of returning early.
Bring the preconditions and the costs: independence of the rows, every step paid, and the fact that the good values are still discarded when any row fails.
Decide the standard across pipelines - typically a cheap collecting validation pass before an expensive fail-fast one - and define what operators are entitled to see.
## What "first bad row only" actually means A traversal that reports one bad row is sequencing its steps: it takes the next step only if the previous one produced a value, so the first failure ends the batch and is returned as the whole result. This is the right default - it is cheap, it needs nothing of the failure type beyond existing, and for dependent work it is the only correct behaviour. It is also useless to the operator staring at a file with forty malformed rows, who will otherwise fix one, rerun for twenty minutes, and find the next. ## The two changes required 1. **Give the failure channel a merge.** One failure plus one failure must produce one failure. The everyday shape is a failure that carries a list, so merging is concatenation; any associative combine will do - that monoid-like structure is the whole requirement, and it must have a sensible result for "no failures at all". 2. **Stop short-circuiting.** The traversal must apply the step to every row, keep the values on one side and the failures on the other, and at the end return the merged failures if there are any, or the collected values if there are none. Nothing about the *rows* changed; what changed is how the per-row results are combined. ## The precondition people skip: independence Collecting every failure is only meaningful when each row's step can run without the previous row's value. Validating a meter reading is independent; resolving row `k` against an identifier minted while importing row `k-1` is not. When the steps are genuinely dependent, there is nothing to run after the first failure, and demanding "all the errors" is asking for errors that could not have been produced. Say this before you promise an operator a complete report. ## Fail-fast against collect-all | | fail-fast | collect-all | |---|---|---| | step calls when 3 of 200 rows fail | up to the first failing row | 200 | | failures reported | the first one | all three, merged | | successful values when any row fails | dropped | dropped | | needs a merge on the failure type | no | yes | | requires independent rows | no | yes | | fits | expensive or dependent steps | validating a whole submitted batch | ## What collecting still does not give you - **It is not a partition.** The result is either every value or the merged failures. If you want the 197 good rows imported and the 3 bad ones quarantined, that is a different operation, chosen deliberately, with its own contract about what happens to the rejects. - **It does not undo anything.** If the step wrote as it ran, the writes from the successful rows have happened. A merged failure value has no transactional power. - **"Every failure" means one per row.** A row that violates three of its own checks still contributes a single failure unless the per-row check itself accumulates. Batch-level collecting and per-row collecting are separate decisions. - **It is not free.** Every step runs, including on rows after the point where a fail-fast traversal would have stopped. When the step is a network call or an expensive computation, collecting can cost far more than the report is worth. ## Choosing per operation - **Collect** when a human is on the other end and a round trip is expensive: a submitted batch, a validation pass whose only job is to say what is wrong. - **Fail fast** when steps are dependent, when they are expensive, when a failure makes the rest meaningless, or when the caller is another system that will retry the whole batch anyway. - **Decide once, per pipeline, and write it down.** The most common production shape is a cheap collecting validation pass over the whole batch first, then a fail-fast processing pass over rows that already passed - so the expensive work never runs on a batch that was going to be rejected. ## Saying it precisely in an interview The answer that lands names the change in the **combining**, not in the rows or the error messages. A fail-fast traversal sequences its steps, so the second step exists only because the first succeeded and there is exactly one failure to report. A collecting traversal treats the rows as independent results to be merged, which is why it needs a merge on the failure side and why it cannot be applied to dependent work. Everything else - the list of failures, the friendlier report, the extra step calls - follows from that one difference. Candidates who describe it as "catching the error and continuing the loop" have the behaviour roughly right and the requirement on the failure type missing, which is where the design work actually is.
- What has to be true of the per-row steps before you can promise every failure?They must be independent - no step may need a value produced by an earlier one. With dependent steps there is nothing to run once a step fails, so the later failures do not exist to be collected, and promising a complete report is promising something the shape cannot deliver.
- What does a collecting traversal return when every row is valid?Exactly what the fail-fast one returns: a success holding one value per row, in order. The accumulation shows up only on the failure side, which is why switching styles does not change the happy path for callers.
- Why not collect everywhere, given how much friendlier the report is?Because collecting always pays every step call, including on rows a fail-fast traversal would never have reached. When the step is expensive, touches another system, or is meaningless once an earlier row failed, that cost buys a report nobody needed.
- Where does a collecting pass and a fail-fast pass sit in the same pipeline?A cheap collecting validation over the whole batch runs first and reports everything wrong at once; a fail-fast pass then performs the expensive work only on a batch that already passed. The expensive steps never run on a batch that was going to be rejected.
A proofreader who stops at the first typo hands the page back with one mark on it; one who reads to the end hands it back with every mark. Both refuse the page - only the second saves you the next round trip.
saying these in an interview costs you the question
- Thinks you can accumulate failures while still stopping at the first
- Says the collecting version returns the good rows alongside the failures
- Assumes any failure type can be merged without defining the merge
- Believes collecting is free because the rows are read anyway
- Claims fail-fast and collect-all differ only in the message wording
- Promises a full report over steps that depend on earlier rows