A nightly address import writes validated rows in one batch at the end, and a malformed row at position 9,998 signalled a failure — why were the earlier good rows lost?
answer
- the flush had one branch only
- a failure is a different ending
- nothing was rolled back, nothing was written
- bounded chunks make work durable
- the open chunk is still lost
basics
~10 sThe write was keyed to the sequence ending normally. A failure is a different ending, so the stage that had accumulated 9,997 validated rows was torn down with the subscription and never wrote anything.
solid answer
~40 sAccumulate-then-write makes the whole run's output conditional on a normal ending. A terminal failure is not that ending: it ends the sequence a different way, so the branch that flushes the accumulated rows is never reached and everything held in memory goes away with the released subscription. One malformed record at position 9,998 therefore destroys 9,997 rows of successful work. The shape hides in testing because fixtures are clean, and it gets worse as files grow, since a bigger file is more likely to contain at least one bad record. The structural fixes are to make the bad record non-terminal by carrying rejections as data, and to write in bounded chunks so work already done is durable before any later failure arrives.
code
pseudocode · 11 linesaccumulated = []
function onEnding(kind):
if kind == NORMAL:
writeAll(accumulated) // the only branch that writes
subscribe(
map(source(rows), validate),
function (row): accumulated.append(row),
function (failure): onEnding(FAILURE), // writes nothing
function (): onEnding(NORMAL))go deeper
The point to carry away: if the write only happens when the sequence ends normally, a failure means no write at all, however much valid work was accumulated first.
Explain the conditional: the flush has one branch, keyed to a normal ending, and a failure is the other ending. Be careful to say nothing was rolled back, because nothing had been written.
Diagnose and repair it: identify that the job's whole output depends on an ending no long run can promise, move to bounded chunks, state that the open chunk is still lost, and raise rerun safety before anyone asks.
The trade-off is blast radius against write cost and atomicity. Decide where all-or-nothing is a real requirement and where it is an accident, and make the safe shape the default others inherit.
## Two endings, one flush A sequence ends once, in one of two ways: **normally**, or with a **failure**. The accumulate-then-write shape attaches the entire side effect of the job to the first of those. Read it as a conditional and the defect is obvious: *if the sequence ends normally, write everything*. There is no branch for the other ending, so when a failure arrives at position 9,998 the accumulated rows are simply discarded along with the subscription that held them. Notice what this is **not**. It is not a rollback: nothing undid the work, because the work had not happened yet. The rows only ever existed in the accumulator. That distinction matters, because the two failure modes need opposite fixes — a rollback problem needs transaction boundaries, while this needs the write to stop depending on an ending it cannot guarantee. ## Why it survives testing - **Fixtures are clean.** A test file of twenty well-formed rows never raises, so the normal ending always fires and the flush always runs. - **The defect is conditional on a failure landing mid-run**, which is precisely the case the happy-path test omits. - **It scales against you.** If any single record has a small chance of being rejected, the chance that a file of ten thousand contains at least one approaches certainty. The shape is most fragile exactly where it is most used — big nightly loads. - **The symptom is confusing.** The destination is untouched, so operators report "the import did nothing" rather than "the import stopped at 9,998", and the investigation starts in the wrong place. ## What is durable when the failure lands | write shape | durable after a failure at 9,998 | memory held | write cost | |---|---|---|---| | accumulate all, write on normal ending | nothing | the whole file | one large write | | bounded chunks written as they fill | every completed chunk; the open chunk is lost | one chunk | one write per chunk | | write each row as it passes | every row before the failure | one row | one write per row | None of these is universally right. Accumulating gives the largest single write and the cleanest all-or-nothing story, and it is defensible when the destination can accept the whole set atomically and you genuinely want all-or-nothing. It is indefensible when it was chosen by accident, which is the usual case. The chunked middle is where most bulk imports belong: work already done survives, memory is bounded, and the write cost stays reasonable. Be precise about what it costs — the rows in the **open** chunk when the failure arrives are still lost, so the guarantee is "every completed chunk", not "everything before the failure". ## The rerun question Any shape that makes partial work durable creates a second obligation: a rerun must be safe over records already written. That usually means writing by a stable identity so that re-processing a record already present changes nothing, rather than appending blindly and producing duplicates on the second attempt. Engineers who have run these jobs mention this unprompted, because the first production incident after switching to chunked writes is almost always duplicates. And the deeper fix stands above all of these: if a malformed postcode is an expected per-record outcome, it should never have ended the sequence at all. Carry the rejection as a value, and position 9,998 costs you one record on a rejects report instead of a whole night's run. Chunked writing limits the blast radius of a *genuine* stop condition; making expected rejections non-terminal removes most of the stop conditions in the first place. ## What an interviewer wants to hear 1. The mechanism: the flush was conditional on a normal ending, and the failure was a different ending. 2. The correction of vocabulary: nothing was rolled back, because nothing had been written. 3. The structural fix with its cost stated honestly, including the open chunk. 4. The rerun consequence, unprompted. An answer that only says "add error handling" has not identified the defect: the defect is that the entire output of the job was made conditional on an ending no long-running pipeline can promise.
- Did the terminal failure undo rows that had already been written to the destination?No. A terminal failure ends the sequence; it is not a transaction boundary and reverses nothing. In this shape the question does not even arise, because no row had been written yet — which is exactly why the loss is total rather than partial.
- After switching to chunked writes, what new obligation does the job take on?Reruns must be safe over records already written. Writing by a stable record identity so a repeat write is a no-op keeps a partial run harmless; appending blindly turns every retry into duplicates, which is the usual first incident after the change.
- Why did this shape pass every test before it reached production?Test fixtures are clean, so the sequence always ended normally and the flush always ran. The defect needs a failure to land mid-run, and the odds of that grow with file size, so it first appears on the largest real load.
- Is accumulate-then-write ever the right choice here?Yes, when all-or-nothing is a genuine requirement and the destination can accept the whole set as one unit — then losing everything on failure is the intended semantics. The defect is choosing the shape by accident and discovering its semantics during an incident.
saying these in an interview costs you the question
- Expects a completion-keyed flush to run after a failure ending
- Says the terminal failure rolled back the writes
- Assumes accumulating the whole file is safe because tests passed
- Thinks a larger accumulator would have prevented the loss
- Claims chunked writing preserves everything before the failure
- Treats the stream's ending as a transaction commit