A nightly meter batch has grown to millions of rows; when should a whole-batch all-or-nothing traversal stop being its shape?
answer
- pick the atomic unit first
- one bad row discards millions
- values held until the last row
- writes done are not undone
- chunk, then traverse inside the chunk
basics
~20 sWhen the batch stops being the unit the business accepts or rejects. At millions of rows, one bad row discarding every good one, memory held to the end, effects already performed and a late verdict argue for a smaller unit.
solid answer
~50 sA whole-batch traversal encodes one claim: the batch is the unit of success. That claim is worth keeping while it is true - a statement that must load whole, a file that is meaningless in part. At millions of rows it usually stops being true, and four things push back. Every produced value is held until the last row, so peak memory scales with the batch. One bad row discards millions of good ones. If the step wrote as it ran, those writes happened and a returned failure has no power to undo them. And a single verdict after a six-hour run is feedback nobody can act on in time. The move is not to abandon the traversal but to shrink what it spans: pick the unit the business can accept or reject on its own, traverse inside that unit, and build the quarantine and re-run machinery the smaller unit now needs.
go deeper
Understand that a whole-batch traversal is all-or-nothing, and that one bad row therefore throws away every good row in the file.
Name what grows with the batch - values held until the end, the blast radius of one bad row - and how traversing smaller units changes each.
Reason about the effects already performed when a late row fails, and what safe re-running and quarantining actually cost to build.
Choose the unit of atomicity from what the business can accept or reject, then defend the stop rule and re-run semantics that a smaller unit now demands.
## What the whole-batch shape is promising A traversal over the entire nightly file says something specific: **this batch succeeds or fails as one thing**. Every row's result is folded into a single wrapped collection, and nothing is visible to the caller until the last row has been stepped. That is a real and sometimes correct guarantee. The question is not whether it is elegant, but whether it is the guarantee the business actually wants at this size. ## Four things that change as the row count grows 1. **Peak memory.** The accumulated values are alive from the first row to the last, because the result cannot be handed over until the traversal finishes. Memory therefore scales with the batch, not with the work per row. 2. **Blast radius.** One malformed row discards the other million. For a statement that must be internally consistent, that is the point. For meter readings, where each row stands alone, it is an outage caused by a typo. 3. **Effects already performed.** If each step wrote as it ran, a failure at the last row leaves every earlier write in place. The failure value the caller receives carries a description, not a rollback; undoing requires a transaction spanning the run or explicit compensation, and neither comes from the traversal's shape. 4. **Feedback latency.** A six-hour run that returns one verdict at the end tells the operator nothing at hour one. The information existed at minute three; the shape withheld it. ## Choosing the atomic unit | unit traversed | one bad row costs | feedback | fits when | |---|---|---|---| | the whole batch | the whole night | once, at the end | the file is one indivisible document | | a chunk of `k` rows | one chunk | per chunk | rows group naturally, order within a chunk matters | | a single row, partitioned | one row | per row | rows are independent and a quarantine exists | The decision is made in one order, and getting the order right is most of the judgment: 1. **Name the unit the business can accept or reject on its own.** That is a product question, not a technical one - ask whether importing 999,999 of a million readings is better or worse than importing none. 2. **Make that unit the span of one traversal.** All-or-nothing then applies exactly where it is meaningful, and nowhere else. 3. **Build what the smaller unit now requires** - because a smaller unit moves work outward, it does not remove it. ## What you have to build when you shrink the unit - **Somewhere for the rejects to go.** A quarantine with the row, the reason and enough context to fix it, or the operator has traded one useless failure for a million invisible ones. - **A re-run that is safe.** Once part of a batch has landed, re-running the rest must not double-count what already did; that usually means an idempotency key per row or per chunk. - **A report that aggregates.** Per-chunk results are only usable if something sums them into "accepted, rejected, why". - **A stop rule.** Independent chunks will happily import a corrupt file one chunk at a time. A threshold - abort the run when the rejection rate crosses some fraction - restores the protection that all-or-nothing gave for free. - **A definition of "done".** With one traversal, done was obvious. With many, the run needs its own completion and failure semantics. ## When whole-batch all-or-nothing stays right - **Partial acceptance is incorrect**, not merely awkward - the rows are parts of one document and half of it is wrong rather than incomplete. - **The batch is bounded and small enough** that memory and latency never enter the conversation. - **The step is pure**, so a failure leaves nothing behind and re-running the whole batch costs only time. - **The downstream consumer cannot express partial state**, and would treat a half-loaded batch as a full one. The mistake to avoid in both directions: keeping whole-batch atomicity out of habit after the rows became independent records, and abandoning it out of scale-anxiety while the rows are still parts of one indivisible thing. The traversal is not the thing being chosen here - the unit it spans is.
- What makes a chunk the right atomic unit rather than an arbitrary size?A chunk is right when it is something the business can accept or reject on its own - a meter, a day, a site. An arbitrary size chosen for memory alone gives partial acceptance that nobody can reason about, and re-runs that land on different boundaries.
- If each step writes as it runs, what does a failure on the last row leave behind?Every earlier write, in place and unannounced. The returned failure value describes what went wrong; it has no transactional power. Undoing needs a transaction spanning the run, or compensating writes you have written yourself and tested.
- What protection do you lose by splitting into independently accepted chunks, and how do you get it back?You lose the guarantee that a systematically corrupt file is rejected as a whole - chunks will import it piece by piece. A stop rule restores it: abort the run once the rejection rate crosses a threshold, so widespread corruption still halts the import.
saying these in an interview costs you the question
- Says a larger batch only needs more memory and nothing else changes
- Assumes a failed traversal undoes the writes its steps performed
- Keeps whole-batch atomicity without asking what the unit of meaning is
- Thinks chunking removes the need to report rejected rows
- Treats one verdict after a six-hour run as adequate feedback
- Splits on an arbitrary size chosen only to fit memory