Why is a row count either side of a cleanup step a better check than scrolling its output?
answer
- structure, not values
- a claim a machine can fail
- before and after one step
- assert the intended difference
basics
~20 sA row count is a claim a machine can fail on, while scrolling only samples the rows you happen to see. Counts either side of one step catch rows silently lost or multiplied anywhere in the table.
solid answer
~50 sScrolling the output inspects a handful of rows you did not choose, and the defects that matter here are structural rather than visible: rows removed by a condition nothing absent can satisfy, rows multiplied by a match, records the reader refused before the data ever became a table. All of those move the row count and none of them make a visible cell look wrong. So I write the claim down beside the step — `after == before - removed`, with the intended removals counted first — as an inline assertion: a claim in the same file, next to the step it is about, that fails the run where it is false. Stating the intended change rather than plain equality is what keeps the check alive. Where the table is already materialised the count is bookkeeping the object maintains; on a deferred pipeline asking for it is real work, so budget it.
code
pseudocode · 9 linesbefore = row_count(orders)
cancelled = row_count(where(orders, status == "cancelled"))
orders = drop_where(orders, status == "cancelled")
expected = before - cancelled
if row_count(orders) != expected:
raise_error("cancellation step: expected " + expected +
" rows, got " + row_count(orders))go deeper
Remember the habit: record the row count before a step and compare it after, in the file, as something that fails. Structural damage is what this catches; a wrong value needs a different kind of check.
Explain why plain equality is the wrong claim for a step that removes rows, and show the computed expectation instead: count the intended removals first, then assert the difference. Say why a printed number is not a check.
Show where the claim belongs so it names the step that broke rather than the end of the run, and be honest that the count is bookkeeping on a materialised table and real work on a deferred one.
The angle is what evidence a result must carry before anyone acts on it, and what the standing rule costs: which steps must carry a claim, how many passes over the data the team is willing to spend on evidence, and who notices when a claim is deleted.
## What the claim actually says A row count assertion is one sentence about structure: **this step received N rows and returned M**, where the relationship between N and M is something you decided before the step ran. It says nothing about any value in any cell. The **row and column counts** — how many rows and columns a table holds at this point in the file — are the cheapest claim there is to make about a result, and on a materialised in-memory table the object already knows them. The second half of the idea matters as much as the first. An **inline assertion** is a claim written in the same file, beside the step it is about, that **fails the run at the point where it is false**. A claim that lives anywhere else is worth much less, and one that nothing acts on is worth nothing. ## Why a structural claim finds more than reading values - **The screen is a sample you did not choose.** A viewer shows the head and the tail of a table. A defect concentrated in one region, one source system or one day is invisible there, and a defect spread thinly is invisible everywhere. - **Most damage in this kind of work is structural.** Rows multiplied because a match found several partners; rows quietly removed because a condition cannot be satisfied by an absent value; records the reader skipped before the data became a table at all. Every one of those moves the row count. None of them makes an individual visible cell look wrong. - **A value has no baseline.** Reading `4173.22` tells you nothing unless you already know what it should be. Reading `1,000,000 -> 1,412,033` tells you immediately that a step you believed removed rows added four hundred thousand. - **One number compares mechanically.** A machine can compare two integers on every run, for ever, without getting bored on the fiftieth run. None of this makes the count a check on correctness of values. It is a check on structure, and structure is where the cheap wins are. ## State the intended change, not equality The common mistake is to assert that the count is unchanged across a step whose whole purpose is to change it. That claim fails on every correct run, so somebody deletes it, and with it goes the only evidence the step ever produced. The pattern that survives is: 1. Measure what the step is supposed to do — count the rows that match the removal condition **first**. 2. Run the step. 3. Assert the computed expectation: `after == before - removed`. Now an unintended loss of eleven rows fails, and an intended removal of nine hundred thousand passes. The same shape works for a step that adds rows: count what should arrive, then assert the sum. ## Beside the step, not at the end of the run | where the claim lives | what a failure tells you | what you still have to do | |---|---|---| | beside the step | this step, this input, this expectation | open the step that was named | | once at the end of the run | one of eleven steps is wrong | re-run, inserting counts, to bisect it | | nowhere — you read the output | something looks off, or nothing does | start from the top with no evidence | Failing early has a second benefit beyond diagnosis: the run stops before a wrong result is written somewhere another person will read. ## Printing is not asserting Printing counts to the log is what most people do, and it is not a check. It requires a human to read the number, remember the previous number and decide. On the day the defect appears, that human is running the file for the fortieth time and is looking at something else. Turn the print into a comparison that raises. ## What it costs, and what it does not tell you The cost is not the same everywhere, and asserting otherwise is how the habit gets a bad name. On a **materialised in-memory table** the count is bookkeeping the object already maintains, so the claim is effectively free. On a **deferred pipeline** — one that builds a plan and computes nothing until a result is demanded — asking for a count is asking the plan to run, and asking again for the next count can run it again. On a reader that walks a file in chunks, a count can mean another pass over the input. Where the claims are not free, collect the ones you need in as few passes as you can. And be honest about the limit: a failed count says **rows moved**, not which rows and not why. That is the start of the next investigation, not the end of this one — but it is an enormously better starting point than a number at the bottom of a file that somebody thinks looks low.
- The step lives inside a function called from three places. Where does the claim go?Inside the function, beside the step, so every caller is covered by one claim. Express it in terms of the function's own inputs — the count it received, minus the removals it measured — rather than a literal taken from one caller's data, which would be wrong for the other two.
- What do you do with a count assertion that fires every month on legitimate input changes?Treat it as a wrong claim, not a noisy one. A literal row count is almost never the right expectation; a relationship usually is — after equals before minus what was measured for removal, or after equals the number of distinct entities. If the relationship genuinely varies, you have not yet understood what the step guarantees.
A load of parcels changes hands at four depots. Counting them at each handover takes seconds and tells you which depot lost one; opening parcels at the far end tells you only that something is missing. The count is near-free where the load is already stacked and counted, and costs a re-weigh where it is not.
saying these in an interview costs you the question
- Says reading the first rows of the output is enough
- Prints counts to the log instead of failing the run
- Collects every check at the end of the run
- Asserts equality across a step meant to remove rows
- Deletes the assertion because it keeps firing
- Treats a matching count as proof the values are right