skip to content

Trusting the Result

The number on screen came from a program that may no longer exist, and the worst failures here raise nothing at all. Proving a result is separate work from producing it.

on this pageshow

questions

21

Why is a row count either side of a cleanup step a better check than scrolling its output?

level: juniorimportance: must knowfreq 70%

answer

  1. structure, not values
  2. a claim a machine can fail
  3. before and after one step
  4. assert the intended difference

basics

~20 s

A row count is a claim a machine can fail on, while scrolling only samples the rows you happen to see. Counts either side of one step catch rows silently lost or multiplied anywhere in the table.

solid answer

~50 s

Scrolling the output inspects a handful of rows you did not choose, and the defects that matter here are structural rather than visible: rows removed by a condition nothing absent can satisfy, rows multiplied by a match, records the reader refused before the data ever became a table. All of those move the row count and none of them make a visible cell look wrong. So I write the claim down beside the step — `after == before - removed`, with the intended removals counted first — as an inline assertion: a claim in the same file, next to the step it is about, that fails the run where it is false. Stating the intended change rather than plain equality is what keeps the check alive. Where the table is already materialised the count is bookkeeping the object maintains; on a deferred pipeline asking for it is real work, so budget it.

code

pseudocode · 9 lines
pseudocode
before    = row_count(orders)
cancelled = row_count(where(orders, status == "cancelled"))

orders = drop_where(orders, status == "cancelled")

expected = before - cancelled
if row_count(orders) != expected:
    raise_error("cancellation step: expected " + expected +
                " rows, got " + row_count(orders))

go deeper

for a junior

Remember the habit: record the row count before a step and compare it after, in the file, as something that fails. Structural damage is what this catches; a wrong value needs a different kind of check.

for a middle

Explain why plain equality is the wrong claim for a step that removes rows, and show the computed expectation instead: count the intended removals first, then assert the difference. Say why a printed number is not a check.

for a senior

Show where the claim belongs so it names the step that broke rather than the end of the run, and be honest that the count is bookkeeping on a materialised table and real work on a deferred one.

for a principal

The angle is what evidence a result must carry before anyone acts on it, and what the standing rule costs: which steps must carry a claim, how many passes over the data the team is willing to spend on evidence, and who notices when a claim is deleted.

## What the claim actually says A row count assertion is one sentence about structure: **this step received N rows and returned M**, where the relationship between N and M is something you decided before the step ran. It says nothing about any value in any cell. The **row and column counts** — how many rows and columns a table holds at this point in the file — are the cheapest claim there is to make about a result, and on a materialised in-memory table the object already knows them. The second half of the idea matters as much as the first. An **inline assertion** is a claim written in the same file, beside the step it is about, that **fails the run at the point where it is false**. A claim that lives anywhere else is worth much less, and one that nothing acts on is worth nothing. ## Why a structural claim finds more than reading values - **The screen is a sample you did not choose.** A viewer shows the head and the tail of a table. A defect concentrated in one region, one source system or one day is invisible there, and a defect spread thinly is invisible everywhere. - **Most damage in this kind of work is structural.** Rows multiplied because a match found several partners; rows quietly removed because a condition cannot be satisfied by an absent value; records the reader skipped before the data became a table at all. Every one of those moves the row count. None of them makes an individual visible cell look wrong. - **A value has no baseline.** Reading `4173.22` tells you nothing unless you already know what it should be. Reading `1,000,000 -> 1,412,033` tells you immediately that a step you believed removed rows added four hundred thousand. - **One number compares mechanically.** A machine can compare two integers on every run, for ever, without getting bored on the fiftieth run. None of this makes the count a check on correctness of values. It is a check on structure, and structure is where the cheap wins are. ## State the intended change, not equality The common mistake is to assert that the count is unchanged across a step whose whole purpose is to change it. That claim fails on every correct run, so somebody deletes it, and with it goes the only evidence the step ever produced. The pattern that survives is: 1. Measure what the step is supposed to do — count the rows that match the removal condition **first**. 2. Run the step. 3. Assert the computed expectation: `after == before - removed`. Now an unintended loss of eleven rows fails, and an intended removal of nine hundred thousand passes. The same shape works for a step that adds rows: count what should arrive, then assert the sum. ## Beside the step, not at the end of the run | where the claim lives | what a failure tells you | what you still have to do | |---|---|---| | beside the step | this step, this input, this expectation | open the step that was named | | once at the end of the run | one of eleven steps is wrong | re-run, inserting counts, to bisect it | | nowhere — you read the output | something looks off, or nothing does | start from the top with no evidence | Failing early has a second benefit beyond diagnosis: the run stops before a wrong result is written somewhere another person will read. ## Printing is not asserting Printing counts to the log is what most people do, and it is not a check. It requires a human to read the number, remember the previous number and decide. On the day the defect appears, that human is running the file for the fortieth time and is looking at something else. Turn the print into a comparison that raises. ## What it costs, and what it does not tell you The cost is not the same everywhere, and asserting otherwise is how the habit gets a bad name. On a **materialised in-memory table** the count is bookkeeping the object already maintains, so the claim is effectively free. On a **deferred pipeline** — one that builds a plan and computes nothing until a result is demanded — asking for a count is asking the plan to run, and asking again for the next count can run it again. On a reader that walks a file in chunks, a count can mean another pass over the input. Where the claims are not free, collect the ones you need in as few passes as you can. And be honest about the limit: a failed count says **rows moved**, not which rows and not why. That is the start of the next investigation, not the end of this one — but it is an enormously better starting point than a number at the bottom of a file that somebody thinks looks low.

  • The step lives inside a function called from three places. Where does the claim go?
    Inside the function, beside the step, so every caller is covered by one claim. Express it in terms of the function's own inputs — the count it received, minus the removals it measured — rather than a literal taken from one caller's data, which would be wrong for the other two.
  • What do you do with a count assertion that fires every month on legitimate input changes?
    Treat it as a wrong claim, not a noisy one. A literal row count is almost never the right expectation; a relationship usually is — after equals before minus what was measured for removal, or after equals the number of distinct entities. If the relationship genuinely varies, you have not yet understood what the step guarantees.

A load of parcels changes hands at four depots. Counting them at each handover takes seconds and tells you which depot lost one; opening parcels at the far end tells you only that something is missing. The count is near-free where the load is already stacked and counted, and costs a re-weigh where it is not.

saying these in an interview costs you the question

  • Says reading the first rows of the output is enough
  • Prints counts to the log instead of failing the run
  • Collects every check at the end of the run
  • Asserts equality across a step meant to remove rows
  • Deletes the assertion because it keeps firing
  • Treats a matching count as proof the values are right
open as a page

In a long-lived interactive session, why can the result on screen come from a program not in the file?

level: juniorimportance: must knowfreq 72%

basics

~20 s

An interactive session keeps every value any block ever bound, including blocks later edited or deleted. The number on screen reflects the order blocks were actually executed, and the file records neither that order nor the code that produced it.

open as a page

Eight chained steps take 10,000 rows in and return 9,860 with no error raised — how do you find the step that lost them?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Record the number of rows after every step in one instrumented run, then read the sequence for the first adjacent pair where it falls. Two endpoint numbers give the size of a loss and never its location.

open as a page

A column of numbers now sorts with 10 before 9. What does that tell you, and why did no step raise?

level: juniorimportance: must knowfreq 74%

basics

~20 s

A column ordered character by character is being held as text rather than as numbers: '10' comes before '9' because '1' comes before '9'. Nothing raised because putting text in order is a legal operation.

open as a page

Two versions of a transform return the same 40,000 rows in a different order, yet a row-by-row comparison reports 39,000 differences. Why?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Nothing promises that two versions emit rows in the same sequence, so a position-by-position comparison is measuring order rather than correctness. Align both outputs on the columns that identify a row, then compare the matched pairs.

open as a page

Before matching a table to a lookup table on a key, what must you assert about the lookup side, and why beforehand?

level: middleimportance: must knowfreq 62%

basics

~20 s

Assert that the key columns identify a row: the number of distinct combinations of those columns equals that side's row count. Stated before the match, a failure names the table, the columns and the step instead of a row count that grew later.

open as a page

A tolerant comparison of a rewritten transform's 12-million-row output against the old one returns "not equal". What do you produce next?

level: seniorimportance: must knowfreq 58%

basics

~20 s

A verdict is not evidence. Produce the differing rows themselves: keys only in the old output, keys only in the new one, and matched keys whose values disagree, each with a magnitude and a direction, plus a sample a human can read.

open as a page

A layout change is expected to change a table's row count — so what claim replaces the row count on that step?

level: middleimportance: should knowfreq 38%

basics

~20 s

Assert a quantity the step is supposed to preserve: the number of non-absent value cells, the total of each measure, and the number of distinct entities. Rearranging a table moves measurements around; it must not create or destroy them.

open as a page

Executing a block a second time in the same session changes the numbers — what makes a block safe to re-execute?

level: middleimportance: should knowfreq 58%

basics

~20 s

A block is safe to execute again when it reads only names bound earlier in the file and does not write a name that also appears on its own right-hand side. Blocks that accumulate into, grow or adjust their own input drift on every run.

open as a page

A random-drawing step records its seed beside the result — what does the recorded seed fix, and what does it not?

level: middleimportance: should knowfreq 45%

basics

~20 s

A recorded seed fixes the sequence of draws taken from the generator that was seeded, in the order they are taken. It reaches no other generator, no parallel worker drawing from its own stream, and nothing whose result depends on an arbitrary iteration order.

open as a page

A pipeline deliberately removes test accounts and cancelled orders — how do you keep those removals out of an unexplained-loss investigation?

level: middleimportance: should knowfreq 40%

basics

~20 s

Give every intended removal its own named step that reports how many rows it removed, so the run produces a balance: input equals output plus each named removal plus a remainder. Anything in that remainder is unexplained by construction.

open as a page

140 records vanished at one step — why is the difference of two counts not enough, and how do you recover the records?

level: middleimportance: should knowfreq 52%

basics

~20 s

A difference of counts is a magnitude you cannot inspect or group. Recover the rows: keep the input rows whose key value is absent from the output, using a key unique in the input and unchanged by the step.

open as a page

A column became text at step two, but the run only fails at step six. Why does the failure surface there?

level: middleimportance: should knowfreq 58%

basics

~20 s

Every step between the change and the failure was defined on text, so each one succeeded and passed the column along. The failing step is simply the first one that required a numeric column and could not get one.

open as a page

A comparison of two transform outputs reports them unequal while every number on both sides agrees. What else can it be objecting to?

level: middleimportance: should knowfreq 46%

basics

~20 s

A strict comparison answers about three things at once — the values, the representation each column is held in, and the structure (column names and order, and whether each side carries row labels). One verdict hides which of the three moved.

open as a page

A rewritten aggregation agrees with the old one to twelve digits but not exactly. Why is exact equality the wrong verdict here?

level: middleimportance: should knowfreq 52%

basics

~20 s

Two correct implementations that combine the same fractional numbers in a different order do not produce bit-identical results. Exact equality is a tolerance of zero, chosen by omission; a diff has to state how close counts as the same.

open as a page

You add a row count and a distinct-key count either side of ten steps and the run slows badly — which of those claims is free and which is not?

level: seniorimportance: should knowfreq 44%

basics

~20 s

On a materialised in-memory table the row count is bookkeeping the object maintains and costs nothing, while a distinct-key count must read every key value, so it is a full pass. On a deferred pipeline even the row count forces the plan to run.

open as a page

A clean run from empty reproduced the result, but a colleague's run does not — what did the clean run not prove?

level: seniorimportance: should knowfreq 50%

basics

~20 s

A clean run from empty proves only that no value living inside the previous process was holding the result up. Anything the work left outside the process — written files, caches, inserted records, local paths, installed versions — survives the restart untouched.

open as a page

You recovered the 140 lost records — why is grouping them by source or day more useful than their count, and what could mislead you?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A loss concentrated in one source, region or day is a structural defect with a bounded repair; a scattered loss is a per-record property. Grouping the residue separates them — but compare each group's share against its share of the input.

open as a page

Where in a multi-step file is a column's representation worth asserting, and what must the assertion claim?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Assert at each boundary where the representation can change and just before the first step whose correctness depends on it. The claim must be that the column is held as a number, not merely that its values could be read as one.

open as a page

A report's total is wrong, the column prints normally, and it is held as text. How do you find the step that changed it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Record the column's declared representation at each step and find the first point where it is no longer numeric. The values cannot answer this: a column of digits held as text prints and aggregates much like a numeric one.

open as a page

A diff of two transform outputs flags 4,000 cells where both sides hold no value at all. Why, and what must the comparison state?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Absence is not a value that equals itself. Depending on the design, comparing two absent cells yields false, or yields absence rather than a verdict — neither is "equal" — so a diff has to state its rule for absent cells rather than inherit one.

open as a page