Eight chained steps take 10,000 rows in and return 9,860 with no error raised — how do you find the step that lost them?
answer
- two numbers measure, they do not locate
- a number after every step, one run
- first adjacent pair that falls
- unchanged count means net zero only
- counts are free only when materialised
basics
~20 sRecord the number of rows after every step in one instrumented run, then read the sequence for the first adjacent pair where it falls. Two endpoint numbers give the size of a loss and never its location.
solid answer
~50 sThe two numbers you have measure the loss and say nothing about where it entered. Instrument the chain: after each of the eight steps, record the step's name and the number of rows it produced, in a single run, and print the sequence in execution order. The first adjacent pair where the number falls localises the loss to one step, and that is all this pass is for — *why* those rows left is a separate question about whatever that step does. Two cautions. Equal counts either side of a step rule out only a **net** change, because a step can remove some rows and multiply others. And on a pipeline that defers work until a result is asked for, each count can force the plan to run, so gather the counts you need in one instrumented pass rather than sprinkling them.
go deeper
Recall the move rather than the theory: when rows go missing with no error, print how many rows exist after every step, in order, and look for the first fall. Location first, cause second.
Explain why a single pass matters and what an unchanged count does and does not rule out, and be able to say what asking for a count costs on a materialised table against a pipeline that defers work until a result is requested.
Show that you instrument once, with named steps, on a run whose input cannot shift underneath you, and that you treat a clean bisect with wrong downstream totals as a signal to compare key sets rather than to count again.
The angle is what a pipeline should emit as a matter of course so nobody has to reconstruct this by hand, weighed against the measurements that force an extra pass over the data and the runtime they buy back.
## What the two endpoint numbers prove A run that took 10,000 rows and returned 9,860 has established exactly one fact: somewhere between the first step and the last, the row count fell by 140 **net**. It has not established where, and it has not established how many rows actually left — 140 is the difference between everything removed and everything added along the way, so the true number of departures can be far larger. The instinct at this point is to theorise about a cause: the match, the condition, the reader. Resist it. Guessing at a cause before you know the step is how an afternoon disappears into an argument about the wrong operation. ## Turn one number into a sequence of numbers Instrument the chain so that each step, in order, reports how many rows it produced: - record the input count once at the top, then the count **after** each step; - give every step a stable **name**, so the printed sequence is readable by the next person and by you in six months; - keep the numbers in one list in execution order and print them together, rather than scattering them through the output; - collect them in **one run**. Re-running the chain repeatedly with a different count added each time compares numbers from different executions, and if the input can change underneath you, or any step involves randomness, those numbers are not comparable at all; - include the steps you believe are innocent. The cheap ones are exactly the ones nobody instruments, and a loss hides well behind an operation nobody suspects. Reading the result is then trivial. The first adjacent pair where the number falls is where the loss entered. Everything before it is exonerated. ## This pass localises; it does not explain Once the step is known, the reason those rows left is a property of whatever that step does — a match that found no partner on one side, a condition applied to the rows, a reshape that combined several rows into one, a reader that refused some input lines. Each of those is its own subject with its own mechanism. Keeping localisation and explanation apart is what makes the investigation finite: one pass finds the step, the next recovers the rows, and only then does the cause become a question with a small enough surface to answer. ## Equal counts are weaker evidence than they look | What the sequence shows at a step | What it establishes | What it does not establish | |---|---|---| | The count falls | The net loss entered here | Why, or whether this step also added rows | | The count is unchanged | The step's **net** effect on the row count was zero | That the output holds the same rows as the input | | The count rises | The step produced more rows than it consumed | That nothing was lost at the same time | A step that removes 200 rows and turns 60 others into 260 reports an unchanged count. So if the bisect comes back clean but the downstream totals are still wrong, the next move is not another count — it is to compare the *sets* of key values either side of the flat step, which is a different and more expensive measurement. ## What a count costs, and when that matters The cost of asking "how many rows is this" is a property of the design you are working in, not a constant: - on a **materialised** in-memory table, the row count is metadata the object already holds, and asking for it is effectively free; - on a **deferred** pipeline — one that records a plan and executes only when a result is demanded — asking for a count can force the whole plan to run, and asking again for the next step can force it again, so ten counts can mean ten executions of everything upstream; - on a **chunked** read, where the input is walked a piece at a time, a count means another full pass over the input. The practical consequence is not "do not count". It is to gather the counts in one instrumented pass — carrying the running totals alongside the work, or materialising once at the point where you intend to measure — and to budget the measurements that force extra work instead of discovering them as a mysterious slowdown. ## After the bisect A localised loss is a starting point, not a finding. The next deliverable is the lost rows themselves, recovered by excluding the survivors, because a count cannot tell you whether the 140 are one region, one source, one day, or scattered across the whole input — and those are different defects with different repairs.
- The counts either side of one step are identical, yet the totals downstream are still wrong. What could that step have done?Removed some rows and produced extra rows for others in the same operation, so the net change is zero while the contents differ. An unchanged count rules out a net change and nothing more. Confirm it by comparing the distinct key values on each side of that step rather than the counts.
- You have localised the loss to one step. What is the next artefact you produce, and why is it still not an explanation?The lost rows themselves, recovered by taking the step's input and keeping the rows whose key value is absent from the step's output. That gives you something to look at and group. It still says nothing about why they left — that depends on what the step does, and is answered by examining the operation, not the residue.
A parcel posted in one city and never delivered in another. The two facts you start with prove only that it is gone; the depot scans along the route are what tell you which leg it stopped on. The scan does not say why it stopped — it says where to go and ask.
saying these in an interview costs you the question
- Guesses which step is to blame from the final count alone.
- Treats the endpoint difference as the number of rows that actually left.
- Believes an unchanged count proves the step kept every row.
- Re-runs the whole chain repeatedly, adding one count per run, and compares across runs.
- Assumes a step that raised no error cannot have removed rows.
- Sprinkles counts through a deferred pipeline without noticing each one re-executes it.