skip to content

A pipeline deliberately removes test accounts and cancelled orders — how do you keep those removals out of an unexplained-loss investigation?

level: middleimportance: should knowfreq 40%

answer

  1. unexplained means what is left over
  2. one removal, one step, one number
  3. the run should balance
  4. state the band you expect
  5. combined conditions report nothing

basics

~20 s

Give every intended removal its own named step that reports how many rows it removed, so the run produces a balance: input equals output plus each named removal plus a remainder. Anything in that remainder is unexplained by construction.

solid answer

~50 s

The residue has to be defined by subtraction, and that only works if the intended removals are countable. So each deliberate removal becomes its own step, doing nothing else, reporting its own number — 4,120 test accounts, 860 cancelled orders — instead of being folded into a condition that also does real work. The run then emits an identity: the input count equals the output count plus every named removal plus a remainder, and the remainder is the only thing worth investigating. Two further points. An intended removal is not automatically a correct removal: state the band you expect it to fall in, because a step that normally removes 4,000 rows and today removed 40,000 is a defect wearing the clothes of an intention. And a removal folded in with other logic cannot be counted at all, so its number is guessed and the remainder becomes an argument.

go deeper

for a junior

The habit to recall is that deliberate removals should be their own steps that say how many rows they took, so that a later question about missing records has an obvious starting number.

for a middle

Explain the balance the run should produce — input equals output plus each named removal plus a remainder — and why a single combined condition makes that balance impossible to write.

for a senior

Demonstrate the band: an intended removal is also a step that can misbehave silently, so state the proportion it is expected to take and compare every run against it rather than trusting the intent.

for a principal

The tradeoff is how much the accounting is worth: separate steps cost intermediate results and, on a deferred design, extra execution, against investigations that start from a number rather than from a disagreement.

## The residue is defined by subtraction An unexplained loss is whatever is left after everything explained has been accounted for. That definition is only usable if the explained part is a number somebody can produce. In a pipeline that deliberately removes test accounts, cancelled orders, records outside the reporting period and rows from a decommissioned source, the explained part is large — and if nobody counted it, every investigation starts by arguing about how big it should have been. The fix is structural and cheap: **make intended removals explicit, separate and counted**. ## One removal, one step, one number - Each deliberate removal is **its own step**, and it does nothing else. It does not also rename columns, derive a field or reorder rows. - The step is **named for its intent** — "remove test accounts", not "clean" or "prepare". - The step **reports how many rows it removed**, and that number is part of the run's output, not something recovered later from logs. - The removals happen in a **known place** in the chain, so their counts sit at fixed positions in the sequence rather than being interleaved with transformations at random. What this buys is that the run can emit an identity, and identities are checkable: ``` input rows - removed as test accounts 4,120 - removed as cancelled 860 - removed as outside the period 2,190 = expected output rows vs. actual output rows → remainder ``` A non-zero remainder is, by construction, loss nobody intended. That is the whole point: the investigation now begins with a number that means something instead of with a disagreement. ## Why folding removals into other logic destroys this A condition written to keep "active, non-test orders in the period" removes three different populations in one operation and reports none of them. Consequences: - the number removed for each reason is unknown, so the balance cannot be written; - a removal that also takes rows nobody intended is invisible, because there is no expected figure to compare against; - when a rule changes — test accounts get a new marker, the period moves — nobody can tell which part of the fall in output is the rule change and which is a regression; - the next reader cannot tell whether the operation was primarily about removing rows at all. | Written as | What the run can report | What an investigation starts with | |---|---|---| | Three named steps, one reason each | A count per reason plus the remainder | A specific unexplained number | | One combined condition | A single output count | An argument about what the expected count was | ## An intended removal is not automatically a correct removal Counting the intended drops is necessary and not sufficient, because a step whose job is to remove rows will happily remove far too many and still look like it worked. Two habits close that gap: 1. **State the expected band** beside the step — this normally removes 3–6% of the input — and compare against it on every run. A step that usually removes 4,000 rows and today removed 40,000 is a defect, and it is the kind that never raises anything because removing rows is exactly what it was written to do. 2. **Keep the removal's own criteria narrow** so that an unexpected value cannot widen it. A removal expressed as a condition can also take rows that nobody had in mind, and the point of the band is to catch that without needing to reason about why it happened. A removal that fails its band is not an unexplained loss — it is an explained step behaving unexpectedly, which is a faster thing to fix because the step is already named. ## What this looks like in the investigation With the removals explicit, the earlier passes become short: - the sequence of counts now has labelled falls, so the bisect skips straight past the intended ones; - the residue recovered by excluding survivors contains no test accounts and no cancelled orders, so grouping it is not swamped by rows that were supposed to go; - the characterisation of the residue is about the defect rather than about the business rules. Without them, every one of those passes is polluted, and the most common outcome is an investigator concluding there is no problem because the loss "looks like the filters" — which is a conclusion drawn from a resemblance rather than from an accounting. ## The cost Several small steps instead of one condition means more intermediate results, and on a design where each step materialises its own output that is real memory and real time. On a deferred pipeline the extra counts can force extra execution. Both are usually worth paying, and where they are not, the compromise is to keep the removals combined but compute each reason's count once, in the same pass, so the accounting survives even though the steps do not.

  • An intended removal step took 40,000 rows on a run where it normally takes 4,000, and nothing failed. Why is that hard to notice?
    Because removing rows is the step's purpose, so a large removal looks like success. Nothing raises, the output is still well formed, and the only visible symptom is a smaller total downstream. An expected band stated beside the step is what converts it into a detectable event.
  • The pipeline runs as one combined condition and splitting it is too expensive. What is the minimum you keep?
    The per-reason counts, computed once in the same pass over the data even though the removals stay combined. The accounting is what the investigation needs; separate steps are just the most convenient way to get it. Record the numbers as part of the run's output so the balance can still be written.

saying these in an interview costs you the question

  • Assumes any fall in the row count is fine because filters exist.
  • Folds several unrelated removals into one condition and reports one total.
  • Counts intended removals but never states what count was expected.
  • Treats a removal step as correct simply because removing rows is its job.
  • Reconstructs the intended counts afterwards by rerunning parts of the chain.
  • Names removal steps after their mechanics rather than their intent.