Across a pipeline of a dozen matches, what standing rule would you set for the default match shape and rows with no partner?
answer
- where should the cost of an unpartnered row land
- drop, keep and mark, or stop
- twelve one-percent losses are not one percent
- publish an unpartnered count per seam
basics
~20 sThere is no single right rule, only three defensible defaults with different costs: drop unpartnered rows, keep the driving table whole and mark them, or stop the run. Whichever you pick, the disposition should be decided at each seam and written down, not inferred later from a number.
solid answer
~50 sThe real decision is not which shape is best but **where the cost of an unpartnered row should land**. Three defaults are defensible. Dropping unpartnered rows keeps every output narrow and clean, but each match becomes an undeclared filter and the losses compound invisibly across a dozen of them. Keeping the driving table whole preserves the grain and turns loss into visible absence, but every downstream step now has to mean something by that absence, and manufactured absence is indistinguishable from source absence unless you also mark provenance. Failing the run on any unpartnered row gives the strongest guarantee, and it is only affordable where unpartnered is genuinely impossible — on an external feed it converts a data problem into an availability problem. My rule: default to keeping the driving table whole with provenance marked, choose per seam rather than per pipeline, and publish the unpartnered count so the decision is visible.
go deeper
The takeaway at this level is simply that the choice exists and is a choice: an unpartnered row can be dropped, kept and marked, or treated as a reason to stop, and somebody decided which.
Be able to say what each option does to the row count and to the work downstream, and why dropping rows is the option whose cost is hardest to see afterwards.
Argue for a default at a specific seam and name the instrument that makes it safe — provenance marked on kept rows, a published count of rows that found no partner, and a stated expectation next to the call.
Own the compounding and the ownership: a dozen seams each losing a little lose a lot together, and a number nobody is told about is not a control. The rule has to be cheap at the point of use or it will be quietly abandoned.
## The decision you are actually making A dozen matches in a row is not twelve independent choices; it is one policy applied twelve times, and its subject is not "which shape" but **where the cost of a row that found no partner should land**. Every option puts it somewhere: - Drop it, and the cost lands on whoever reads a total later and does not know it is a total over matched rows only. - Keep it, and the cost lands on every downstream step, which must now decide what an empty cell means. - Stop, and the cost lands on whoever is on call. There is no option that has no cost, which is what makes this a standing decision rather than a technical one. ## Three defensible defaults | Default | What it buys | What it costs | |---|---|---| | Keep only partnered rows | narrow, complete outputs; nothing downstream handles absence | each match is an undeclared filter, and across a dozen seams the loss multiplies and is invisible | | Keep the driving table whole | the grain is preserved; loss becomes visible absence rather than vanished rows | every later step must handle absence, and manufactured absence looks exactly like source absence | | Fail on any unpartnered row | the strongest guarantee, and violations surface at the seam that caused them | only viable where unpartnered is truly impossible; on an uncontrolled feed it stops the pipeline for a data problem | The second column is why reasonable teams disagree. An internal pipeline over reference tables you own can afford to fail; a pipeline ingesting an external feed cannot, because the failure mode you are choosing is "no report at all" and that is often worse than a report with a stated gap. ## What compounding means here, concretely The argument against the first default is arithmetic. Twelve matches, each quietly dropping one percent of rows, do not lose one percent — they lose over eleven, and nothing in the pipeline reports either the per-step figure or the total. Worse, the losses are correlated: the entity missing from the customer reference is often the same entity missing from the region reference, so one class of record silently disappears from the whole system while every individual step looks healthy. A single match that drops one percent is a rounding error someone can argue about; a dozen of them is a systematic exclusion nobody chose. ## The rule I would argue for 1. **Choose per seam, not per pipeline.** Each match is a statement about the relationship between two datasets, and those relationships genuinely differ. A blanket rule is easier to police and wrong more often than it is right. 2. **Default to keeping the driving table whole, with provenance marked.** Preserving the grain means row counts stay predictable through the pipeline, and marking which side each row came from is what keeps "had no partner" distinguishable from "partner had nothing there". Predictable row counts are worth a great deal in a long chain. 3. **Publish the unpartnered count at every seam.** One number per match, emitted with the run. It costs almost nothing and it converts an invisible failure class into a visible one, which is the single highest-value thing on this list. 4. **Reserve failing the run** for seams where the reference data is under your control and an unpartnered row genuinely means a bug rather than a fact about the world. 5. **Write the expected disposition where the match is,** in a comment or a named constant next to the call. Not in a document that will be out of date, and certainly not left to be inferred from the shape argument by whoever reads the code next. ## What makes a rule like this survive contact - **It has to be cheap at the point of use**, or people will route around it. Emitting one count per match is cheap; requiring a written justification for every shape is not. - **It has to name a default**, because most seams are unremarkable and a policy that demands a judgment at every one of them spends the team's attention on the eleven that do not matter. - **It has to say who acts on the number.** An unpartnered count nobody owns is a metric, not a control. Name the person or team who is told when it moves, and what "moves" means. - **It should be uniform in style even where it differs in choice.** Two teams choosing differently at two seams is fine; two teams expressing the same choice in two different ways is what makes a pipeline unreadable. ## The trap to avoid The seductive rule is "always keep everything, it is safer". It is not safer, it is deferred. Absence carried forward through twelve steps, each of which decides for itself what an empty cell means, produces a result whose provenance nobody can reconstruct — and the numbers will still look plausible. Keeping unpartnered rows is only a good default **if** something downstream is genuinely committed to handling them. If nothing is, you have chosen the slowest way to get the same wrong answer.
- What makes failing the run affordable at one seam and not at another?Who owns the reference data and what the pipeline is for. Where you own both sides, an unpartnered row means a bug and stopping is the cheapest way to find it. Where one side arrives from a system you do not control, unpartnered rows are a fact about the world, and stopping converts someone else's data problem into your availability problem — usually at the worst possible hour.
- Why publish a count per seam rather than one number for the pipeline?Because a single total tells you something is wrong but not where, and the whole value of the measure is that it points at one match. Per-seam counts also make correlated loss visible: the same entities disappearing at three different seams is a very different finding from one seam degrading alone, and only the per-seam view shows it.
- Is 'keep everything and decide later' ever the right default?Only when something downstream is genuinely committed to deciding. Carrying absence through a dozen steps that each interpret it independently produces a plausible result nobody can reconstruct. If no step owns the decision, keeping the rows has bought nothing except a longer route to the same wrong answer.
saying these in an interview costs you the question
- Declares one match shape correct for the whole pipeline without exceptions
- Calls keeping every row the safe default, with nothing downstream handling absence
- Treats a one percent loss per seam as negligible across a dozen seams
- Proposes a rule with no owner for the number it produces
- Would fail the run on an external feed nobody on the team controls