skip to content

A cut that keeps every counterpart across three chained joins grows toward the whole input - where do you stop, and what do you tell the team that absence in the cut means?

level: principalimportance: nice to knowfreq 25%

answer

  1. completeness is transitive
  2. upward cheap, downward unbounded
  3. close along the joins the job walks
  4. reference inputs copied whole
  5. absence is never evidence

basics

~20 s

Close upward always and downward selectively: every kept row's parent must be present, while children are kept only where the job actually joins them. Then publish a manifest saying which joins are complete, so a thin result during development is read as the cut, not a defect.

solid answer

~50 s

Referential completeness is transitive, and that is what makes it expensive. Select customers, keep their orders, keep those orders' payments, keep those payments' settlements - each hop widens the set, and a chain that eventually reaches a shared reference pulls in most of it. The workable rule separates two directions. **Upward closure** - every row's parent is present - is cheap, bounded and non-negotiable, because a missing parent produces orphans that never occur in production and teaches the team a false failure. **Downward closure** - keeping all children of a kept parent - is unbounded, so it is applied only to the joins the job actually performs, and only as far as the join chain the job actually walks. Everything beyond that is knowingly incomplete and must be written down, because a thin result is otherwise indistinguishable from a bug.

go deeper

for a junior

Understand that inputs reference each other in chains, so cutting one input forces decisions about all the others it points at.

for a middle

Explain why closing upward is bounded while closing downward is not, and what an orphan row does to a join result during development.

for a senior

Show that you bound the closure by the join graph the job actually walks, copy small reference inputs whole, and cut at edges that fan back out.

for a principal

Set the standing rule and the artefact that carries it: a manifest naming which joins are complete, and the agreement that absence in the cut is never evidence about production.

## Why closure expands Cutting one input by key is easy. Cutting a *graph* of inputs is not, because referential completeness is transitive. Suppose the job reads four inputs: customers, orders keyed by customer, payments keyed by order, and a settlement record keyed by payment. Select one percent of customers and the cut has to contain: - their orders (bounded by those customers) - the payments of those orders (bounded by those orders) - the settlements of those payments (bounded again) So far each hop is bounded by the previous one, and the cut stays small. The expansion begins as soon as a hop leaves the tree. A payment references a merchant; the merchant references a contract; the contract references every customer on it. One or two such hops and the selected key set has grown back toward the whole input, at which point the cut has stopped being a cut. ## Two directions, two rules The distinction that makes this tractable is between closing *upward* and closing *downward*. | direction | meaning | cost | rule | |---|---|---|---| | upward | every kept row's referenced parent is in the cut | bounded - at most one parent per reference | always close | | downward | every child of a kept parent is in the cut | unbounded - a parent can have any number of children, and each child has its own references | close only along joins the job actually performs | **Upward closure is not optional** because its absence manufactures a condition that does not exist in production. A kept row whose parent was not selected is an orphan. Depending on how the job joins, the orphan either disappears - so counts in the cut are wrong in a way nobody notices - or surfaces as a null, which looks exactly like a source data-quality problem and sends someone to investigate an upstream system that is fine. Both outcomes teach a false lesson, which is worse than no lesson. **Downward closure is where you spend judgment.** It is bounded only by the data itself, so it has to be bounded by the job instead: close along the joins this job performs, stop at the first reference the job never follows, and accept that the cut is incomplete there. ## Where to stop, concretely 1. **Draw the join graph the job actually walks.** Not the source schema's full set of relationships - the joins in this job's code. That graph is usually far smaller than the schema, and it is the only part whose completeness changes an output number. 2. **Anchor on one key, the one most joins hang off.** Select there, and derive every other input's key set from the rows already kept rather than drawing independently. 3. **Cut the graph at any edge that fans back out.** A reference to a widely shared entity - a reference table, a merchant, a contract - is closed upward only: keep the referenced rows, do not keep their other children. 4. **Small reference inputs go in whole.** If a lookup input is a few thousand rows, sampling it buys nothing and every missed row is a false null. Copy it entirely. 5. **Write down what you cut.** The edges you stopped at, and what an empty or thin result on each of them means. ## The manifest is the deliverable The cut is not the only artefact; the statement of its limits is the other half, and on a shared cut it is the more valuable one. Somebody who did not build it will eventually see a join return a tenth of what they expected, and the manifest decides whether that costs them ten seconds or a day. It should say, per input: how its key set was derived, whether it is complete upward, which downward edges were closed and which were cut, which reference inputs were copied whole, and what was capped and by what factor. There is a standing rule worth agreeing once, rather than per person: **absence in the cut is never evidence.** A row that is not there proves nothing about production, a count that is low proves nothing, and an empty result is a question rather than a finding. What the cut *can* prove is positive and local: for the keys it kept completely, the values the job computes are the values production will compute. ## What this never buys Even a perfectly closed cut is a shape, not a size. It says nothing about how long the real run takes, how much memory it needs, or what happens when a worker is lost mid-run, and nothing about the cost of bringing all records for a key together across the cluster - a cost whose mechanics differ between runtimes anyway. Keeping that line visible is part of the same discipline: the team that understands what absence means in the cut is also the team that will not quote its runtime.

  • Why is a missing parent worse than a missing child in a development cut?
    A missing child makes the cut thinner, which is what a cut is for. A missing parent makes it *wrong*: it creates an orphan that production never contains, so the job either silently drops a row - corrupting counts nobody is checking - or emits a null that looks like an upstream data-quality failure. One costs volume you did not want; the other costs a false investigation.
  • What do you do with a small reference input that everything joins to?
    Copy it whole. Cutting an input of a few thousand rows saves nothing measurable and every row omitted becomes a missing counterpart somewhere, producing nulls that read as defects. The size rule is practical rather than principled: if an input fits comfortably in the cut in its entirety, take it entirely and remove a whole class of false failures.
  • Two teams share one cut and disagree about which joins must be complete. How is that settled?
    By making the manifest the contract rather than the cut. Each team names the joins its jobs walk; the union of those edges is closed, and anything outside it is documented as incomplete. If the union grows so large that the cut stops being small, that is the signal to build a second anchored cut rather than to weaken the completeness guarantee of the first.

saying these in an interview costs you the question

  • Keep every related row and the cut will be correct
  • A missing counterpart is a source data-quality problem worth investigating
  • Sample each input independently and reconcile the differences later
  • If the cut returns few rows the transformation must be wrong
  • The cut is complete, so a low count in it predicts a low count in production