When should the expected one-to-many relationship at a nightly match be a declared property that fails the run rather than a number someone reads?
answer
- the expectation exists either way
- cost of a wrong number against a stopped run
- how often the pairing legitimately changes
- a ladder, not two choices
- a check that always fires gets deleted
basics
~20 sDeclare it when a wrong number is more expensive than a stopped run, and when the pairing the step depends on is stable enough that a violation really is a defect. Where it is expected to change, measure and publish the number instead of failing.
solid answer
~50 sThe decision is not about matching technique, it is about what a violation should cost. Declaring the relationship - one row per key here, many allowed there - converts a silently inflated table into a failure at the seam where the assumption lives, with the offending step named. That is worth a great deal when the output feeds money, a regulator, or a model, and it is worth less when the pairing legitimately changes and the run stops for something nobody considers wrong. So weigh three things: how expensive a wrong number is downstream and how long it survives before anyone notices; how often the expected pairing genuinely shifts; and whether anyone is awake to act on a failure at the hour it fires. The two ends of the ladder are silence and a hard stop, but the useful middle exists too - measure the relationship, publish the number with the output, and fail only on the part of it that is genuinely invariant.
go deeper
Know that the pairing a step assumes can be written down and checked rather than remembered, and that a check placed next to the step tells you which step was wrong.
Be able to name both mechanisms - a uniqueness assertion with a predicted count around the step, or the relationship declared on the match where the tool accepts one - and say what each catches.
Show that you have operated one of these: what the check does to the run, what it leaves behind for diagnosis, and why a check that fires for legitimate reasons gets disabled.
Reason about the whole ladder rather than the two ends, price a wrong number against a stopped run for this specific consumer, and say who is allowed to decide a violation is the new normal.
## What is actually being decided Every match carries an expectation about how the two sides pair: one row per key on the reference side and many on the transaction side, or one to one, or an expansion that is deliberate. That expectation exists whether or not anyone writes it down. The decision here is only about **where it is recorded and what happens when reality disagrees** - not about how the pairing works. The usual failure is not choosing wrongly. It is not choosing: the expectation stays in the head of whoever wrote the step, the data changes, the table inflates, and a number reaches a report with nobody in the path who knew what it was supposed to be. ## The ladder of responses | response | when a violation surfaces | what it costs | |---|---|---| | nothing recorded | when someone disbelieves a downstream number | days or never; the cause is far from the symptom | | count measured and logged | when someone reads the log | needs a reader; usually has none | | count published beside the output | when a consumer looks at the run's own numbers | cheap, and creates a record to compare across runs | | declared on the step, failing the run | immediately, at the step, named | a stopped pipeline, including for legitimate change | | declared, failing, with the output quarantined | immediately, and yesterday's output still stands | the most operational machinery to build | Most teams should be further down this ladder than they are, and very few should be at the top of it everywhere. ## The questions that place a given seam on that ladder 1. **What does a wrong number buy?** A total that drives an invoice, a payout or a regulatory filing is in a different class from an exploratory count. The cost of being wrong, multiplied by how long wrong survives, is the whole case for failing hard. 2. **How stable is the pairing?** A reference maintained deliberately by a team changes shape rarely, and a change is genuinely news. A feed assembled from several upstream systems may add a second row per key for perfectly good reasons, and a hard failure there trains people to disable the check. 3. **Who is awake?** A failure at 03:00 with nobody on call is only better than a wrong number if the failure leaves the previous output intact. If it takes the whole pipeline down until morning and the wrong number would have been caught at 09:00 anyway, you have converted a data defect into an availability incident. 4. **How reversible is publication?** Where consumers pull from a location you can leave untouched on failure, failing is cheap. Where the step's output is broadcast, failing late is worse than not publishing at all. 5. **What does the tool actually support?** Some tools accept the expected relationship on the match itself and raise when the data violates it. Where that exists, using it removes the distance between the assumption and its test. Where it does not, the portable form is a uniqueness assertion before the step and a predicted count checked after, and the decision above is unchanged - only the mechanism differs. ## Where the check belongs At the seam, with the step that depends on it. The same two tables can be matched in three places with three different expectations: one step wants exactly one row per key, another is content with an expansion it then reduces. A single global rule about a table cannot express that, and it fails in the wrong place, blaming a table rather than the step whose assumption was violated. Separately from anything agreed with whoever produces the data, the consuming step still has its own expectation and still has to decide what to do about it. ## What a violation should do When you do choose to fail, decide these before writing the check, not during the incident: - **Does the previous output stand?** Leaving yesterday's result in place is usually better than publishing an inflated one and almost always better than publishing nothing at all. - **What is kept for diagnosis?** The offending key values and their occurrence counts on each side are small, and they are the difference between a five-minute fix and a re-run to reproduce. - **Who decides it is the new normal?** A violation that turns out to be legitimate needs a person to change the declaration. If the fastest route back to green is to delete the check, the check will be deleted. ## The failure mode of declaring everything A check that fires often for reasons nobody considers defects is worse than no check, because it teaches a team to route around the mechanism. Reserve the hard stop for seams where the expectation is genuinely invariant and a wrong number is expensive, publish the measured relationship everywhere else, and re-examine the placement when a check has fired three times and been overridden three times - that is data about the expectation, not about the pipeline.
- Is there a defensible middle between silence and failing the run?Yes, and it is usually the right answer: measure the relationship on every execution and publish the numbers beside the output - rows and distinct key values on each side, and the output count against the prediction. Consumers get something to compare across runs, a trend becomes visible before it becomes a defect, and you keep the hard stop for the part of the expectation that is genuinely invariant.
- What makes a hard failure at a match cheaper to operate than it first looks?Leaving the previous output in place. A failure that publishes nothing new but keeps yesterday's result serviceable turns a 03:00 stop into a morning task rather than an outage, and it removes the pressure to disable the check at the moment someone is least able to judge whether the change is legitimate.
- A declared expectation has fired and been overridden three times in a quarter. What does that tell you?That the declaration, not the data, is wrong. The pairing you encoded is not the pairing the source actually has, so the check is measuring an obsolete belief and training the team to route around the mechanism. Change the declaration to the relationship that is genuinely invariant, and move the rest to a published measurement.
saying these in an interview costs you the question
- Argues every match should fail hard on an unexpected relationship
- Treats logging the row count as equivalent to acting on it
- Puts the expectation on the table globally rather than at the step that holds it
- Ignores that a stopped nightly run has a real cost of its own
- Keeps a check that fires monthly and is overridden every time