A schema-drift mismatch in a fleet telematics ingest reached the field. How do you judge which phase should have contained it?
answer
- List the checkpoints it walked past
- Earliest link that could have caught it
- Absent, wrong oracle, or never ran
- Ask what the test data made possible
- One containment change, then verify recurrence stops
basics
~20 sWalk the chain of checks the change passed through, find the earliest one that could cheaply have caught it, and decide whether that check was absent, wrong, or skipped. That answer, not the phase that reported it, is the escape phase.
solid answer
~40 sReconstruct the checkpoint chain the change actually passed - interface agreement, review, unit, integration, system, acceptance, monitoring - and at each one ask a single question: **could this check have caught the mismatch at reasonable cost?** The escape phase is the earliest one where the answer is yes. Then separate three very different gaps: no check of that kind existed, a check existed but asserted the wrong thing, or a check existed and did not run. Record what the environment made possible too - a test environment fed only clean synthetic payloads cannot catch drift no matter how good the assertions are. Finish with one containment change at the phase you named, plus the origin classification pointing at prevention. An escape review that produces six actions usually delivers none of them.
code
pseudocode · 11 linesfor check in [CONTRACT, REVIEW, UNIT, INTEGRATION, SYSTEM, MONITORING]:
if could_have_caught(check, defect) and cost(check) is acceptable:
escape_phase = check
gap_kind = ABSENT | WRONG_ORACLE | DID_NOT_RUN
break
# example outcome
escape_phase = UNIT
gap_kind = ABSENT # no case fed a text-shaped reading
env_note = "all 340 pack cases use generated numeric payloads"
action = one change at UNIT + a loud boundary rejectiongo deeper
Focus on the idea that a defect passes several checkpoints and the interesting question is which one had a fair chance. Be ready to name the usual chain from review through to field monitoring.
Explain the mechanics: walk the chain, ask at each link whether the check could have caught it cheaply, and distinguish a missing check from a check with a weak assertion or one that never ran.
This is the level the question targets. Show that you reason about what the test data and environment made possible at all, pick one high-leverage containment change rather than six, and verify later that the class stopped recurring.
Own the pattern across escapes rather than the single case: which link is repeatedly thin, whether the checks sit at the right cost tier, and how to keep the review blameless so the classification stays honest enough to aggregate.
## The concrete case A fleet telematics ingest accepts periodic vehicle payloads. After a provider firmware roll-out, roughly 0.8% of devices began sending the odometer reading as quoted text rather than a number. The ingest coerced the unparseable value to zero, stored it, and reported nothing. Distance-travelled reports for those vehicles read zero for **11 days** before a customer noticed. The regression pack for the ingest holds **340 cases** and all of them passed on every build during those 11 days. That last fact is the whole subject: the checks ran, were green, and the defect walked past all of them. Judging the containment gap means deciding **which of those checks should have stopped it**, and being specific about why it did not. ## Walk the chain, ask one question at each link List the checkpoints the change actually passed through, earliest first, and at each one ask: *could this check have caught this defect at reasonable cost?* * **Interface agreement.** Was the numeric form of the reading ever written down as part of the contract with the provider? If not, there was nothing for any later check to compare against, and the origin is a requirement gap rather than anything in the code. * **Review.** Would a reader of the ingest change plausibly have noticed that an unparseable reading becomes zero silently? Silent coercion is a review-catchable pattern; a drift in someone else's firmware is not. * **Unit level.** Could a case with a text-shaped reading have been written cheaply? Almost always yes - this is usually the earliest link where the answer flips to yes, and that makes it the strongest candidate for the escape phase. * **Integration level.** Was there any check comparing what the provider actually sends against what the ingest expects? A recorded sample of live traffic replayed against the parser would have caught it the day the firmware shipped. * **System level.** Did any end-to-end run use payloads shaped like production, or only clean synthetic ones? A system phase fed exclusively by generated clean data has no chance at drift, and naming it as the escape would send effort somewhere it cannot help. * **Field monitoring.** Was there any signal for the rate of coerced or rejected values? Eleven silent days is itself a containment gap, at the last checkpoint of all. The escape phase is the **earliest link where the answer was yes**, because that is where the check would have been cheapest. Do not simply record "field" because that is where it was reported; found-in and escaped-at are different fields, and only escaped-at tells you where to invest. ## Absent, wrong, or skipped Three gaps look identical in a report and need completely different fixes: 1. **No check existed.** Nobody ever wrote a case for a malformed reading. This is a test-design gap; the fix is a new case and, more importantly, a rule that new input handling ships with malformed-input cases. 2. **A check existed and asserted the wrong thing.** A case fed a text reading and asserted that ingest "did not throw" - which the coercion satisfied. This is worse than no check, because it bought false confidence. The fix is the oracle: assert the stored value and the rejection, not the absence of an exception. 3. **A check existed and did not run.** The case sat in a quarantined group after an earlier flake, or the phase was cut for schedule. This is a process gap, and adding another case fixes nothing. ## What the environment made possible Always record the data and environment the checks ran against, because it caps what any of them could have found. If every one of the 340 cases uses generated payloads that are numeric by construction, the pack is structurally incapable of seeing drift, and the honest containment finding is not "we needed case 341" - it is "no phase consumed a real-shaped payload". That is a much bigger, much more valuable conclusion, and it is invisible if you stop at the first missing case. ## Turning it into prevention The pair now does its work. Origin says whether to prevent the class - if the payload contract never specified the reading's form, prevention is an explicit, versioned contract with the provider, plus a validating boundary that rejects loudly instead of coercing silently. Escape says where to add detection - if the answer flipped to yes at the unit link, a case with a text-shaped reading belongs there, not another slow end-to-end case. Two disciplines keep this useful. **Pick one containment change**, at the highest-leverage link you can afford; a review that emits six actions across five teams delivers none. And **check back**: the only evidence the change worked is that this class stops appearing, which is a question you can only answer if the class was classified consistently in the first place. Keep the whole exercise blameless. The finding is "an input the system could not represent was accepted silently, and no phase used production-shaped data", never "the ingest author missed a case". Escape data goes dishonest the first time it costs someone a review.
- The pack was green throughout. Does that make the pack worthless?No, it makes its coverage boundary visible. A pack built entirely on generated well-formed payloads is doing its job on the behaviour it encodes and is structurally blind to input drift. The finding is about the missing class of input, not the pack's value, and it is why 'all green' is never on its own evidence of containment.
- How do you decide between adding a case at the unit link and adding one end to end?Prefer the earliest link that can hold the assertion honestly. A unit-level case for a text-shaped reading is fast, deterministic and runs on every change; an end-to-end case for the same thing costs run time forever and fails for a dozen unrelated reasons. Reserve the end-to-end slot for the wiring that only exists once assembled.
- Should silent field coercion count as a defect if no report ever came in?Yes, and the containment finding is about the monitoring link. A value the system cannot represent should be visible - a counter of coerced or rejected inputs turns eleven silent days into an alert on day one. Silent recovery is a detection gap even when the recovery itself was reasonable.
- What if two people disagree about which link should have caught it?Make the disagreement explicit as a cost question rather than an opinion: at which link would the check have been cheap enough to be worth having, given what that phase can see? Record the earliest link where both would say yes, and note the dispute. A written tiebreak beats re-arguing it every escape.
Water gets into a basement past a gutter, a membrane and a sump pump. Knowing it came in tells you nothing useful; knowing which of the three had a fair chance and failed tells you what to repair.
saying these in an interview costs you the question
- Recording the phase that reported it as the escape phase
- Naming a phase that could never have seen the defect
- Treating a green suite as proof of containment
- Adding one end-to-end case and calling it prevention
- Ignoring whether the check existed but never ran
- Framing the finding as an individual's missed case