A defect can be fixed at the failing value, at a missing guard, or in the design behind both. How do you choose the layer?
answer
- Reach, cost, time to payback
- The cheapest fix covers one occurrence
- Ship the fast one, commit the deep one
- Out of reach means contain and make it loud
basics
~20 sChoose by reach, cost and how soon each pays back. Refusing the failing value is cheap but covers one occurrence; a guard covers the family by the next release; a design change removes the class over months.
solid answer
~50 sFor each available layer I ask three things: what it covers, what it costs to build and to live with, and how soon it is worth anything. Refusing the value that failed is cheap and immediate, and covers one occurrence. A guard at the boundary the whole family crosses costs more, covers the family, and pays back at the next release. Making the bad state impossible to build costs the most and pays back over months, so it needs real recurrence or real severity behind it. In practice that means two moves, not one: ship the fastest safe fix to stop the harm, then commit the deeper change with an owner and a date. When the deepest layer is out of reach, because it sits in something you do not control, contain instead: guard your side, make recurrence loud, and record the family as contained rather than removed.
code
pseudocode · 8 linesfix_options for DEF-4127:
A refuse_the_value reach = 1 occurrence effort = hours pays_back = today
B guard_at_boundary reach = the family effort = 2 days pays_back = next release
C state_not_constructible reach = the class effort = 3 weeks pays_back = months
decision: ship A today; B this cycle, owner = payments team
C only if the family arrives again
record: "family contained by B, not removed; C open, owner assigned"go deeper
Know that a defect usually has more than one possible fix and that they differ in how much they cover. Be ready to say which one you took and why, not only what you changed.
Explain the trade between reach and cost, and describe a fix that covers the family rather than the single occurrence you happened to be shown.
Show the two-move answer: stop the harm now, commit the deeper change with an owner and a date, and keep the record honest about which layer actually shipped.
Own when the expensive layer is worth buying and how the team argues for it, on the size of the family and the drag of the workarounds rather than on the drama of one occurrence.
## The ladder of fixes under one defect A defect that has been understood, with its observation, the condition that produced it and the weakness underneath, usually has more than one fix available, and they sit at different depths: - **The failing value.** Refuse, correct or special-case the exact input from the report. - **The missing guard.** Put the rule where every path into that code crosses it, so the whole family is covered. - **The design underneath.** Change the structure so the bad state cannot be built at all, and the guard has nothing left to catch. They are not three candidates for the one true fix. They are three purchases at three prices. ## Three axes, not one Judge each option on: 1. **Reach.** One occurrence, the family, or the whole class of failures, permanently. 2. **Cost.** The effort to build it, plus the cost of living with it: reviewers to convince, code to keep, the risk of breaking something that currently works. 3. **Time to payback.** The fast fix is worth something this afternoon. A guard is worth something at the next release. A structural change is worth something over months, and only if the class keeps arriving. | | Refuse the value | Guard at the boundary | Change the design | |---|---|---|---| | Covers | This occurrence | The family | The class, permanently | | Typical effort | Hours | Days | Weeks | | Worth something | Today | Next release | Over months | | Main risk | The family reads as closed | A new path bypasses the guard | Wide change, wide blast radius | | Justified by | Harm happening right now | A family you can enumerate | Recurrence, severity, or a whole area slowed by workarounds | ## The usual answer is two moves, not one Under real pressure the answer is rarely a single choice. It is: **ship the fastest safe fix to stop the harm, and commit the deeper one with an owner and a date.** That is not a compromise, it follows directly from the axes. Harm is happening now, so buy the thing that pays back now. The family is still open, so buy the thing that pays back at the next release, on a schedule somebody owns. What makes this honest rather than a way of never doing the deeper fix is the record. The defect says which layer actually shipped, what the family still admits, and where the deeper change is tracked. Without that sentence the two-move plan quietly degrades into one move and an intention. ## Two failure modes to avoid - **Always deepest.** Every defect becomes a redesign, delivery slows to a crawl, and users keep being hurt while the perfect fix is built. - **Always narrowest.** The code fills with special cases, each a small tax on every future reader, and no family is ever closed. Both are recognisable by the absence of a decision. Nobody weighed reach against cost; a habit was applied. ## Arguing for the expensive layer One occurrence almost never justifies weeks of work, and arguing from the drama of that occurrence is how these proposals lose. What does justify it: - **The size of the family** the weakness admits, enumerated rather than asserted. - **The worst reachable member** of that family: the same missing guard on a path that touches money, safety or personal data. - **The drag already being paid**: the workarounds in the code, the guards other teams wrote for the same weakness, the changes that take longer because everybody is careful around this area. Framed that way it is a capacity argument rather than a plea, and it can be scheduled against other work instead of competing with it emotionally. ## When the deepest layer is out of reach Sometimes it genuinely is. The weakness lives in a component another team owns, in a supplied dependency you cannot change, or in a stored structure with years of history behind it. The move then is **containment, stated as containment**: 1. **Guard on your side of the boundary**, so the condition cannot reach the weakness through any path you control. 2. **Make recurrence loud.** A check that fails, an alert, an assertion, something that tells you the day the condition returns rather than the day a user notices. 3. **Record the truth.** The family is contained, not removed. Say what remains reachable and who owns it. 4. **Hand the owner the evidence.** The occurrence, the family you enumerated and the containment you had to build make a far stronger case than a request does. Containment is a legitimate destination and often the only available one. Silently treating containment as removal is not, because the next person reads a closed record, assumes the structure underneath is sound, and builds on it.
- The deepest fix sits in a component your team does not own. What do you do?Contain it and make it visible. Guard on your side so the condition cannot reach the weakness through any path you control, add a check that fails loudly if it recurs, and record that the family is contained rather than removed. Then take the occurrence, the enumerated family and the containment you had to build to the owning team as evidence.
- How do you argue for the expensive layer when the defect has occurred only once?Not on the single occurrence. Argue on the size of the family the weakness admits, on what the same weakness would cost in the worst place it is reachable from, and on the drag already being paid in workarounds and slow changes around that area. One occurrence rarely buys weeks; an enumerated family touching every path into a component sometimes does.
- What tells you a team is defaulting to a layer rather than choosing one?The reasoning is missing from the records. Every defect closes with a special case, or every defect turns into a redesign proposal, and in neither case does anyone write down what the change covers versus what it cost. A team that is choosing leaves a trail: reach, effort and what remains open, on the record.
saying these in an interview costs you the question
- Always reaches for the deepest fix regardless of cost
- Always ships the narrow fix and never revisits it
- Judges the layer by effort alone, ignoring reach
- Treats an unreachable layer as a reason to do nothing
- Closes the record when only the narrow fix shipped