With one sprint of capacity, do you fix the likely-but-bounded threat or the remote-but-catastrophic one?
answer
- The rating is an input, not the decision
- Fix cost is on neither axis
- Different shapes, different response types
- Attack the assumption under a low likelihood
- Acceptance needs an owner and an expiry
basics
~20 sReject the framing that one rating decides both. Ship the cheap fix for the likely bounded threat this sprint, and treat the catastrophic one as an architecture decision with a named owner, a date and a written acceptance.
solid answer
~50 sTwo threats: a course-registration service whose bulk-enrolment endpoint accepts an unbounded date range, so any student can exhaust it during the two-day enrolment window; and a container-terminal crane telemetry relay, reachable only from the port's operational network, where forged readings could endanger a physical process. The first is near-certain and bounded; the second is remote and catastrophic. I do not average them into one number, because the number hides the shape and both responses would be wrong. The enrolment bound is a small, well-understood change with a hard deadline attached to it, so it ships this sprint on cost alone. The crane path is not sprint work — it needs authenticated telemetry, segmentation and plausibility checks, which is a design commitment with an owner and a date. Meanwhile I challenge the likelihood hardest there, because "only from the operational network" is an assumption, not a property.
go deeper
Be ready to see that a common threat with limited damage and a rare threat with enormous damage are not the same kind of problem, even when a rating puts them close together.
An interviewer expects you to keep fix cost separate from the rating and to notice that a cheap, contained change with a calendar deadline should not be queued behind a redesign.
Demonstrate that you attack the assumption underneath a low likelihood, especially when the impact is catastrophic, and that you can propose interim detection or limiting measures while the real fix is scheduled.
Own the routing and the accountability. Show how you keep three response lanes visible in one prioritized list, and define what a legitimate risk acceptance requires: a named owner, an expiry, a compensating measure and a written, testable assumption.
## The question behind the question An interviewer asking this is not looking for a winner. They are checking whether you understand that a likelihood-impact rating is an input to a decision and not the decision itself, and whether you know that the two shapes call for different *kinds* of response rather than different positions in one queue. ## Separate the rating from the fix decision Nothing in either axis mentions cost, and cost is decisive at sprint granularity. Bounding a date range on the enrolment endpoint is a contained change with a test; nobody should be arguing about its rank against anything, because the correct comparison is against its own trivial cost and a hard calendar deadline. Meanwhile the crane relay's fix is not a task — signed or authenticated telemetry, network segmentation and a fail-safe on implausible readings is a design programme. Putting it on the same list as a parameter check invites the worst outcome available: the sprint spends its capacity on a fraction of the big job, ships nothing usable, and the small certain problem is still there when enrolment opens. So the answer is that both move, in different lanes: - **Sprint fix**: the enrolment bound, plus any cheap partial measure available on the crane path — rate limiting, a plausibility ceiling on readings, alerting on telemetry from an unexpected source. - **Scheduled design work**: the telemetry authentication and segmentation, with a named owner and a date. - **Explicit acceptance**: whatever exposure remains between now and that date, written down with the assumption it rests on. ## Challenge the low likelihood hardest where impact is catastrophic A remote-but-catastrophic rating is only as good as the assumption underneath it, and here the entire rating rests on one clause: the relay is reachable only from the operational network. That clause is invalidated by a vendor maintenance link, an engineering laptop that also reaches corporate email, a flat VLAN nobody has re-checked since commissioning, a contractor with physical access, or a future integration that quietly bridges the two networks for a reporting dashboard. Rare, high-consequence estimates are the least reliable ones in any model, because there is no experience to calibrate them against and the failure modes are correlated in ways nobody enumerated. So the deliverable is not just a rating, it is a **written assumption with a test**: state that the rating depends on network isolation, and put a periodic check of that isolation next to it. There is also a floor rule worth stating in an interview: where the credible worst case is physical harm to people, likelihood does not license a shrug. Those threats get escalated out of the engineering queue regardless of how the arithmetic lands, because the organisation's tolerance for that outcome is not an engineering parameter. ## Why not just combine the two ratings A single combined score sorts the queue, which is exactly the problem: it makes two threats that need completely different responses look like neighbours on one list. Collapsing them loses the information a decision-maker needs — whether the risk is a steady drip you can bound with a code change, or a rare event that would end a business line. Present the rating, but present it with the shape and the response type attached. One list is fine for a director who wants one list; it just needs three lanes in it. ## Who decides, and what acceptance means An engineer or a lead can decide to defer a bounded availability threat. Nobody in the engineering organisation should be quietly accepting a life-safety or company-ending risk on their own authority — the point of surfacing the shape is to route the decision to the person whose job it is to carry it. A legitimate acceptance has four parts: a named accountable person, an expiry or review date, a compensating measure in the meantime, and the written assumption that justified the low likelihood so it can be re-tested. An acceptance without those is not a decision, it is the threat quietly leaving the model. ## What a weak answer looks like Multiplying the two ratings and fixing the larger product; or the opposite reflex, always dramatising the catastrophic one and letting the certain, cheap, deadline-bound fix slip. Both replace judgment with a rule, and both are visibly wrong the moment the enrolment window opens or someone plugs a laptop into the wrong switch.
- What would change your mind about the crane relay's low likelihood?Anything that dissolves the required-access assumption it rests on. A vendor maintenance link or engineering laptop bridging corporate and operational networks, a flat segment nobody has verified since commissioning, contractor physical access to the plant, or a new reporting integration that reaches into the operational side. Required access was the entire basis for calling it remote, so I would write that assumption next to the rating and schedule a periodic check of the isolation rather than trusting it indefinitely.
- A director wants a single prioritized list. How do you present these two without flattening them?Give the single list, but annotate each row with its shape and its response type: fix this sprint, scheduled design work with an owner and a date, or accepted with an expiry. The ranking then carries the information the ranking alone would have destroyed, and the director can see that the top item is a two-hour change while the second is a quarter of engineering work. One list, three lanes.
- Is accepting a catastrophic-impact threat ever legitimate?Yes, but only as an explicit decision with four parts: a named accountable owner senior enough to carry it, a review or expiry date, a compensating measure running in the meantime, and the written assumption that justified the low likelihood so somebody can re-test it. Where the worst case is harm to people, the acceptance is not an engineering decision at all and must be escalated. An acceptance missing those parts is just the threat leaving the model quietly.
saying these in an interview costs you the question
- Multiplies the two ratings and fixes the higher product
- Treats reachable only from the internal network as permanent
- Accepts a life-safety risk on an engineer's own authority
- Delays a two-hour fix because a scarier threat exists
- Rates likelihood low because it has never happened here
- Puts a design programme and a parameter check on one queue