A marketplace bidding engine has produced silent data corruption in an area your risk register ranked low. How do you re-rank mid-release?
answer
- The register is not fixed at planning
- A realised failure is likelihood evidence
- Silence raises impact through detection time
- Re-score the neighbours, not one row
- Fund new depth by cooling something
basics
~20 sTreat the incident as evidence about the ranking method, not just one row. Raise that item on both axes, re-score every item resting on the same assumption, and fund the new depth by cooling an area the evidence has cleared.
solid answer
~50 sFirst fix the item's own scores: a realised failure proves the likelihood estimate wrong, and silent corruption usually means impact was under-scored too, because undetected wrong data spreads and recovery is a backfill, not a retry. Then ask what the miss says about the **method**. If the low rank came from an assumption like 'this subsystem is stable, so likelihood is low', every other item resting on that assumption is mis-scored, and re-scoring the cluster together is worth more than re-scoring the one row. Fund the new depth explicitly by cooling an area the evidence has cleared, rather than asking for time you will not get. Because the failure is silent, add detection alongside tests — a reconciliation check that compares totals against an independent source gives you an oracle where none existed. Finally, record the re-rank and its evidence so the trail survives the release.
go deeper
Know that a risk ranking is an estimate that gets revised, and that a failure found in an area you ranked low is information rather than an embarrassment. Be able to say what you would raise on the register after such a find.
Explain which evidence moves which axis: a realised failure and defect clusters move likelihood, while time-to-detection and blast radius move impact. Show that re-ranking must change planned depth or it changed nothing.
Demonstrate the cluster instinct — finding every item that rested on the same broken assumption — and the funding move, where new depth comes from an area evidence has cleared rather than from a request for extra days.
Own the feedback loop across releases: how re-ranks are recorded, how last release's misses inform the next scoring workshop, and how you prevent the register from becoming a document that only ever grows and is never revisited.
## Why the register is expected to move A risk ranking made at planning time is a set of estimates made with the least information anybody will have all release. Execution, defect data and real usage are all information about likelihood, and treating the register as a document fixed at kick-off is the single most common way risk-based selection quietly stops working. Re-ranking is not an admission that the plan was wrong; it is the plan working as designed. ## What this particular incident tells you Take a marketplace bidding engine that has written corrupted values for a stretch of bids without any failing check. Three separate corrections follow. **Likelihood was wrong, and you now know it empirically.** An estimate has been replaced by an observation. Whatever score the area carried, it now carries a higher one — not as a punishment, but because a demonstrated failure is the strongest likelihood evidence available. **Impact was probably wrong too.** Impact scoring tends to imagine a loud failure. Silence changes the arithmetic in three ways: the blast radius grows with time-to-detection, downstream systems consume the bad values and propagate them, and the recovery is a data-repair exercise with its own risk rather than a redeploy. If the register's impact descriptor never mentioned detectability, that is a gap in the scale, not just in this row. **The area's neighbours are now suspect.** This is the part that separates a senior answer. Ask *why* the item was ranked low. If the reason was "the settlement path has been stable for two years", then every item scored low for the same reason inherits the doubt. Re-score the cluster: items sharing the component, the assumption, the author, or the same silent-write shape. ## Gathering the evidence properly Re-ranking on one incident and vibes is barely better than not re-ranking. Cheap evidence worth pulling in the same afternoon: - **Defect clustering.** Where have the release's defects actually landed, by component and by cause? Clusters relocate likelihood far better than opinion does. - **Real usage.** Which flows carry the volume, and which parameter values dominate? A path that the plan treated as an edge case may be the mainline. In the bidding engine the corruption appeared for bids sitting at a 92nd-percentile budget — a value the plan had treated as an outlier, while the traffic showed a heavy shoulder there. That single fact moves both axes: the input is common, so likelihood rises, and it belongs to the highest-value bidders, so impact rises. - **Change history.** Which areas moved late in the release, and did anything touch the corrupted path recently? - **Support and operations signal.** Quiet defects are often visible first as odd manual corrections rather than as tickets. ## Reallocating, not just re-ranking A re-ranked register that does not change what anybody does is theatre. The credible move is a swap, stated out loud: *the settlement path moves to the top band and gets boundary work around the budget distribution plus an exploratory session against real value shapes; the notifications area, ranked high at planning because it was new but now two cycles clean with no defects, drops a band and releases that time.* This is a decision the release owner can accept or overrule, and it does not depend on winning an argument for extra days. ## Add an oracle, not only cases Silent corruption is a detection failure as much as a correctness failure: even a strong test can pass when nothing distinguishes a plausible-looking wrong value from a right one. So part of the response belongs outside the case list — an independent check that recomputes or reconciles the value from a different source and shouts on divergence. This does double duty: it acts as an oracle for the tests you are about to write, and it survives into production as a monitor, which lowers the impact score of the whole family of risks by shrinking time-to-detection. ## Close the loop on the record Write the re-rank into the register with a date, the evidence that caused it, the depth traded away and the person who agreed the trade. Two things depend on that line existing. First, the retrospective question "why did we rank this low?" has an answer that is not memory. Second, the next release's ranking starts from what was learned rather than from the same stale assumption — which is the only mechanism by which the register gets better over time. ## What a weak answer looks like Re-scoring only the failing row; demanding extra time as the sole response; treating the incident as a defect to be fixed rather than as evidence about the model; and, most commonly, promising a fuller register next release instead of changing anything in this one.
- How do you decide which area to cool so the re-ranked area can get more depth?Look for items whose original score rested on uncertainty that evidence has now removed: an area ranked high because it was new, which has since run two clean cycles with no defects and no late changes, is the strongest candidate. State the trade explicitly to whoever owns the release rather than absorbing it quietly, because the cooled area's exposure is going up and that should be a decision someone made, not a side effect of your calendar.
- How often should the ranking be revisited, and what triggers a re-rank outside that rhythm?A light review each iteration or cycle boundary is usually enough for the routine drift. Out-of-band triggers are events that invalidate an assumption rather than events that are merely bad: a failure in an area scored low, a late architectural or dependency change, a shift in real usage that moves the traffic, a new regulatory constraint, or the loss of the person whose knowledge justified a low likelihood score.
- The corrupted values are already in stored data. Does that change the risk ranking or only the fix?Both. Existing bad data raises the impact score for the whole family, because the exposure is no longer bounded by future writes and the repair itself carries risk of making things worse. It also creates a new risk item for the repair — a one-off correction run against live data deserves its own likelihood and impact scores and usually its own verification, which teams routinely forget because it feels like remediation rather than change.
saying these in an interview costs you the question
- Re-scores only the row that failed
- Asks for more time as the only response
- Says the register is fixed after planning
- Ignores detection time when scoring impact
- Adds tests but no independent check
- Leaves the re-rank undocumented and unowned