A student dropout early-warning model starts flagging the wrong cohort while error rates and latency stay flat - how do you size this incident's severity?
answer
- nothing errored, so nothing paged
- impact is decisions, not requests
- count what already went out
- start is earlier than detection
- irreversibility drives the escalation
basics
~20 sSeverity is the count of wrong decisions already delivered and acted on, not a failure rate. Bound the start at the last known-good change, count the flags published since, and escalate on how irreversible the adviser action was.
solid answer
~50 sNothing failed, so the availability signals do not move: every request returned a well-formed score inside its budget. Severity therefore cannot come from a failure rate. The impact number is the volume of decisions already published times what was done with each one - a name on an adviser's weekly list is recoverable, a student moved onto a remediation track or told they are at risk is much less so. Because the failure is silent, the incident is older than its detection, so you bound the start with the last known-good change to inputs, features or configuration and count every flag published since. Two properties push severity up that an availability matrix has no column for: the detection lag is unbounded until you find the start, and exposure keeps growing after you mitigate, because the lists already sent are still being worked.
go deeper
Remember the core fact: a model can return a perfectly valid answer that is wrong, so success rate and latency say nothing about whether its predictions are right.
Be able to explain why the usual severity dimensions read green here, and what replaces them: the count of decisions already published rather than the share of requests failing.
Show that you would bound the start at the last known-good change, count the delivered flags, and weigh how reversible the downstream action was before naming a severity.
The open call is how much silent wrongness the organisation accepts before an automatic shutoff fires without a human, and who carries the cost of the false shutoffs that policy buys.
## The failure that returns a valid answer An early-warning model scores every enrolled student each week and hands advisers a ranked caseload. A change to the term calendar shifts what "week 3" means, so the features counting attendance and assessment events now describe a different slice of time than the ones the model was trained on. The model is given a well-formed input and returns a well-formed score. Every request succeeds, the weekly job finishes on time, and every downstream consumer of the list is happy. What has changed is that the scores are about the wrong students. This is the shape of ML incident that the outage everyone has rehearsed does not prepare you for. Three properties follow from it: 1. **The availability signals stay flat.** Request success rate, tail latency, saturation and queue depth all report health, because the system is doing exactly what it was asked to do. 2. **The incident is older than its detection.** A wrong answer does not announce itself. The gap between the input change landing and someone noticing is the *detection lag*, and until you find the start you do not know its size. 3. **Exposure keeps growing after you mitigate.** Switching the model off stops new wrong flags. It does not recall the three weekly lists advisers are already working through. ## Why the usual severity dimensions read green | Dimension a severity matrix usually asks for | What it reports in this incident | Why it misleads | |---|---|---| | Share of requests failing | Zero | Wrong answers are successful answers | | Users affected right now | Zero sessions erroring | The harm landed in a batch, days ago | | Time to restore service | Minutes - flip a switch | The service was never down | | Blast radius by component | One scoring service | The radius is decisions, not components | Sizing from these produces a low severity for an incident that may have mis-directed a term's worth of advising capacity. The matrix is not wrong in general; it is measuring concurrency, and this failure's impact is not concurrent. ## The numbers that do size it Three quantities, in the order you can usually get them: 1. **Volume delivered.** How many scored students were published on a list since the suspected start. If a decision record exists per published flag, this is a query; if not, you are reconstructing it from the weekly exports, and that reconstruction gap is itself a finding. 2. **Churn.** Of those, the share whose flag would have been different under the last known-good configuration. A calendar shift that moves a handful of borderline students is not the incident a wholesale cohort swap is. 3. **Irreversibility.** What was done with each flag. Advising work splits into steps that are cheap to undo (a name on a list nobody reached yet), awkward to undo (a conversation that told a student they were at risk), and effectively permanent (an enrolment change, a referral, a record another system now carries). Severity is the product of volume and irreversibility, not the size of either alone. ## Pulling the start time backwards You cannot count what you cannot bound, so finding the start is part of triage rather than part of the postmortem: - Take the **last known-good change** to any input the model reads - upstream schema, calendar, feature definition, model version, configuration - and treat everything after it as suspect. - Compare the **published flag volume and its composition** against the same weeks in prior terms. A silent failure that swaps a cohort usually moves the composition long before anyone reports it. - Look for a **caseload that stopped moving**: if the same students recur week after week, or the overlap between consecutive weekly lists jumps, something upstream froze or shifted. - If a monitoring alarm fired earlier and was dismissed, its timestamp is a candidate start, not a nuisance. ## What the severity is actually for Severity decides three things here, and each is ML-specific. It decides whether you switch the model off now or wait for a cause, which is a decision about the cost of a wrong answer against the cost of losing the model entirely. It decides whether the already-delivered decisions need **active remediation** - re-running the affected weeks under the restored configuration and telling advisers which names to drop - or whether letting the next cycle correct them is enough. And it decides who is woken: a silent quality failure that has been running for three weeks rarely needs a 02:00 page, while one on a path that takes irreversible action does. The habit to build is stating the impact as a sentence with a number in it: *"roughly 400 students have appeared on adviser lists under the wrong cohort mapping since the calendar change nine days ago, of whom about 60 have had a conversation."* That sentence is the severity. A percentage of failing requests, in this incident, is not.
- A drift alarm from monitoring crossed its threshold overnight. Is that alarm the page?No. A drift alarm is a signal that an input or prediction distribution moved; on its own it fires on seasonality, on a marketing push, on a new intake. The page needs a decision rule laid over it: the alarm plus an independent corroborating signal - published flag volume outside its baseline band, or a decision rate that jumped - plus a stated action for whoever answers. Monitoring owns producing the statistic; this layer owns deciding it is an incident.
- The team suspects the model is wrong but cannot prove it. Do you pull the switch before you have a cause?Usually yes, if the downstream action is hard to undo. Every cycle you spend proving the cause delivers another batch of suspect decisions, while the fallback path's cost is bounded and known: fewer at-risk students found, none of them wrong for a reason you cannot explain. Pull first, diagnose with the traffic off. The judgment flips when the fallback would itself flood the advising team.
- How does this incident differ from the same model returning errors for every request?An erroring model is a conventional outage: it is loud, its start is timestamped, its blast radius is the requests that failed, and the downstream product usually degrades visibly. The silent version inverts all four - it is quiet, its start must be inferred, its radius is decisions already consumed, and the product looks normal throughout. Only the second one needs this playbook.
saying these in an interview costs you the question
- Flat error rate and latency mean the model is healthy
- Severity is the share of requests failing right now
- The incident started when the alert fired
- Switching the model off ends the harm already delivered
- No user complained, so no incident should be declared