The dropout early-warning model is switched off mid-term because every version reads the same broken term calendar - what serves advisers instead?
answer
- leave the model path, not the version
- no shared dependency with the failure
- same contract, same weekly volume
- worse recall, explainable every time
- cap at adviser capacity, rank below it
basics
~20 sA rules-based path built from inputs the broken one does not touch: a small deterministic score over recent attendance and failing grades, capped at the same weekly caseload and published in the same output contract, with the quality loss stated in advance.
solid answer
~40 sFalling back to the previous model version is useless when the bad input feeds every version, so the fallback has to leave the model path entirely. Build a deterministic rules path - a handful of thresholds over attendance in a rolling window and current failing grades - and hold three properties. It must not read the dependency that is suspected: a heuristic computed from the same broken calendar is not a fallback. It must emit the **same output contract and the same weekly volume**, so the advising workflow downstream does not also break. And its quality loss must be named before the incident, not discovered during it: it finds the obvious cases and misses the subtle ones the model was bought for. That loss is the price of correctness you can explain.
code
pseudocode · 21 lineson weekly_adviser_run(week):
if kill_switch.engaged("dropout-early-warning"):
source = "rules"
else if model_scores.age(week) > 8 days:
source = "rules"
else:
source = "model"
if source == "rules":
scored = []
for each student in enrolled(week):
risk = 0
if attendance_rate(student, last_4_weeks) < 0.70: risk = risk + 2
if failing_grades(student, current_term) >= 1: risk = risk + 2
if unpaid_balance_blocks_enrolment(student): risk = risk + 1
if risk >= 2: scored.append(student, risk)
else:
scored = model_scores(week)
caseload = take(sort_desc(scored, by = risk), WEEKLY_CAPACITY)
publish(caseload, contract_version = 3, source_tag = source)go deeper
Recall that a system depending on a model needs a non-model path to serve from, and that the path has to produce the same shape of output.
Explain the difference between reverting to a prior model version and leaving the model path entirely, and when only the second one mitigates anything.
Show that you would check the fallback shares no dependency with the suspected cause, hold the output contract and volume fixed, and state the quality loss up front.
The judgment is how much detection quality the organisation will trade for explainability, and for how long, before the degraded path becomes the status quo nobody revisits.
## Why a version rollback does not always help There are two different things people call "rolling back" an ML system, and the incident decides which one you need. **Reverting to the prior model version** undoes a change your team made. **Falling back off the model entirely** undoes your dependence on the model at all. When the term calendar changes meaning under the feature pipeline, every retained version reads the same corrupted input and every one of them is wrong in the same way. The pointer flip is motion without mitigation. So the kill switch's job is not to select a different model. It is to select a **different path** - one whose correctness does not depend on whatever you have not diagnosed yet. ## What the fallback path has to satisfy - **No shared dependency with the suspected cause.** A rules path that derives "weeks into term" from the same calendar feed inherits the failure. Compute over rolling windows (the last four weeks of attendance) rather than term-relative ones wherever you can. - **Same output contract.** The adviser tool, the weekly export and any reporting downstream parse a shape. Changing that shape during an incident turns one failure into two. - **Same volume.** Advisers have a fixed weekly capacity. A fallback that ranks by a crude score and publishes everyone above a threshold can triple the caseload, which fails the team as surely as a wrong list. Cap it at the capacity the model path was capped at. - **Deterministic and explainable.** The reason a rules path is acceptable at all is that anyone can say why a student is on the list. During an incident that explainability is the product. - **Independently sourced, independently monitored.** Its inputs need their own freshness check, or you have swapped a silent model failure for a silent rules failure. - **Engaged by a flag, not a deploy.** If switching paths requires a release, the fallback is not a switch. ## The quality you agree to lose Write this down before you need it, because arguing about it during an incident guarantees the switch does not get pulled. | Property | Model path | Rules fallback | |---|---|---| | At-risk students found | Includes subtle, early cases | Obvious cases only, found later | | False positives | Tuned against a target caseload | Higher, and unevenly distributed | | Explanation to an adviser | Indirect, needs supporting detail | Direct - the rule that fired | | Behaviour under a bad input | Silently wrong | Wrong only if that input feeds a rule | | Cost of being wrong | Hard to detect | Visible and arguable | The honest framing is that the fallback is **worse on the metric the model exists to move and better on the property the incident took away**: you can say why every name is on the list. An organisation that has agreed to that trade in advance can run on the fallback for weeks; one that has not will spend the incident relitigating whether to switch at all. ## Holding the caseload constant The volume cap does more work than it looks like it does. It converts an arbitrary threshold problem into a ranking problem: instead of asking "what attendance rate means at risk", the rules path scores everyone, sorts, and takes the top N where N is the advising team's weekly capacity. The threshold then only has to be roughly right, because the cap absorbs its error. It also makes the switch's effect legible to everyone downstream - the list is the same size on Monday whether the model produced it or the rules did, and the only visible change is a source tag on the run. ## Coming back The path back is the mirror image of the switch and needs its own decision rule. Restore the model only when the suspected input is verified good, a cycle has been scored under the restored configuration and compared against the fallback's list, and someone has decided whether the weeks served by the fallback need re-scoring. Re-enabling silently, because the alarm stopped firing, is how the same incident recurs with nobody watching for it - which is why the on direction is gated more tightly than the off direction.
- Why cap the rules path at the same weekly caseload instead of publishing everyone above its threshold?Because advising capacity is the real constraint and a crude threshold has no idea where it sits. Ranking and taking the top N makes the threshold's error harmless - it only has to order students roughly correctly - and it keeps the switch invisible to everyone downstream except as a source tag. Publishing an unbounded list swaps a quality incident for a capacity one.
- The fallback reads one feature the model also reads. Is that disqualifying?Only if that feature is the suspected cause or shares its source. Total independence is rarely achievable and not the goal; independence from the failing dependency is. The practical rule is to enumerate the fallback's inputs at design time and check none of them traces back to the pipeline you are switching away from - then re-check that list whenever either path's inputs change.
- How long can the system reasonably run on the rules path?As long as the stated quality loss stays acceptable, which is usually weeks rather than hours - the fallback is a degraded service, not an outage. What forces a deadline is drift in the rules themselves: thresholds calibrated on an older cohort slowly mis-size the caseload, and the cap hides that until someone compares the fallback's list against what the model would have produced.
A kitchen whose supplier fails plates a short standing menu it can cook entirely from the stock room. Fewer dishes go out, and the point is that none of them depends on the delivery that failed.
saying these in an interview costs you the question
- Falling back means loading the previous model version
- A fallback computed from the same feed is still a fallback
- The fallback must match the model's precision to be usable
- Publish every student the rules flag, whatever the volume
- Switching paths can wait for the next release