How do you set the abstention threshold for a chest-radiograph triage model that routes uncertain cases to a radiologist?
answer
- two numbers, never one
- trade retained-case error against workload
- the humans have finite reading capacity
- check the deferral rate per subgroup
- rising deferrals warn of drift before labels do
basics
~20 sTreat it as a coverage-versus-risk trade, not a modelling choice. Sweep the threshold on held-out data, then pick the point where the deferred volume fits real radiologist capacity and the residual error is acceptable — and check that trade per subgroup.
solid answer
~50 sThe threshold is a business decision expressed in a model quantity. Rank held-out cases by the uncertainty signal and sweep the threshold to trace risk against coverage: as you defer more, accuracy on retained cases rises and the human workload rises with it. Two constraints bound the choice — the residual error the clinical owners accept on autonomous cases, and the studies radiologists can actually re-read per day. If the point satisfying the first exceeds the second, the model is not ready for that workflow, and no threshold hides that. Then check the trade per subgroup: a global 5% deferral rate can be 40% at one site, silently withdrawing the service from one population. In production, monitor deferral rate at a fixed threshold as a drift alarm, and audit deferred cases so abstention never becomes a way of hiding hard cases from the metric.
go deeper
Know what abstention means: the system declines to predict on some inputs and passes them to a human. Remember that the fraction it answers, called coverage, must always be reported next to its accuracy.
Explain how the threshold traces a trade: as more cases are deferred, error on the retained cases falls and human workload rises. Be able to describe building that curve on held-out data.
Show operational judgment: capacity limits, subgroup deferral rates, and monitoring the deferral rate as an early drift alarm. Be ready to say when a model is simply not ready for an autonomous path.
Own the threshold as a contract term rather than a hyperparameter. Drive the stakeholders to state an explicit error tolerance per error type, decide whose fairness objective governs subgroup thresholds, and account for the joint human-plus-model service.
## Abstention is a decision layer, not a model layer A model that can say *I do not know* is doing selective prediction: for each input it either predicts or defers. Two numbers describe such a system, and they must always be quoted together. - **Coverage**: the fraction of inputs on which it predicts. - **Selective risk**: the error rate *among the covered cases only*. The threshold on the uncertainty signal moves you along a curve trading one against the other. Raise the bar and coverage falls while selective risk falls with it, because the cases you shed are disproportionately the ones you would have got wrong. Report either number alone and you have said nothing: 99% accuracy at 20% coverage is a system that answers one study in five. ## Constructing the curve On a held-out set that resembles deployment, score every case with the uncertainty signal, sort by it, and sweep the cut. At each cut record coverage and the error rate on what is retained. This risk-coverage curve is the only honest artefact for the conversation, and it is the thing to bring to the clinical stakeholders rather than a single proposed number. Two properties of the curve matter as much as any point on it: how steeply risk drops at low deferral (a good uncertainty signal sheds errors fast), and where it flattens (past that point you are deferring cases you would have got right, buying workload for nothing). The curve must be built on data the threshold was not chosen on, and on data drawn from the deployment population — a curve traced on a curated research set will be optimistic on the mixture of portable, repeat and poorly positioned studies a real queue contains. ## Picking the point The choice is bounded from two sides. **From above, by clinical tolerance.** Somebody must state the residual error rate acceptable on cases the system handles without a human, and it is asymmetric: a missed finding and a false alarm are not the same cost, so the tolerance is usually stated per error type. That statement is not the machine learning team's to make; the team's job is to force it to be stated explicitly. **From below, by capacity.** Every deferral is a study a radiologist must read. Departments have finite reading capacity and it is already fully committed. If meeting the error tolerance requires deferring 45% of studies and the department can absorb 10%, the honest report is *this model is not deployable in this workflow yet*, with the risk-coverage curve as the evidence. Choosing a threshold that meets capacity while violating the tolerance is the failure mode to refuse. ## What the aggregate curve hides A single global threshold produces different deferral rates in every subgroup, because uncertainty is not uniformly distributed. Portable equipment, one imaging site, paediatric studies, or an under-represented condition can sit far above the global rate. Two harms follow. Operationally, one site drowns in deferrals while another sees none. Ethically, if abstention concentrates on one population, the automated benefit is withdrawn precisely from the group already least well served — and the aggregate coverage number conceals it perfectly. Trace the curve per subgroup, and be prepared to argue for group-specific thresholds or for restricting deployment to the subgroups where the trade actually holds. ## Living with it in production Three monitoring commitments make abstention safe rather than cosmetic. 1. **Deferral rate at a fixed threshold is a drift alarm.** It is often the earliest available signal, because it moves without needing any labels: a new scanner, a protocol change or a seasonal case-mix shift raises uncertainty long before anyone has outcome data to measure accuracy against. Alert on it. 2. **Audit the deferred cases.** Abstention can quietly become a way of removing hard cases from the headline metric. Sample the deferred set and check that it is genuinely harder for the model — and check what the humans did with it. If the model defers exactly the cases the readers also find hardest, it is adding review load without adding safety. 3. **Measure the joint system, not the model.** The deliverable is the model-plus-human pipeline. Its metrics are end-to-end accuracy, total human minutes consumed, and turnaround time. A model that improves selective risk while doubling reading time has not improved the service. ## The framing to own The recurring senior mistake is to treat the threshold as a hyperparameter to tune against a validation metric. It is a contract term: it encodes how much residual error the institution accepts and how much human work it is willing to spend to avoid the rest. The technical contribution is producing a trustworthy uncertainty ranking and an honest risk-coverage curve; the choice of operating point belongs to the people who carry the clinical and operational risk, and the discipline is to make them choose it explicitly rather than to inherit a number nobody agreed to.
- What would you put on the production dashboard for this system?Coverage and selective risk together, never either alone, plus the deferral rate broken out by site, equipment and patient subgroup. Add total reader minutes consumed and end-to-end turnaround, because the deliverable is the combined pipeline. Alert on a moving deferral rate at fixed threshold, since it detects shift without waiting for outcome labels.
- The error tolerance can only be met by deferring 45% of studies. What do you report?That the model is not deployable in this workflow, and show the risk-coverage curve as the evidence. Then offer the alternatives explicitly: narrow the scope to the subgroups where the trade does hold, use the model to reorder the queue rather than to decide, or invest in data from the regions driving the deferrals. What you do not do is move the threshold to fit capacity and quietly ship the higher error rate.
- Should the abstention threshold be the same for every subgroup?Not automatically. One global threshold yields wildly different deferral rates across sites and populations, and the aggregate number hides that. Per-subgroup thresholds can equalise either the deferral rate or the residual risk, but not both, so somebody must choose which is the fairness objective. That choice needs the clinical owners in the room, and whichever way it goes it should be stated and monitored rather than left implicit.
saying these in an interview costs you the question
- Quotes accuracy on retained cases without quoting coverage
- Tunes the threshold as a validation hyperparameter
- Ignores whether reviewers can absorb the deferred volume
- Never checks deferral rate by site or subgroup
- Treats abstention as always safe regardless of rate