skip to content

How do you set the coverage level for a skin-lesion classifier's conformal prediction sets in a triage workflow?

level: principalimportance: nice to knowfreq 18%

answer

  1. alpha follows the cost of a miss
  2. tighter coverage buys larger sets
  3. set size varies with case difficulty
  4. singleton, ambiguous pair, empty: three routes
  5. coverage of the set is not top-1 accuracy

basics

~20 s

Choose the miscoverage level from the cost of missing the true diagnosis, not from convention. Higher coverage means larger label sets and more escalations, and the workflow must handle sets of any size, including two labels or none.

solid answer

~50 s

In classification, conformal outputs a *set* of labels rather than an interval: score each label by one minus its predicted probability, calibrate a threshold on held-out data, and include every label clearing it. Set size then varies with case difficulty. Setting alpha is therefore a decision about the cost of a miss: at 90% coverage one case in ten excludes the true diagnosis, which for melanoma is likely unacceptable, and 99% is defensible even though sets grow. Quantify the price before choosing — measure average set size and the share of single-label sets at each candidate alpha, and translate those into escalation volume and reviewer hours. Then commit the workflow: a singleton can auto-route, a set containing both melanoma and nevus means the model cannot separate them and belongs with a dermatologist, and an empty set means nothing cleared the bar and must never render as a negative result.

go deeper

for a junior

Know that a conformal classifier can return several labels or none, and that the number of labels reflects how hard the case is. A set of two is not an error message.

for a middle

Explain how the set is formed: score labels by one minus predicted probability, calibrate one threshold on held-out data, keep every label above it. Note that the classifier's probabilities need only rank labels, not be calibrated.

for a senior

Show how you would quantify the alpha choice — average set size, singleton share and empty-set share measured per candidate level on fresh data — and how each set size maps onto a concrete routing decision in the workflow.

for a principal

Own the tradeoff as a risk and capacity decision made with clinical stakeholders, plus the commitments it creates: recurring recalibration, per-cohort coverage reporting, an interface that renders abstention honestly, and re-approval whenever the pipeline changes.

## How a prediction set is built The conformal machinery is unchanged from regression; only the nonconformity score differs. A common choice scores each calibration row by `s = 1 - p(true label | x)`, where `p` is the model's predicted probability for the label that actually occurred. Sort those scores, take `q` at the rank `ceil((n+1)(1-alpha))`, and for a new image output the set of every label whose predicted probability is at least `1 - q`. The guarantee carries over: the set contains the true label with probability at least `1 - alpha`, marginally, in finite samples, for any classifier. Two consequences follow immediately. First, **set size is not fixed**. An easy lesion produces one label; an ambiguous one produces two or three; a case unlike anything in training may produce none. That variation is information, and it is the main reason to prefer sets over a bare top-1 label. Second, the procedure does not require the probabilities to be well calibrated — it only uses them to rank labels against a single learned threshold. Better probabilities give smaller sets, not more valid ones. ## Alpha is a cost decision, not a convention The default reflex is 90% or 95% because those numbers are familiar. They carry no clinical meaning. The right question is what a miss costs relative to an escalation. - **Cost of a miss.** At `alpha = 0.10`, one case in ten has a set that excludes the truth. If the excluded truth can be melanoma, that rate is hard to defend, and pushing to `alpha = 0.01` is reasonable even though it multiplies set sizes. - **Cost of coverage.** Higher coverage widens sets. More multi-label sets means more cases a human must adjudicate. Past some point the sets contain most of the label space and the system tells the clinician nothing while consuming their time. The deliverable is a curve, not a number: for each candidate alpha, measure on a fresh held-out sample the average set size, the proportion of singleton sets, and the proportion of empty sets. Convert the middle number into escalation volume, and escalation volume into reviewer hours. Now the choice is a conversation with clinical leadership about acceptable risk and available capacity, backed by numbers, rather than a modelling preference. ## What each set size must mean in the workflow An uncertainty output that the workflow cannot act on is decoration. Decide in advance: - **Singleton.** The model is confident enough that only one label cleared the threshold. This is the case that can be auto-routed or fast-tracked, and it is where the throughput gain lives. - **Two or more labels — say melanoma and nevus together.** The model cannot separate two diagnoses with very different consequences. This is the most valuable output the system produces: it is a targeted request for expert attention, and it names *which* distinction is at issue. Route to a dermatologist with both candidates shown. - **Empty.** No label passed the threshold, meaning the classifier assigned low probability to everything — typically an image unlike the training distribution. This must be surfaced as an abstention or an out-of-distribution flag. Rendering it as "no finding" is the most dangerous possible interface bug on such a system. Some deployments force non-empty sets by always including the highest-probability label, which trades a clean guarantee for a predictable interface; if you do that, say so, because the coverage claim is no longer exact. ## Two claims to keep separate Coverage is not accuracy. A classifier with 60% top-1 accuracy can produce perfectly valid 90% sets — they will simply contain more than one label more often. Anyone reading "90% coverage" as "90% correct" will over-trust singletons and under-use multi-label sets, so the phrasing in the interface and in any regulatory documentation matters as much as the calibration itself. Coverage is also marginal. The 90% is an average over the case mix. A rare lesion type, a skin tone under-represented in training, or an unusual imaging device can sit far below it while the headline holds. For a clinical system this is not a technicality: commit in advance to the cohorts you will calibrate within, size the calibration data so each has enough rows, and report per-cohort coverage rather than only the aggregate. ## The organisational commitments Choosing alpha commits you to a recurring calibration process with fresh labelled data, to a staffing level implied by the escalation rate, to an interface that renders variable-size sets and abstentions honestly, and to a monitoring and re-approval story every time the classifier or the imaging pipeline changes. Those commitments, not the threshold arithmetic, are the substance of the decision, and they are the reason it belongs to a lead rather than to whoever is tuning the model.

  • What does an empty prediction set mean, and how should the product handle it?
    It means no label's probability cleared the calibrated threshold — the classifier is unusually unsure, often because the input is unlike anything in training. Route it to human review or an explicit abstain path, and never render it as a negative finding. If the interface cannot show an empty state, force the top label in and document that the exact coverage claim no longer holds.
  • Does 90% coverage mean the classifier is right 90% of the time?
    No. Coverage says the returned set contains the true label about 90% of the time; it says nothing about the top-1 prediction. A weak classifier reaches the same coverage by returning larger sets. Reporting them as the same number is how a system gets over-trusted, so keep set-based coverage and top-1 accuracy as separate published metrics.
  • How do you present the cost of moving from 90% to 99% coverage to clinical leadership?
    As a curve measured on held-out data: for each alpha, the average set size, the share of single-label sets, and the share of empty sets. Convert the multi-label share into escalation volume and reviewer hours per week. That turns an abstract threshold into a staffing and risk tradeoff they can actually decide on.

saying these in an interview costs you the question

  • Picks 95% because it is the usual number
  • Treats set coverage as top-1 accuracy
  • Renders an empty set to a clinician as a negative result
  • Assumes every case returns exactly one label
  • Reports only aggregate coverage for a clinical system

context