skip to content

In an inpatient deterioration-risk service, which decisions do the offline sweep metric and the online launch metric each gate?

level: middleimportance: must knowfreq 72%

answer

  1. two numbers, not one
  2. one is swept, one is launched
  3. frozen cohort against live shifts
  4. fast score filters, slow read decides
  5. alert load reported beside escalations

basics

~20 s

The offline sweep metric ranks candidate models during training - a ranking score over a frozen retrospective cohort. The online launch metric decides whether the service ships: what the deployed worklist changed in care, read on live shifts.

solid answer

~40 s

Framing names two numbers, not one. The **offline sweep metric** is the score you optimise against - `ROC AUC` or `precision@20` over a frozen cohort of admissions whose outcomes are already recorded - and its only job is to rank hundreds of candidates cheaply enough to sit inside a training loop. The **online launch metric** is what the deployed service is supposed to change: escalations that changed a patient's care per shift, reported beside the alert load per nurse-shift that produced them. It is slow, needs chart review to attribute, and accumulates only as fast as real deterioration events do - so it can never be swept on, and it is the only one a launch turns on. Writing both down early is what makes their later disagreement diagnosable instead of a surprise.

go deeper

for a junior

Remember that two different numbers exist: a fast score computed on historical data, and a slow read taken on the live ward. They are not the same measurement.

for a middle

Explain why the fast one is a proxy: it is chosen for turnaround, so it scores the whole ranking, a historical population, and a predicted event rather than a changed outcome.

for a senior

Show that you fix k at the deployed list length, that you log escalation attribution from day one, and that you settled which number has launch authority before any candidate existed.

for a principal

Frame the pair as an evidence policy for the service, not one project's choice: what the organisation lets a cheap proxy authorise, and what it insists on paying live-shift time to learn.

## Two numbers, named in the first ten minutes An ML design round for an inpatient deterioration-risk service opens with a business ask - help the rapid-response team reach patients before they crash - and should leave the framing stage holding **two** metrics, not one. The **offline sweep metric** is a score computed on a frozen retrospective cohort: admissions whose outcomes are already recorded. Its job is mechanical. It has to be cheap and repeatable enough to run hundreds of times inside a training loop, so that feature sets, model families and hyperparameters can be ranked against each other before anything reaches a ward. For a ranked worklist that means a ranking score - `ROC AUC` across all scored patient-days, or `precision@20` when the deployed list is twenty rows per shift. The **online launch metric** is the change the deployed service is supposed to produce on the ward: *escalations that changed care* per shift, reported beside the alert load per nurse-shift that produced them. It is read on live shifts, it needs chart review to tie an action to an alert, and it accumulates only as fast as real deterioration events do. ## What each one gates | | offline sweep metric | online launch metric | |---|---|---| | computed on | a frozen historical cohort | live shifts | | turnaround | minutes, repeatable | weeks to a quarter | | cost of one read | a compute job | clinician attention and chart review | | what it decides | which candidates are contenders | whether the service replaces the incumbent | | optimised against | yes - that is its purpose | no | ## Why the fast number cannot be the launch number - **It scores a slice nobody consumes.** `ROC AUC` integrates over every ordered pair in the ranking; a shift only ever sees twenty rows. An ordering improvement at rank 300 raises the offline score and changes nothing on the ward. - **It scores a population the service does not serve.** The retrospective cohort contains admissions the live path never scores - patients discharged before the first score, admissions missing the inputs the model needs, wards not yet onboarded. - **It scores a different event.** Predicting deterioration correctly is not the same as a clinician managing the patient differently because of the prediction. - **It cannot see redundancy.** A candidate that promotes patients the team was already watching is indistinguishable, offline, from one that surfaces patients nobody had noticed. - **It was chosen for speed.** Every property above is the price of a metric that returns in minutes. That is a trade made knowingly, not a defect to be fixed by finding a better offline score. ## Why the launch number cannot be the sweep number 1. **Turnaround.** One honest read spans weeks of shifts; a sweep needs hundreds of reads. 2. **Cost.** Every read spends clinician attention inside a safety-critical path, plus the chart review that decides whether care actually changed. 3. **Exposure.** You cannot put a ward in front of a hundred candidate rankers to discover which one is best. ## What naming the pair buys the design - It fixes `k` in the offline metric at the deployed list length, so the number you sweep on at least measures the slice the ward will consume. - It tells you what the post-launch monitoring has to carry, because the online read is not something you can bolt on afterwards - the escalation attribution has to be logged from day one. - It makes the *proxy gap* - offline up, online flat or down - a diagnosable event rather than a crisis, because both numbers were declared before either moved. - It settles authority in advance: the offline number filters, the online number decides. Teams that skip this argue about authority under deadline pressure, with a candidate already built. ## What a weak framing looks like A framing that names only `ROC AUC` has quietly promoted a training convenience to a launch criterion. A framing that names only *escalations that changed care* has no way to choose between two candidates this week. A framing that names an online number with no alert load beside it has not stated what the ward is paying for those escalations - twenty alerts per shift and eighty alerts per shift are different services, and the same escalation count means opposite things under each. In the round, the answer an interviewer is listening for is not what `ROC AUC` means; it is which number has authority over which decision, and why a single metric cannot hold both roles.

  • Why does the offline sweep metric have to be computable inside a training loop?
    Because the sweep is a search. Ranking feature sets, model families and hyperparameters means scoring hundreds of candidates, and a metric that costs a compute job can support that while one that costs weeks of live shifts cannot. Cheapness is the requirement the offline metric exists to satisfy; fidelity to the launch decision is the thing it trades away to get there.
  • What does 'escalations that changed care' require that a raw alert count does not?
    An attribution rule and a definition of the change. You need a window linking an alert to a subsequent action, a record of what the team did, and a judgment - usually chart review - that the action altered management rather than restating the plan already in place. That human step is why the metric reads in weeks, not minutes.
  • Should the offline metric be ROC AUC or precision@20 here?
    `precision@20` when the deployed list is twenty rows per shift, because it measures the slice the ward will actually consume. `ROC AUC` is steadier across candidates and useful early in a sweep, but a candidate can win on it purely by reordering ranks nobody reads, so it makes a poor final filter for a worklist product.

saying these in an interview costs you the question

  • Says the candidate with the best offline score ships.
  • Treats the offline metric as a reliable predictor of the online one.
  • Proposes sweeping hundreds of candidates directly on the live read.
  • Calls the offline number wrong when the two numbers disagree.
  • Reports escalations per shift without the alert load that produced them.