Two candidate models' ROC curves cross - how do you decide which one to ship?
answer
- no dominance, each wins somewhere
- one scalar cannot express both regions
- area weights unreachable thresholds equally
- compare inside the reachable band
- partial area, or TPR at fixed FPR
basics
~20 sCrossing curves mean neither model dominates - each wins over a different range of false positive rates. Compare them where you will actually operate, such as true positive rate at a fixed low false positive rate, not by total area.
solid answer
~50 sA crossing says there is no dominance: one model separates better at strict thresholds, the other at loose ones. Total AUC still gives one number each, but it integrates true positive rate uniformly over the whole false-positive-rate axis, including regions no product would run in - so the higher AUC can belong to the model that is worse everywhere you can actually operate. Decide where the crossing does not matter: fix the operating region first, then compare true positive rate at the reachable false positive rate, or partial AUC restricted to that range. Add three checks: whether the gap survives resampling rather than being curve noise, whether the ordering holds across segments and over time, and whether combining the two scores beats both. Then ship on the region-specific comparison and record why the scalar AUC was overridden.
go deeper
Understand what a crossing means geometrically: neither curve is above the other everywhere, so neither model is better at every threshold. Knowing that a higher AUC does not settle it is enough at this level.
Explain why a single area can mislead - it weights every false-positive-rate region equally - and name the two region-restricted alternatives, true positive rate at a fixed false positive rate and partial AUC.
Show the diagnostic sequence: bound the reachable region, compare inside it, confirm the gap survives resampling, and recompute per segment and per time window before trusting the ordering.
Own the decision and its record. Set the operating region with the product side, decide whether combining beats choosing, weigh serving cost and explainability, and document why the headline AUC was overridden so nobody reverses the call from a dashboard.
## What a crossing actually tells you One ROC curve **dominates** another if it lies weakly above it at every false positive rate. Under dominance the choice is easy: the dominating model is at least as good at every possible threshold, and its AUC is necessarily higher. A **crossing** is the explicit statement that dominance does not hold. Model A is above model B up to some false positive rate; past that point model B is above A. Each model is the better model somewhere. This is common in practice, and it usually has a mechanical cause. A model with a very confident, sparse high-score head - a boosted tree ensemble that isolates a small set of near-certain positives - can win decisively at the strict end while flattening out later. A smoother model - a regularised linear score over broad features - can lose the strict end and win the middle, because it never assigns extreme scores it cannot back up. ## Why the AUC comparison is the wrong instrument here AUC integrates true positive rate over false positive rate with **uniform weight**. That means it treats a gain at `FPR = 0.9` as worth exactly as much as the same-sized gain at `FPR = 0.001`. For a fraud queue, a moderation pipeline, or a credit decision, false positive rates above a few percent are unreachable - the alert volume or the customer harm is unacceptable long before you get there. So the total area is partly a summary of behaviour in a region the model will never see. When curves cross, this stops being a philosophical objection and becomes a live risk: the model with the higher AUC can be the one that is worse at every operating point you could ever choose, because it banked its area in the far right of the plot. A single scalar cannot express "better here, worse there", and a crossing is precisely the case where that distinction is the whole decision. ## How to decide instead **1. Fix the operating region before comparing.** Establish, from the product side, the band of false positive rates that is reachable at all. This is not the same as picking the exact threshold - it is bounding the plot. Everything outside that band is irrelevant evidence. **2. Compare inside the region.** Two standard readouts: - **TPR at a fixed FPR** - "at a 1% false positive rate, model A catches 62% of positives and model B catches 54%". This is the most legible comparison to non-specialists and it is exactly a point-wise read of the curve. - **Partial AUC** - the area restricted to a false-positive-rate interval, typically rescaled to run from 0 to 1 over that interval. It preserves AUC's summary character while discarding the unreachable region. The cost is variance: fewer negatives fall inside a narrow low-FPR window, so the estimate is noisier than full AUC and needs a bigger evaluation set to be trusted. **3. Check that the crossing is real.** On a finite evaluation set two curves will jitter around each other; a crossing near the middle of the plot with only a handful of examples separating them may be pure sampling noise. Whether an observed gap between models is statistically meaningful is its own discipline - the point here is that a visible crossing is not automatically a structural fact about the models, and a resampled comparison should be run before a shipping decision rests on it. **4. Check stability, not just the average curve.** Recompute the region-restricted comparison per segment - by geography, device, tenure, traffic source - and across several time windows. A model that wins the operating region on average but loses it on your largest segment, or that changes places from month to month, is a worse ship than the average curve suggests. **5. Consider not choosing.** Crossing curves are a hint that the two models make different errors. Averaging their scores, or averaging their ranks - which is the natural combiner when only ordering matters - frequently produces a curve that beats both, precisely because the disagreement carries information. Validate the combination on held-out data rather than assuming it. **6. Weigh the non-curve factors.** Latency, retraining cost, feature dependencies, explainability obligations and operational familiarity are real inputs. When two models are close inside the operating region, these decide - and pretending the curve settled it is a way of hiding the actual reason. ## The judgment being tested The question is not really about ROC geometry. It is about whether you will let a convenient scalar make a decision that the scalar cannot see. The expected answer refuses the framing of "which AUC is higher", replaces it with "which model is better in the region we can operate in, and is that difference stable and real", and is explicit that the conclusion must be written down - including the fact that the headline AUC pointed the other way, so that the next person to read the dashboard does not silently reverse the call.
- What is partial AUC, and when is it the better summary?Partial AUC is the area under the ROC curve restricted to a range of false positive rates, usually rescaled so it runs from 0 to 1 over that range. It is the better summary when only part of the curve is operationally reachable - a screening system that cannot exceed a 1% false positive rate, say. The tradeoff is variance: fewer negatives fall inside a narrow window, so the estimate is noisier than full AUC.
- Can two models with identical AUC behave very differently in production?Yes, and it is common. AUC is a single integral, and many different curve shapes integrate to the same area. One model may be sharply better at low false positive rates and worse in the middle, the other flat and even. Their production behaviour at a strict threshold can differ enormously while the reported scalar is indistinguishable, which is why the curve is worth plotting rather than summarising.
- Would combining the two crossing models help rather than choosing one?Often, yes. A crossing suggests the models rank different examples well, so their errors are partly decorrelated. Averaging their scores - or averaging ranks, which is the natural combiner when only order matters - can produce a curve above both across the operating region. It is not guaranteed, so validate on held-out data, and weigh the extra serving and maintenance cost of running two models.
saying these in an interview costs you the question
- Picks the higher AUC without plotting either curve
- Assumes AUC ordering holds at every threshold
- Treats a 0.005 AUC gap as a decisive result
- Ignores which false-positive region is operationally reachable
- Believes a genuinely better model can never have a crossing curve