skip to content

What do the two axes of a binary classifier's ROC curve show, and what traces the curve?

level: juniorimportance: must knowfreq 88%

answer

  1. two rates, never raw counts
  2. y-axis: share of positives caught
  3. x-axis: share of negatives falsely flagged
  4. one point per decision threshold
  5. sorted-score walk: up, right, up

basics

~20 s

An ROC curve plots true positive rate on the y-axis against false positive rate on the x-axis. Each point is one decision threshold, and sweeping the threshold from strictest to loosest traces the curve from (0,0) to (1,1).

solid answer

~50 s

A scoring classifier outputs a continuous score, and a threshold turns that score into a yes/no decision. For each possible threshold you can compute two rates: the true positive rate, `TPR = TP / (TP + FN)`, the share of actual positives you catch, and the false positive rate, `FPR = FP / (FP + TN)`, the share of actual negatives you wrongly flag. Plotting TPR against FPR for every threshold gives the ROC curve. A very strict threshold sits at the bottom-left corner (0,0) where you flag nothing; a very loose one sits at the top-right corner (1,1) where you flag everything. Loosening the threshold can only move you up and to the right, so the curve is non-decreasing. Up and to the left is better; the diagonal from (0,0) to (1,1) is what a coin-flip ranking looks like.

go deeper

for a junior

Be ready to name both axes without hesitating and to say that one point equals one threshold. Knowing that up-and-left is good and the diagonal is random ranking covers most screening versions of this question.

for a middle

You are expected to derive the curve: sort by score, step up on a positive, right on a negative, and explain why ties create a diagonal segment and why the endpoints are (0,0) and (1,1).

for a senior

Show that you read curves diagnostically. A curve pinned to (0,1) usually means leakage; a curve hugging the diagonal means the score carries no ordering signal, and no threshold will rescue it.

for a principal

Own the framing question: the curve describes a family of classifiers, so the real decision is which region of it your product can operate in at all. Push teams away from reporting a curve with no reachable operating point marked on it.

## The setup: scores, not labels Most classifiers do not natively output a class. They output a **score** - a number where larger means "more likely positive". A **threshold** converts that score into a decision: flag everything at or above the threshold as positive, everything below as negative. Change the threshold and you get a different classifier out of the same model, with a different error profile. The ROC curve is the picture of every one of those classifiers at once. ## The four counts and the two rates Fix a threshold and compare decisions against truth. You get four counts: - **TP** - actual positives you flagged - **FN** - actual positives you missed - **FP** - actual negatives you flagged - **TN** - actual negatives you correctly left alone The ROC curve uses exactly two rates built from them: - **True positive rate**: `TPR = TP / (TP + FN)`. The share of the positive class you caught. Also called recall or sensitivity. Plotted on the **y-axis**. - **False positive rate**: `FPR = FP / (FP + TN)`. The share of the negative class you wrongly flagged. Equal to `1 - specificity`. Plotted on the **x-axis**. Notice that each rate has a denominator drawn from a single class: TPR looks only at real positives, FPR only at real negatives. Nothing in either rate mixes the two groups. ## Sweeping the threshold Start with the threshold above the highest score. Nothing is flagged, so `TP = 0` and `FP = 0`, giving `TPR = 0` and `FPR = 0` - the point (0,0). Now lower the threshold gradually. Every time it crosses the score of a real **positive**, that example flips from missed to caught: TP goes up by one, and the point moves **up** by `1 / (number of positives)`. Every time it crosses the score of a real **negative**, FP goes up by one and the point moves **right** by `1 / (number of negatives)`. Keep going until the threshold is below the lowest score: now everything is flagged, `TPR = 1` and `FPR = 1` - the point (1,1). That is why the curve on real, finite data is a **staircase**, not a smooth arc. It only changes at observed score values. If several examples share the same score, they must all be flagged or none of them, and the step becomes a **diagonal segment** whose slope reflects the mix of positives and negatives inside the tie group. A compact recipe: sort all examples by score descending, walk the list, step up for a positive and right for a negative, and join the points. ## Reading the picture - **Up and to the left is better.** The ideal is the point (0,1): a threshold that catches every positive with zero false alarms, meaning the two score distributions do not overlap at all. - **The diagonal `y = x` is the coin-flip line.** A support-ticket escalation score that ranks tickets no better than random produces a curve that hugs this diagonal, because at any threshold it flags the same fraction of urgent tickets as of routine ones. A curve near the diagonal is the visual statement "this score carries no ordering information". - **The corners are not choices, they are the degenerate extremes** - flag nothing, or flag everything. Real operating points live in between. - **Monotonicity.** Loosening the threshold can never reduce TP or FP, so the curve never goes down or left as you move along it. ## What the curve does not show The axes are rates, not counts, so the curve does not tell you how many examples there were, how many were positive, or how many alerts a given point would produce in a day. It also does not show the score values themselves: two models whose scores are on totally different numeric ranges can produce the identical curve, because only the **ordering** of the scores matters. And a single point on the curve is not a model property - it is a threshold you chose. ## Why interviewers open here The axes question is a screening question because getting it wrong usually means something deeper is wrong. Candidates who answer "precision against recall" are describing a different curve entirely; candidates who say each point is a different model have missed that a scoring model plus a threshold is what defines a point. Being able to say "y is TPR, x is FPR, one point per threshold, sorted-score walk traces it" in a single breath is the expected baseline.

  • Why is an ROC curve a staircase on real data rather than a smooth line?
    Because the data is finite. The confusion counts only change when the threshold crosses an actual observed score, so the curve is piecewise constant between scores. Crossing a positive steps the point up by one over the number of positives; crossing a negative steps it right by one over the number of negatives. Ties, where several examples share a score, produce a single diagonal segment instead of separate steps.
  • What does a point at (0, 1) on an ROC curve mean?
    Perfect separation at that threshold: every actual positive scored at or above it, and no actual negative did. The two score distributions do not overlap, so a single cut-off catches all positives with zero false alarms. On real data this almost always signals leakage - a feature that encodes the label - rather than a genuinely perfect model.
  • How does false positive rate relate to specificity?
    They are complements: `FPR = 1 - specificity`. Specificity is `TN / (TN + FP)`, the share of actual negatives correctly left alone. Medical and diagnostic literature usually plots sensitivity against `1 - specificity`, which is exactly TPR against FPR - the same curve under different names.

Think of a metal detector with a sensitivity dial. Turn it all the way down and it never beeps: no weapons found, no false alarms. Turn it up slowly and you start catching weapons - but also belt buckles. The ROC curve is the record of every dial setting.

saying these in an interview costs you the question

  • Says the x-axis is precision rather than false positive rate
  • Thinks each point on the curve is a different trained model
  • Describes the curve as accuracy plotted against threshold
  • Believes an ROC curve can be drawn from hard labels alone
  • Confuses TPR with the share of flagged cases that are correct

context