What does a precision-recall curve plot, and what does average precision measure?
answer
- one model, every possible cut-off
- two axes, both about positives
- recall on x, precision on y
- each threshold contributes one point
- precision weighted by recall gained
basics
~20 sA precision-recall curve plots precision against recall as the decision threshold sweeps from strict to permissive. Average precision compresses that curve into one number: the recall-weighted mean of the precisions along it, estimating the area beneath.
solid answer
~50 sA binary classifier outputs a score per row; a threshold turns scores into positive predictions. Sweep that threshold from very strict to very permissive and each setting gives one pair — recall (the share of true positives you caught) and precision (the share of your positive predictions that were right). Plotting precision on the y-axis against recall on the x-axis traces the curve. It typically starts high-precision/low-recall on the left and falls as recall rises, because loosening the threshold admits more false positives. Average precision summarises the curve as `AP = sum over thresholds of (R_n - R_{n-1}) * P_n` — each precision value weighted by the recall it bought — which is a step-wise estimate of the area under the curve. The point of the curve is that it shows the whole tradeoff at once, before any single operating point is fixed.
code
python · 17 lines# labels of ten rows, sorted by model score, highest score first
ranked = [1, 0, 1, 1, 0, 0, 1, 0, 0, 0]
total_pos = sum(ranked)
hits = 0
ap = 0.0
for rank, y in enumerate(ranked, start=1):
hits += y
if y == 1: # recall stepped up here
precision = hits / rank
recall = hits / total_pos
ap += precision # weight is one recall step = 1/total_pos
print(f"rank {rank}: recall {recall:.2f} precision {precision:.2f}")
ap /= total_pos
print(f"average precision = {ap:.3f}")
print(f"no-skill baseline = {total_pos / len(ranked):.3f}")go deeper
Be ready to name both axes without hesitating and to say that the curve comes from sweeping the threshold across a single model's scores, not from one set of predictions.
Explain how each threshold yields one point, why admitting a false positive drops precision while recall stands still, and how average precision weights each precision by the recall it bought.
Show that you treat average precision as an estimate from a finite sample: the high-precision left end rests on very few rows, and the number is meaningless unless the evaluation set's positive rate is quoted with it.
Own which curve your organisation tracks. A precision-recall summary frames performance in terms of the flagged workload, and that framing quietly defines what every review afterwards calls an improvement.
## What the curve is made of A binary classifier normally produces a **score** for each row — a number where higher means "more likely positive". Nothing is a prediction yet. To get predictions you pick a **threshold**: every row scoring at or above it is predicted positive, everything else negative. Change the threshold and you get a different set of predictions from the *same* model. At any one threshold you can count four things: true positives (TP, actually positive and predicted positive), false positives (FP, actually negative but predicted positive), false negatives (FN, actually positive but predicted negative), and true negatives. Two ratios come out of them: - `precision = TP / (TP + FP)` — of everything you flagged, what fraction was really positive. - `recall = TP / (TP + FN)` — of everything that was really positive, what fraction you caught. A precision-recall curve is what you get by computing that pair at **every** threshold and plotting precision (y-axis) against recall (x-axis). It is a property of the model's *ranking* of rows, not of any one cut-off. ## How the sweep traces the curve Start with a very high threshold: almost nothing is flagged, so recall is near 0, and precision is whatever the top handful of scores happens to be. Lower the threshold one row at a time. Each time the newly admitted row is a true positive, recall steps up and precision ticks up; each time it is a false positive, recall stays put and precision drops. The result is a saw-toothed, generally downward-sloping path from the left (strict, high precision, low recall) to the right (permissive, low precision, recall 1). Two shape facts trip people up: - The curve is **not** required to be monotonically decreasing. Locally it rises whenever a run of true positives is admitted. Practitioners often draw a monotone "envelope" over it, but the raw curve wiggles. - The far-left end is unreliable. Precision there is computed from a handful of predictions, so it swings between 1.0 and 0.5 on the basis of one or two rows, and at recall exactly 0 it is undefined (zero predictions, zero denominator). The right-hand end is pinned: at a threshold low enough to flag everything, recall is 1 and precision equals the fraction of the data that is positive. ## Average precision as the single number Comparing whole curves by eye does not scale, so the curve is usually summarised by **average precision (AP)**: ``` AP = sum_n (R_n - R_{n-1}) * P_n ``` where n indexes the thresholds in order of increasing recall, `R_n` is the recall at that threshold and `P_n` the precision. Read it as: every precision value is weighted by the slice of recall it bought. Equivalently, for a ranked list with no tied scores, AP is the mean of the precision values measured at each rank where a true positive appears. This is a **step-wise (right-hand) estimate of the area under the curve**, and it is deliberately not the trapezoidal area. Trapezoids draw straight lines between measured points, which quietly assumes you could reach intermediate recalls at intermediate precision — an optimistic assumption when the curve is jagged. The step-wise sum makes no such interpolation. The area is bounded between 0 and 1, and higher is better. What counts as good is **not** anchored at 0.5: a model that scores rows at random has precision equal to the positive rate at every recall, so the no-skill line on a PR plot is a horizontal line at the prevalence of the positive class. On a dataset that is 30% positive, AP of 0.35 is barely better than nothing; on a dataset that is 0.2% positive, AP of 0.15 is dramatically better than nothing. AP is always read against that baseline. ## Why you would use it The curve answers a question a single metric cannot: *across the whole range of operating strictness, how well does this model do?* That matters because the threshold is usually chosen later, by someone weighing costs, and because two models can have identical accuracy while one ranks the true positives much closer to the top. It is especially informative when positives are rare and you only ever act on the flagged set, since both precision and recall are computed purely from positives and flagged rows — the vast pool of correctly-ignored negatives never enters either formula. When you report AP, report the positive rate of the evaluation set alongside it. Without that number, the AP value cannot be interpreted at all.
- Why does a precision-recall curve generally slope downward as recall increases?Raising recall means lowering the threshold, which admits rows the model was less confident about. Those extra rows are disproportionately negatives, so the flagged set grows faster than the true positives inside it and precision falls. The slope is steep for a weak ranker and shallow for a strong one.
- Is average precision the same as the trapezoidal area under the precision-recall curve?No. Average precision is a step-wise sum, `sum (R_n - R_{n-1}) * P_n`, with no interpolation between measured points. Trapezoidal area draws straight lines between them, which assumes intermediate recall is reachable at intermediate precision — optimistic on a jagged curve. The two agree only when the curve is smooth and densely sampled.
- What is precision at the extreme left of the curve, where recall is near zero?It is computed from a handful of top-scored rows, so it is extremely noisy — one row changing rank moves it from 1.0 to 0.5. At recall exactly zero nothing is flagged and precision is undefined (0/0); implementations fix a convention there. Never read a business conclusion off that end.
It is a dial, not a snapshot: you turn the strictness dial through its whole range and photograph the precision and recall at every position, then average precision is the one number that describes the whole roll of film.
saying these in an interview costs you the question
- Says the axes are the same as an ROC curve's
- Thinks one threshold produces the whole curve
- Confuses average precision with accuracy averaged over classes
- Insists the curve must decrease monotonically
- Treats 0.5 as the chance level for average precision