skip to content

Lowering a trained classifier's decision threshold from 0.5 to 0.3 changes what, and what stays fixed?

level: juniorimportance: must knowfreq 78%

answer

  1. nothing is retrained
  2. items only cross one way
  3. recall can only move one direction
  4. precision is not guaranteed monotone
  5. ranking summaries sweep every cut

basics

~20 s

More items are labelled positive, so true positives and false positives can only rise, recall can only rise, and precision usually falls. The fitted model, its scores and its ranking of items do not change at all.

solid answer

~50 s

The threshold is applied *after* the model has scored every item, so lowering it from 0.5 to 0.3 changes nothing about the model: same parameters, same scores, same ordering of items from most to least likely positive. What changes is the confusion matrix. Everything scored between 0.3 and 0.5 flips from predicted-negative to predicted-positive, so true positives and false positives can only go up, false negatives and true negatives can only go down. Recall (sensitivity) therefore weakly increases and specificity weakly decreases. Precision usually falls, but on a finite sample it is not strictly monotone. Threshold-free summaries such as ROC-AUC are computed across every possible cut, so they do not move. You are trading missed positives for false alarms, and where you sit on that trade is a business decision, not a property of the model.

go deeper

for a junior

Be ready to state the direction of each count: lowering the cut can only add predicted positives, so true and false positives rise while false negatives fall. Say plainly that the model itself is untouched.

for a middle

An interviewer expects you to explain why recall is monotone in the threshold but precision is not, using the fact that recall's denominator is the fixed number of real positives while precision's denominator grows with every item you admit.

for a senior

Show that you treat the cut as a product decision made after training: one scorer serving two thresholds for two use cases, and every reported precision or accuracy figure tagged with the operating point it came from.

for a principal

Own the organisational consequence: if the threshold is a business lever, decide who is allowed to move it, how a change is reviewed, and what is monitored afterwards -- an unlogged threshold change looks exactly like a silent model regression.

## Two separate objects: the score and the cut A binary classifier does two things that people often blur together. First it produces a **score** for each item -- typically a number between 0 and 1 meant to express how likely the item is to be positive. Second, somebody turns that score into a hard **label** by comparing it against a **decision threshold**: predict positive if `score >= t`, negative otherwise. Only the first of these is learned. The threshold is a knob bolted on afterwards, and 0.5 is nothing more than a convention that many tools default to. Moving it from 0.5 to 0.3 does not retrain anything, does not change a single parameter, and does not change any item's score. It changes only which side of the line each item falls on. ## What changes: the confusion matrix All four cells of the confusion matrix are counts of items on one side of the cut: - **TP** -- positive item, predicted positive - **FP** -- negative item, predicted positive - **FN** -- positive item, predicted negative - **TN** -- negative item, predicted negative Lowering the cut can only move items from the predicted-negative column to the predicted-positive column, never the other way. So: - TP is non-decreasing, FN is non-increasing - FP is non-decreasing, TN is non-increasing The derived rates follow directly: - `recall = TP / (TP + FN)` -- the denominator is the number of real positives, which never changes, and the numerator can only grow. **Recall is monotone non-decreasing as the threshold falls.** - `FPR = FP / (FP + TN)` -- same argument on the negatives. **FPR is monotone non-decreasing**, so specificity (`1 - FPR`) is non-increasing. - `precision = TP / (TP + FP)` -- *both* the numerator and the denominator grow, so this is **not** guaranteed to move in either direction. On a well-behaved model it trends downward as you admit lower-scored items, but on a finite sample it wobbles: if the next item you admit happens to be a true positive, precision ticks up. That last asymmetry is worth remembering, because it is the one people state incorrectly. Recall monotone: yes. Precision monotone: no. ## What does not change - **The model.** Parameters, learned representation, training data -- untouched. - **The scores.** Item A scored 0.62 before and scores 0.62 after. - **The ranking.** Because scores are unchanged, the ordering of items is unchanged. Anything that depends only on the ordering is therefore also unchanged. - **Threshold-free ranking summaries.** ROC-AUC integrates performance over every possible threshold, so by construction it cannot respond to picking one of them. A candidate who says "we lowered the threshold and AUC went up" has either retrained something or is confused about what AUC measures. ## The extreme cases Set the threshold to 0 and everything is predicted positive: recall is 1.0, FPR is 1.0, and precision equals the positive base rate of the data -- 0.5% precision on a 0.5%-positive problem. Set the threshold to 1.0 (above every score) and nothing is predicted positive: recall is 0, FPR is 0, precision is undefined because there are no predicted positives to be right about. Every useful operating point lives between those two, and the curve of achievable (recall, FPR) pairs traced out as the cut sweeps is exactly what an ROC plot draws. ## Why this matters practically Because the cut is decoupled from the model, two things follow that shape a lot of real work. 1. **A single fitted model supports many products.** The same churn scorer can drive an aggressive retention campaign (low cut, high recall, lots of wasted vouchers) and a conservative executive-escalation list (high cut, few names, high precision). You do not need two models; you need two thresholds. 2. **Reporting a single accuracy or F1 number hides the choice.** "Our model is 92% accurate" is a statement about a model *and a threshold*. If someone quietly moved the cut, the number moved with it, and no retraining happened. The honest framing for an interview: the model's job is to rank and score; choosing where to cut is a decision problem driven by what a false alarm costs relative to a miss, how much review capacity exists, and what the product promises. Those are the questions to ask before touching the number.

  • Does precision always fall as the threshold is lowered?
    No. Lowering the cut adds items to the predicted-positive set, and both the numerator (true positives) and the denominator (all predicted positives) can grow, so precision has no guaranteed direction. It trends down for a model that ranks well, but on real samples it is jagged -- admitting one more genuine positive nudges it back up. Recall, by contrast, is genuinely monotone because its denominator is fixed.
  • If the threshold changes nothing about the model, why report metrics at a threshold at all?
    Because the product ships hard decisions, not scores. Somebody is refunded or not, an alert fires or not. Accuracy, precision, recall and F1 are all properties of a chosen operating point, so they describe the shipped system rather than the model. Threshold-free measures describe the model's ranking quality; you need both, and you should always say which threshold a reported number came from.
  • What is the best recall you can reach by lowering the threshold, and what does it cost?
    Recall 1.0, reached by setting the cut low enough that every item is flagged. It is free to achieve and worthless: precision then equals the positive base rate, so on a problem with 1% positives you review 100 items to find one. That extreme is the reminder that recall alone is never a target -- it is only meaningful paired with the precision or the volume it forces.

It is like moving the pass mark on an already-graded exam: nobody's score changes and nobody's rank changes, only who gets labelled a pass.

saying these in an interview costs you the question

  • Says lowering the threshold retrains or changes the model
  • Claims recall can decrease when the threshold is lowered
  • Thinks ROC-AUC improves after moving the cut
  • Treats 0.5 as a property of the model rather than a choice
  • Assumes precision falls by exactly what recall gains
  • Reports accuracy without saying which threshold produced it

context