How do you turn a cost matrix over false positives and false negatives into a decision threshold?
answer
- compare two expected costs per item
- all four cells, not two
- acting still costs when you were right
- the ratio of costs, not the base rate
- equal costs is the only 0.5 case
basics
~20 sFlag an item when the expected cost of flagging is below the expected cost of not flagging. With false-positive and false-negative costs only, the break-even is cost_FP / (cost_FP + cost_FN), which is 0.5 only for equal costs.
solid answer
~50 sWrite down what each of the four outcomes actually costs, then flag an item whenever doing so has the lower expected cost. For an item with estimated positive probability `p`, flagging costs `p*C_TP + (1-p)*C_FP` and not flagging costs `p*C_FN + (1-p)*C_TN`. Setting those equal gives the break-even probability `p* = (C_FP - C_TN) / ((C_FP - C_TN) + (C_FN - C_TP))`. Take predictive maintenance: a callout costs $2k whether or not the machine was really failing, and a missed failure costs a $50k unplanned outage. Then `p* = 2000 / (2000 + 48000) = 0.04` -- send an engineer on a 4% chance of failure, not a 50% one. Two caveats: the rule assumes the score really is a probability, and if it is not, sweep candidate cuts on a validation set and pick the one with the lowest measured total cost instead.
code
python · 13 linesscored = [(1, 0.92), (0, 0.61), (1, 0.55), (0, 0.40), (1, 0.31),
(0, 0.22), (0, 0.15), (1, 0.09), (0, 0.05), (0, 0.02)]
COST = {(1, 1): 2_000, # flagged and it really fails: pay the callout
(0, 1): 2_000, # false alarm: pay the callout anyway
(1, 0): 50_000, # missed failure: unplanned outage
(0, 0): 0}
def total_cost(cut):
return sum(COST[(y, 1 if s >= cut else 0)] for y, s in scored)
for cut in (0.01, 0.04, 0.10, 0.30, 0.50):
print(f"cut={cut:.2f} cost={total_cost(cut):>7,}")
print("break-even probability =", 2_000 / (2_000 + (50_000 - 2_000)))go deeper
Be ready to say that a threshold should reflect what each mistake costs, and that 0.5 is only right when a false alarm and a miss hurt equally. Knowing the direction -- expensive misses push the cut down -- is enough here.
Derive the break-even by writing the expected cost of flagging against the expected cost of not flagging and setting them equal. Expect to be pushed on the cost of a true positive, which is not zero when acting itself costs money.
Demonstrate that you check the assumption before using the formula: if the score is not a real probability, sweep cuts on validation data and minimise measured cost instead, and handle per-item costs by thresholding row by row.
Own the cost model as a governed artefact. Decide who signs off the ratio, how often it is revisited, and how you surface the implied ratio of whatever threshold is running today so a silent default never stands in for a business decision.
## The decision, not the metric Most threshold arguments go in circles because they are arguments about metrics. The expected-value framing sidesteps that: a classifier's output feeds an **action** (send an engineer, hold the payment, open a ticket), each combination of action and truth has a **consequence**, and the job is to choose the action with the best expected consequence. The threshold falls out of that arithmetic rather than being chosen by taste. ## The four cells Write the cost of each outcome. Costs are conventionally positive for bad things; a benefit is a negative cost. - `C_TP` -- you flagged it and it really was positive - `C_FP` -- you flagged it and it was not - `C_FN` -- you did not flag it and it really was positive - `C_TN` -- you did not flag it and it was not People routinely fill in only `C_FP` and `C_FN` and set the other two to zero. That is often wrong, and the predictive-maintenance case shows why: dispatching an engineer costs $2,000 *whether or not the machine was going to fail*. The callout is not free just because the prediction was right. So `C_TP = 2000` as well as `C_FP = 2000`. ## Deriving the break-even probability For a single item, let `p` be the estimated probability it is positive. The two expected costs are: ``` flag: p*C_TP + (1-p)*C_FP do not flag: p*C_FN + (1-p)*C_TN ``` Flag when the first is smaller. Rearranging gives the break-even point: ``` p* = (C_FP - C_TN) / ((C_FP - C_TN) + (C_FN - C_TP)) ``` The two differences have clean readings. `C_FP - C_TN` is **the price of acting when you did not need to**. `C_FN - C_TP` is **the price of not acting when you needed to** -- the outage you take minus the callout you would have paid anyway. The threshold is the first divided by the sum of both. With `C_FP = C_TP = 2000`, `C_FN = 50000`, `C_TN = 0`: ``` p* = 2000 / (2000 + 48000) = 0.04 ``` Act on a 4% chance. If you had lazily used the two-cost shortcut `C_FP / (C_FP + C_FN) = 2000/52000 = 0.038` you would land close here, but the shortcut breaks badly whenever acting on a true positive is expensive or a correct non-action carries its own cost. The sanity checks are easy. Equal costs give `p* = 0.5`, which is exactly where the 0.5 default comes from -- it is the answer to a symmetric problem nobody checked they had. A miss ten times worse than a false alarm gives `p* = 1/11 = 0.09`. As misses get catastrophic, `p*` goes to 0 and you flag almost everything; as false alarms get catastrophic, `p*` goes to 1. ## What the rule assumes **That `p` is a genuine probability.** The arithmetic multiplies the model's output by dollars, so a score of 0.04 has to mean "happens about 4% of the time". Many models output well-ordered scores that are not probabilities, and then a cut of 0.04 means nothing in dollars. The practical escape hatch is empirical: sweep candidate cuts over a labelled validation set, compute the realised total cost at each, and take the minimum. That works on any monotone score and needs no probabilistic interpretation. It also folds in the second assumption, that the validation mix resembles production. **That the base rate does not enter the formula.** This surprises people. The threshold `p*` depends only on costs. Prevalence enters through `p` itself -- a rare-positive problem produces small probabilities, so fewer items clear the same cut. Candidates who try to put the base rate into the threshold formula have muddled the estimate with the decision rule. **That costs are constant per item.** Often they are not: a $50k outage on one machine is a $500k outage on another. The honest fix is a per-item threshold from per-item costs, which is the same formula applied row by row, or ranking by expected saving `p*(C_FN - C_TP) - (1-p)*(C_FP - C_TN)` and acting on the positive ones. ## Getting the numbers The hardest part is rarely the algebra; it is that nobody owns the cost figures. Two moves help. First, only the **ratio** matters, so you do not need exact dollars -- "a miss is about 25 times worse than a false alarm" fully determines the cut. Second, invert the question when finance will not commit: compute the cut the current threshold implies and show it to the business. "Running at 0.5 means you are telling me a missed outage is exactly as expensive as a needless callout -- is that right?" That conversation converges far faster than asking for a cost matrix cold, and it turns a silent default into an explicit, reviewable decision.
- Finance will not give you dollar costs. What do you do?Work with the ratio, since only the ratio sets the cut -- "a miss is roughly 25 times worse than a false alarm" is enough. If even that is contested, invert the question: compute the cost ratio the current threshold already implies and take it back to the business. Being told the running default asserts that a missed outage and a needless callout are equally expensive usually produces an opinion within minutes.
- Does the positive base rate belong in the threshold formula?No. The break-even probability depends only on the cost cells. Prevalence enters through the estimated probability of each item, not through the cut: on a rare-positive problem the scores themselves are small, so fewer items clear the same threshold. Putting the base rate into the formula double-counts it and is a common muddle between the estimate and the decision rule.
- What if the cost of a miss varies per item rather than being a flat figure?Apply the same formula per row. Each item gets its own break-even from its own costs, so a machine whose outage costs $500k gets a far lower cut than one at $50k. Equivalently, rank by expected saving -- probability times the avoided loss, minus the expected cost of acting -- and take every item where that is positive. A single global threshold is only a convenience when costs are roughly uniform.
It is an umbrella decision: carrying one costs a small nuisance whether or not it rains, and the chance of rain at which you bother to carry it is set entirely by the nuisance versus the soaking.
saying these in an interview costs you the question
- Says 0.5 is the correct threshold by default
- Fills in only false-positive and false-negative costs
- Puts the class base rate into the threshold formula
- Applies the rule to uncalibrated scores without noticing
- Insists exact dollar costs are needed when the ratio suffices
- Confuses the cost of acting with the cost of being wrong