For a rare-positive classifier missing its F1 target, do you train soft-F1 or tune the threshold?
answer
- two levers, one number
- does the loss need the operating point?
- summed probabilities replace hard counts
- a ratio of sums does not average over batches
- cut point tuned on held-out data
basics
~20 sTune the threshold first: train with cross-entropy, then pick the operating point on a held-out split. A differentiable soft-F1 loss, built from summed probabilities instead of hard counts, is the fallback when no fixed threshold works.
solid answer
~50 sStart by separating the two levers. Cross-entropy training gives you a good *ordering* of examples; the threshold decides where you cut that ordering. A model that misses its F1 target at 0.5 has very often learned a fine ordering and is simply being cut in the wrong place, and sweeping the threshold on a held-out split to maximise F1 fixes it in minutes without touching training. That separability is the default answer. A **soft-F1** loss makes the metric differentiable by replacing hard counts with expectations over the sigmoid outputs: `tp = sum(p_i * y_i)`, `fp = sum(p_i * (1 - y_i))`, `fn = sum((1 - p_i) * y_i)`, then minimise `1 - 2tp / (2tp + fp + fn)`. It genuinely helps when a fixed threshold is impractical — thousands of multi-label heads, say. The cost is real: F1 does not decompose over examples, so a per-batch value is a biased, batch-size-dependent estimate of the corpus metric, and the loss pushes probabilities toward 0 and 1, destroying calibration.
code
python · 22 linesys = [1, 1, 0, 0, 1]
ps = [0.90, 0.60, 0.55, 0.10, 0.45]
def f1(tp, fp, fn):
return 2 * tp / (2 * tp + fp + fn)
hard = [1 if p >= 0.5 else 0 for p in ps]
tp = sum(1 for h, y in zip(hard, ys) if h == 1 and y == 1)
fp = sum(1 for h, y in zip(hard, ys) if h == 1 and y == 0)
fn = sum(1 for h, y in zip(hard, ys) if h == 0 and y == 1)
print("hard F1 ", round(f1(tp, fp, fn), 4))
def soft_f1(ps, ys):
stp = sum(p * y for p, y in zip(ps, ys))
sfp = sum(p * (1 - y) for p, y in zip(ps, ys))
sfn = sum((1 - p) * y for p, y in zip(ps, ys))
return f1(stp, sfp, sfn)
print("soft F1 ", round(soft_f1(ps, ys), 4))
ps[4] = 0.49 # still below the 0.5 cut, so no hard count changes
print("soft F1 after", round(soft_f1(ps, ys), 4))go deeper
Know that the 0.5 cut point is a default, not a rule, and that moving it trades precision against recall without retraining anything.
Explain how soft counts are built from summed sigmoid outputs to make F1 differentiable, and say why cross-entropy plus a tuned threshold is the cheaper first move.
Show the operating discipline: tune on a split that is not the reported one, check the positive count behind the chosen point, and name batch-size dependence and lost calibration as the price of a metric-shaped loss.
Own the decision of where the operating point lives in the system — baked into the weights, or a monitored, re-tunable parameter with a drift trigger — and be able to justify that choice to the team that owns the precision-recall trade.
## Two levers, one metric F1 is the harmonic mean of precision and recall, computed from hard counts of true positives, false positives and false negatives. Those counts come from comparing a *thresholded* prediction to a label, so F1 depends on the model through a step function twice over: once in the counts and once in the threshold. Neither is differentiable. You have two distinct ways to attack it, and knowing which to reach for first is the substance of this question. **Lever one: leave training alone and move the cut.** Train with cross-entropy, which is a proper scoring rule — its minimiser is the true conditional probability, so it produces a well-ordered, reasonably calibrated score. Then sweep the decision threshold over a held-out split and pick the value maximising F1, or the one hitting a required precision floor. This is a one-dimensional search over a metric you can evaluate directly; no gradient needed. **Lever two: change the objective so the metric itself is differentiable.** Replace each hard count with its expectation under the model's own probabilities: - `soft_tp = sum over i of p_i * y_i` - `soft_fp = sum over i of p_i * (1 - y_i)` - `soft_fn = sum over i of (1 - p_i) * y_i` and define `soft_F1 = 2 * soft_tp / (2 * soft_tp + soft_fp + soft_fn)`, minimising `1 - soft_F1`. Every term is a smooth function of the sigmoid outputs, so gradients flow. Note that no threshold appears anywhere: the objective is threshold-free by construction, which is exactly its appeal and exactly where its problems begin. ## Why the threshold sweep is the default The 0.5 threshold is not a property of the sigmoid or of the problem — it is the value you get by accident when nobody chooses one. Under a 2% positive rate, the score distribution is squeezed hard toward zero and the F1-optimal cut frequently lands somewhere well below 0.5. A model that looks broken at the default is very often a model that has been read at the wrong point. The sweep is also cheap, auditable and reversible. It changes no weights, so you can re-run it when the class balance shifts without retraining. It keeps the probabilities interpretable for anything downstream that needs an expected value rather than a decision. And it decouples the modelling question ('does the score order examples well?') from the product question ('what precision-recall trade do we want?'), which are usually owned by different people. Three disciplines make it trustworthy. Tune on a split that is neither the training data nor the split you report on, or the reported F1 is optimistically biased by the selection. Check how many positives that split contains: with a few dozen, the argmax of the F1-versus-threshold curve is mostly noise, and you should prefer a smoothed or more conservative choice. And treat the threshold as a parameter with a shelf life — it is a function of the prevalence and score distribution at tuning time, and both drift. ## When soft-F1 earns its place There are cases the sweep does not cover. In extreme multi-label settings — thousands of labels, each with its own operating point — per-label threshold tuning is a large, noisy estimation problem in itself, and an objective that already optimises the metric shape can be the cleaner design. Similarly, when the metric is very far from what cross-entropy prioritises, the surrogate can steer the *representation*, not merely the cut point, toward the positives that matter. But price it honestly: - **Non-decomposability.** Mean cross-entropy over a batch is an unbiased estimate of mean cross-entropy over the corpus, because it is an average of per-example terms. F1 is a ratio of sums, so a batch value is a biased and noisy estimate of the corpus value — biased in a way that depends on batch size and on how many positives happen to land in the batch. Batch size becomes a semantic hyperparameter, not just a throughput knob, and a batch with no positives gives a degenerate or undefined objective. - **Calibration loss.** The soft counts reward pushing every probability toward 0 or 1, because that is what makes the soft counts match hard counts. The resulting scores are no longer usable as probabilities, which forfeits the flexibility that made the threshold sweep attractive in the first place. - **Optimisation behaviour.** The objective is non-convex and its gradient for one example depends on the whole batch's counts, which makes it noisier and more finicky early in training than a per-example loss. A common compromise is to warm up with cross-entropy and switch or blend the metric surrogate in later. ## How to answer this in a loop Say the order out loud: check the threshold before changing the loss; verify the ordering quality is actually adequate before blaming the objective; and reach for a metric-shaped surrogate only when the operating point genuinely cannot be chosen post-hoc. Then name the costs of that surrogate unprompted — the batch dependence and the ruined calibration — because that is what shows you have run it rather than read about it.
- Production prevalence drops to a third of what it was at training time. What happens to your tuned threshold?It is now wrong. The threshold was chosen against a specific score distribution and positive rate; with fewer positives, the same cut admits relatively more false positives and precision falls, so the F1-optimal point moves. Treat the threshold as a monitored parameter with a re-tuning trigger on prevalence drift, and keep a labelled recent slice available so you can re-sweep without retraining the model.
- Why is a per-batch soft-F1 value not an unbiased estimate of corpus F1?Because F1 is a ratio of sums, not a mean of per-example terms. The expectation of a ratio is not the ratio of expectations, so averaging batch F1 values does not converge to corpus F1, and the gap depends on batch size and on how many positives fall in each batch. With rare positives, some batches contain none at all, which makes the objective degenerate for that step.
- You tuned the threshold on the test set and F1 looks great. What is wrong?The reported number is optimistically biased: you selected the single best of many thresholds against the same data you are reporting on, so the score includes the selection's luck. Use three splits — train, a tuning split for the threshold, and an untouched test split for the reported figure — and note how far the tuning-split F1 sits above the test figure as a sanity check on the size of that bias.
saying these in an interview costs you the question
- Reaches for a custom loss before sweeping the threshold
- Tunes the decision threshold on the test set
- Thinks 0.5 is required because the output is a sigmoid
- Assumes batch-level F1 is an unbiased estimate of corpus F1
- Ignores that soft-F1 training destroys probability calibration
- Never re-tunes the threshold when class prevalence shifts