How does balanced accuracy differ from plain accuracy on an imbalanced test set?
answer
- accuracy weights rows, this weights classes
- one recall per class, then average
- chance level is one over k
- same as reweighting rows by class frequency
- strong on rates, silent on volume
basics
~20 sBalanced accuracy averages the per-class recalls, giving every class equal weight, while plain accuracy weights each class by how many rows it contributes. On skewed data the two diverge sharply, and a constant predictor scores 1/k instead of the majority share.
solid answer
~50 sPlain accuracy is correct rows over total rows, so a class with 99% of the rows carries 99% of the weight. Balanced accuracy instead computes recall **within each true class** - the share of that class the model got right - and averages those numbers with equal weight. For two classes that is `(sensitivity + specificity) / 2`; for k classes it is the mean of the k per-class recalls. The payoff is a fixed reference point: a constant predictor scores 1/k, so 0.5 on a binary problem, no matter how skewed the data is. It equals plain accuracy after reweighting rows by inverse class frequency. Its blind spot is volume: because each recall is conditioned on the true class, a huge negative class can absorb thousands of false alarms while barely denting its own recall, so balanced accuracy can read 0.94 while the flagged queue is almost entirely false alarms.
code
python · 16 lines# Wafer inspection, 100,000 units: 300 defective, 99,700 good.
tp, fn = 276, 24 # defective units: caught, missed
fp, tn = 3988, 95712 # good units: false alarms, correct
total = tp + fn + fp + tn
accuracy = (tp + tn) / total
recall_defect = tp / (tp + fn)
recall_good = tn / (tn + fp)
balanced = (recall_defect + recall_good) / 2
baseline = (tn + fp) / total # always predict 'good'
print(round(accuracy, 4)) # 0.9599 -- below the baseline
print(round(recall_defect, 4)) # 0.92
print(round(recall_good, 4)) # 0.96
print(round(balanced, 4)) # 0.94 -- chance level is 0.5
print(round(baseline, 4)) # 0.997go deeper
Recall the one-line definition - the mean of the per-class recalls - and that a constant predictor scores 0.5 on a binary problem no matter how skewed the data is.
Explain the mechanics from the confusion-matrix cells: each class is divided by its own size before averaging, which is why it equals accuracy after inverse-frequency reweighting. Compute both numbers from the same four counts.
Show the judgment: quote it beside the plain accuracy and the majority-class baseline, flag the false-alarm volume it cannot see, and refuse to lean on it when one class has only a handful of rows.
Own the framing that equal class weighting is a value judgement, not neutrality. Decide deliberately whether the organisation's headline metric should reflect prevalence, equal classes, or explicit error costs, and make that choice visible in every review.
## Definition **Balanced accuracy** is the unweighted mean of the per-class recalls. For each true class *c*, its **recall** (also called that class's true-positive rate, and for the negative class in a binary problem, *specificity*) is the share of rows genuinely belonging to *c* that the model assigned to *c*. Balanced accuracy averages those k numbers: ``` balanced_accuracy = (1/k) * sum over classes of recall_c ``` Binary case, written with the confusion-matrix cells: ``` recall_positive = TP / (TP + FN) # sensitivity recall_negative = TN / (TN + FP) # specificity balanced_accuracy = (recall_positive + recall_negative) / 2 ``` Compare with plain accuracy, `(TP + TN) / N`. The structural difference is the **denominator of each term**. Accuracy divides by the whole dataset once, so each class contributes in proportion to its size. Balanced accuracy divides each class by *its own* size first, then averages - so a class with three rows counts as much as a class with three million. ## Worked contrast Wafer inspection, 100000 units, 300 defective and 99700 good. The model catches 276 defects (missing 24) and raises 3988 false alarms on good units. - Plain accuracy = (276 + 95712) / 100000 = **0.9599** - Defect-class recall = 276 / 300 = **0.92** - Good-class recall = 95712 / 99700 = **0.96** - Balanced accuracy = (0.92 + 0.96) / 2 = **0.94** Now the diagnostic value: the majority-class baseline for plain accuracy here is 0.997, so **0.9599 is below the do-nothing rule** even though it looks high. Balanced accuracy is 0.94 against a trivial-model reference of 0.5, so it correctly reports that the model has learned a great deal. Two metrics, opposite verdicts, and balanced accuracy is the one whose scale is stable under the class mix. ## The fixed reference point The most practical property is that the trivial model's score does not move with the imbalance. A predictor that always outputs the majority class gets recall 1 on that class and recall 0 on all others, so its balanced accuracy is exactly **1/k** - 0.5 for binary, 0.333 for three classes - no matter whether the majority holds 60% or 99.99% of the rows. A uniformly random guesser lands at the same 1/k in expectation. That means 'how far above chance is this?' is answerable by eye, which is precisely what plain accuracy loses under imbalance. On a three-class vibration-fault log split 92 / 5 / 3, a model that classifies the dominant healthy class flawlessly and never predicts either fault class scores 0.92 plain accuracy and 0.333 balanced accuracy. The second number is the honest one. ## The equivalent framing Balanced accuracy equals plain accuracy computed on a *reweighted* evaluation set, where each row is weighted by the reciprocal of its class's frequency so that every class carries the same total weight. That is the same thing as saying: it is the accuracy you would measure on a test set artificially rebalanced to equal class sizes, assuming the model's behaviour within each class is unchanged. Holding those two framings together is what lets you convert between the numbers when someone hands you the wrong one. ## Where it still misleads Balanced accuracy is not a general-purpose fix; it swaps one bias for another. 1. **It is blind to volume, so it hides false-alarm load.** Each recall is conditioned on the true class. In the wafer example the good class absorbed 3988 false alarms and its recall only fell to 0.96, yet those 3988 alarms sit alongside just 276 real catches - the inspection queue is more than 93% noise. Balanced accuracy reads 0.94 and says nothing about that. Whenever the action taken on a flag is expensive, the flagged-set view is the one that matters and belongs in the report next to this number. 2. **It asserts that classes matter equally.** That is a value judgement, not a neutral one. If missing a defect costs a thousand times a false alarm, equal weighting is as arbitrary as prevalence weighting - just arbitrary in a different direction. 3. **It is still a single operating point.** It summarises the hard labels the model emitted under one cutoff, so it says nothing about how the model would behave if that cutoff moved. 4. **On tiny classes it is noisy.** A class with 8 rows contributes 1/k of the whole score, and one extra mistake shifts its recall by 12.5 points. Report the per-class support alongside, or the average is being driven by a handful of rows. ## How to use it Treat balanced accuracy as the *scale-corrected* headline: quote it when you need one number on skewed data and want a stable reference, always next to the plain accuracy, the majority-class baseline and the per-class supports. When a decision hangs on it, drop the scalar entirely and show the confusion matrix - the four counts contain everything the summaries throw away.
- What does a balanced accuracy of 0.5 mean on a binary problem?Chance. It is what a coin flip gets in expectation and exactly what a constant predictor gets, because one class scores recall 1 and the other recall 0. Unlike plain accuracy, that reference does not move with the class mix, so 0.5 always means the model carries no class-discriminating information.
- Why can balanced accuracy look strong while the flagged queue is mostly false alarms?Because each recall is conditioned on the true class. A majority class of 99700 rows can absorb 3988 false alarms and still show recall 0.96, even though those alarms outnumber the 276 real catches more than fourteen to one. Volume information lives in the flagged set, not in class-conditional rates.
- How is balanced accuracy related to reweighting the evaluation set?It is exactly plain accuracy after weighting each row by the inverse of its class frequency, so every class contributes equal total weight. Equivalently, it is the accuracy you would read on a test set resized to equal class sizes, provided the model's within-class behaviour is unchanged.
saying these in an interview costs you the question
- Defines it as the average of precision and recall
- Thinks its chance level shifts with the class mix
- Believes it accounts for false-alarm volume
- Averages the per-class recalls weighted by class size
- Presents it as a fix for imbalance rather than a rescaled metric
- Ignores per-class support when one class has a handful of rows