skip to content

Confusion Matrix Measures

The four cells of a confusion matrix and the rates built on them - accuracy, precision, recall, specificity, F1 - and where each one quietly misleads. Screens open here.

on this pageshow

explore

questions

12

In a binary classifier's confusion matrix, what are the four cells and how is accuracy computed?

level: juniorimportance: must knowfreq 84%

answer

  1. two axes: predicted versus actual
  2. four cells, one per combination
  3. second word is what the model said
  4. correct predictions over the total
  5. error rate is the complement

basics

~20 s

The four cells count true positives, false positives, true negatives and false negatives - one per pairing of predicted and actual label. Accuracy is (TP + TN) divided by all four counts; the error rate is 1 minus accuracy.

solid answer

~50 s

A confusion matrix cross-tabulates the predicted label against the true label. For a binary problem that gives four counts: true positives (predicted positive, actually positive), false positives (predicted positive, actually negative), true negatives (predicted negative, actually negative) and false negatives (predicted negative, actually positive). Every scored row lands in exactly one cell, so the four counts sum to the size of the evaluation set. Accuracy is the share of rows on the correct diagonal: `accuracy = (TP + TN) / (TP + TN + FP + FN)`. The error rate is the off-diagonal share, `(FP + FN) / total`, which is exactly `1 - accuracy`. The important thing to say out loud is that accuracy pools both error types into one number and weights every row equally, so which cell your mistakes fall into is invisible in the total.

go deeper

for a junior

Be ready to draw the two-by-two table, name all four cells without hesitating, and compute accuracy and error rate from raw counts on a whiteboard. Practise the reconstruction drill: flags issued, flags correct, true positives total.

for a middle

Explain that accuracy pools both error types and weights every row equally, and show what that hides by giving two matrices with the same accuracy but opposite error profiles.

for a senior

Insist on seeing the full matrix, not a scalar, in any model review, and name which class you call positive before quoting numbers. Point out that the off-diagonal cells drive the operational cost, not the diagonal.

for a principal

Own the convention: fix across the organisation which class is positive, and require the raw four counts plus the evaluation-set class mix alongside any headline number so reviewers can recompute whatever metric they care about.

## What the matrix is A **confusion matrix** is a small table that cross-tabulates what a classifier predicted against what was actually true, for every row in an evaluation set. For a binary problem - two possible labels, conventionally called *positive* and *negative* - the table has two rows and two columns, giving four counts. | | actually positive | actually negative | |---|---|---| | **predicted positive** | true positive (TP) | false positive (FP) | | **predicted negative** | false negative (FN) | true negative (TN) | The naming rule is worth internalising because candidates fumble it under pressure: the **second word is what the model said**, and the **first word says whether the model was right**. A *false negative* is therefore a row the model called negative and was wrong about - an actual positive that slipped through. A *false positive* is a row the model called positive and was wrong about - a false alarm. Every scored row falls into exactly one cell, so `TP + FP + TN + FN = N`, the size of the evaluation set. That single fact is what makes the matrix the common ancestor of essentially every classification metric: each metric is just a different ratio built from these four numbers. ## Accuracy and error rate **Accuracy** is the share of rows the model got right - the diagonal of the table over the whole table: ``` accuracy = (TP + TN) / (TP + TN + FP + FN) ``` **Error rate** (sometimes called misclassification rate) is the complement, the off-diagonal share: ``` error_rate = (FP + FN) / (TP + TN + FP + FN) = 1 - accuracy ``` They carry identical information; teams differ only in which they prefer to quote. Error rate reads better when you are describing improvement, because relative change is meaningful: dropping from 4% error to 2% error is a halving, while '96% to 98% accuracy' sounds like a two-point nudge. ### A worked example Score 200 rows. 40 are actually positive, 160 actually negative. The model flags 50 rows as positive; 32 of those flags are correct. - TP = 32 (flagged and truly positive) - FP = 50 - 32 = 18 (flagged but actually negative) - FN = 40 - 32 = 8 (truly positive but not flagged) - TN = 160 - 18 = 142 (correctly left alone) Check the total: 32 + 18 + 8 + 142 = 200. Accuracy = (32 + 142) / 200 = 174 / 200 = 0.87. Error rate = (18 + 8) / 200 = 0.13. Being able to reconstruct all four cells from partial information like this - 'the model flagged 50, 32 were right' - is a common whiteboard drill. ## The three properties that decide where accuracy is safe 1. **It weights every row equally.** One misclassified row moves the number by 1/N regardless of which row it was. If a missed case costs a thousand times more than a false alarm, accuracy does not know that. 2. **It pools the two error types.** A model with 18 false positives and 8 false negatives, and a model with 8 false positives and 18 false negatives, score identically. Two very different operational behaviours, one number. 3. **It depends on the class mix of the evaluation set.** Because it sums over rows, the class that supplies most of the rows supplies most of the score. On a set that is 99% negative, accuracy is essentially a report on the negative class with a rounding error attached. Property 3 is what makes accuracy collapse under class imbalance, and it is the reason interviewers almost always follow this question with a low-base-rate scenario. Properties 1 and 2 are why the rest of the confusion-matrix family exists at all: the four cells are kept separate precisely so you can ask which kind of mistake the model is making, not just how many. ## Which label is 'positive' The choice of which class to call positive is a convention, not a property of the data - usually the rarer, more interesting or more actionable class (the defect, the churner, the alert). Swapping the convention swaps TP with TN and FP with FN. Accuracy is unchanged, because the diagonal is the diagonal either way; every asymmetric metric flips. Say the convention out loud when you present a matrix, because a reader who assumes the other one will read your false positives as false negatives. ## Multiclass For k classes the matrix is k-by-k: rows are predictions, columns are truth (or the transpose - label your axes). The diagonal holds correct predictions, and accuracy generalises directly as the sum of the diagonal over the sum of the whole matrix. The off-diagonal cells are the useful part, because they tell you *which* pairs of classes are being confused with each other - information no single scalar preserves.

  • A model flags 50 rows out of 200 as positive, 32 flags are correct, and 40 rows are truly positive. Fill in the four cells.
    TP = 32. FP = 50 - 32 = 18. FN = 40 - 32 = 8. TN = 200 - 32 - 18 - 8 = 142. The four counts must sum to 200, which is the arithmetic check. Accuracy is (32 + 142) / 200 = 0.87 and the error rate is 0.13.
  • What changes in the matrix if you relabel which class counts as positive?
    True positives and true negatives swap, and false positives and false negatives swap. Accuracy and error rate are unchanged because the diagonal is unchanged. Every asymmetric quantity built on the cells flips, which is why you should always state which class you call positive before showing numbers.
  • Why do teams quote error rate instead of accuracy?
    They carry the same information, but error rate makes progress legible on a relative scale. Going from 4% error to 2% is a halving of mistakes; the same change described as 96% to 98% accuracy sounds like a marginal two-point move. For near-perfect models the error rate is the only one with usable resolution.

It is a two-by-two seating chart: rows for what the model claimed, columns for the truth. The people on the diagonal are in the right seat; accuracy is just the fraction seated correctly.

saying these in an interview costs you the question

  • Calls a false negative a row the model predicted positive
  • Thinks the four cells can overlap or need not sum to N
  • Says accuracy is TP over all predicted positives
  • Cannot reconstruct the cells from flag counts and true totals
  • Believes accuracy changes when you swap which class is positive

context

open as a page

Why is 99.7% accuracy uninformative for a defect detector when 0.3% of units are defective?

level: juniorimportance: must knowfreq 88%

basics

~20 s

Because a model that never predicts 'defective' already scores 99.7% on that data. The majority class supplies almost every row, so accuracy measures performance on good units and says nothing about whether a single defect was ever caught.

open as a page

How do you read per-class precision and recall off a 10-class confusion matrix?

level: juniorimportance: must knowfreq 62%

basics

~20 s

With true classes as rows and predicted classes as columns, a class's recall is its diagonal cell divided by its row total, and its precision is that same diagonal cell divided by its column total.

open as a page

What is the difference between precision and recall, and why do they trade off?

level: juniorimportance: must knowfreq 88%

basics

~20 s

Precision is the share of predicted positives that are truly positive; recall is the share of actual positives the model catches. Loosening the decision rule labels more examples positive, which raises recall and usually lowers precision.

open as a page

How do macro, micro and weighted averaging of multiclass F1 differ?

level: middleimportance: must knowfreq 78%

basics

~20 s

Macro averages per-class F1 scores with equal weight, so a rare class counts as much as a common one. Micro pools every class's true positives and errors first, so frequent classes dominate. Weighted averages per-class scores by class support.

open as a page

How does balanced accuracy differ from plain accuracy on an imbalanced test set?

level: middleimportance: should knowfreq 60%

basics

~20 s

Balanced accuracy averages the per-class recalls, giving every class equal weight, while plain accuracy weights each class by how many rows it contributes. On skewed data the two diverge sharply, and a constant predictor scores 1/k instead of the majority share.

open as a page

Why is F1 the harmonic mean of precision and recall rather than the arithmetic mean?

level: middleimportance: should knowfreq 62%

basics

~20 s

The harmonic mean is dominated by the smaller of the two values, so F1 stays low unless precision and recall are both decent. Precision 0.90 with recall 0.10 gives F1 0.18, while the arithmetic mean would flatter it at 0.50.

open as a page

Why can a sepsis alert with 95% specificity still be wrong most times it fires?

level: middleimportance: should knowfreq 48%

basics

~20 s

Specificity is measured only among patients who do not have sepsis, so it ignores how rare sepsis is. At a 2% rate, 5% of the huge non-sepsis group yields far more false alarms than true cases, leaving precision near 25%.

open as a page

A defect model reads 94% accuracy on a test set rebalanced to 50/50, but live defects run at 0.3%. What does that number tell you?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Almost nothing about production. Accuracy depends on the class mix it was measured on, so a 50/50 test set reports balanced accuracy, not the live figure at a 0.3% defect rate. Recombine the per-class rates at the true base rate.

open as a page

Why report Cohen's kappa instead of raw agreement for a 5-grade severity classifier?

level: seniorimportance: should knowfreq 38%

basics

~10 s

Cohen's kappa subtracts the agreement expected by chance from the observed agreement and rescales the remainder, so a model that merely mimics the grade distribution scores near zero even when raw agreement looks high.

open as a page

What does Hamming loss measure for a news tagger that assigns several topics per article?

level: middleimportance: nice to knowfreq 27%

basics

~20 s

Hamming loss is the fraction of individual label decisions that are wrong across all articles and all candidate topics. Each missed tag and each spurious tag counts once, lower is better, and partly-correct tag sets earn partial credit.

open as a page

How do you choose beta in F-beta when a missed case costs far more than a false alarm?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Beta sets how many times more recall matters than precision: F2 weights recall twice as heavily, F0.5 half as heavily, F1 equally. Derive beta from the cost ratio of a missed case to a false alarm.

open as a page