How do you read per-class precision and recall off a 10-class confusion matrix?
answer
- rows are truth, columns are guesses
- one class against all the rest
- diagonal cell over its row total
- same cell, column total instead
- the row shows where errors leak
basics
~20 sWith true classes as rows and predicted classes as columns, a class's recall is its diagonal cell divided by its row total, and its precision is that same diagonal cell divided by its column total.
solid answer
~50 sFix the convention first: row = the true class, column = the predicted class. For class `c`, the diagonal cell (c, c) is its true positives. Everything else in row c is a true c that leaked out, so recall(c) = cell(c,c) / row-total(c). Everything else in column c is some other digit wrongly called c, so precision(c) = cell(c,c) / column-total(c). That is exactly a one-vs-rest 2x2 table for class c: TP is the diagonal cell, FN is the rest of the row, FP is the rest of the column, TN is everything outside both. Reading ten such tables tells you far more than one headline score: on handwritten digits you typically find eight classes near 0.98 recall and the errors piling up in a couple of rows, with 5 leaking into 3 and 8. The off-diagonal pattern, not just the per-class number, is what points at the fix.
go deeper
Memorise the layout: rows are true classes, columns are predictions. Recall divides the diagonal cell by its row; precision divides it by its column. Be ready to compute both for a named class from a small matrix on a whiteboard.
Explain how each class collapses to its own one-vs-rest 2x2, naming which cells become TP, FP, FN and TN. Expect to be asked why ten such tables exist inside one matrix and why every one of them is imbalanced.
Show that you diagnose from the matrix, not just report it. Distinguish a class leaking out of its row from a class contaminating its column, say what each implies about the data or the boundary, and refuse to quote rates on classes with tiny support.
Own what the team looks at. Decide whether per-class tables with support are part of every model review, set a minimum test support per class so the numbers mean something, and push back when a single headline score is used to sign off a many-class system.
## The layout A multiclass confusion matrix for K classes is a K x K table of counts. The near-universal convention is **rows are the true class, columns are the predicted class**, so cell (i, j) counts samples whose true label is i and whose predicted label is j. The diagonal holds the correct predictions; everything off the diagonal is an error, and *which* off-diagonal cell it lands in says what the model confused it with. Two totals matter: - **Row total for class c** = how many samples truly are class c. This is the class's *support*, and it is fixed by the data, not by the model. - **Column total for class c** = how many samples the model *called* class c. This is chosen by the model and moves when you retrain or shift a decision boundary. ## Collapsing to one-vs-rest Per-class metrics are defined by pretending, one class at a time, that the problem is binary. For class c: - `TP = cell(c, c)` - `FN = row-total(c) - cell(c, c)` — true c's predicted as something else - `FP = column-total(c) - cell(c, c)` — other classes predicted as c - `TN = N - TP - FN - FP` — everything that neither is nor was called c From that 2x2 the familiar definitions apply unchanged: ``` recall(c) = TP / (TP + FN) = cell(c,c) / row-total(c) precision(c) = TP / (TP + FP) = cell(c,c) / column-total(c) F1(c) = 2 * precision * recall / (precision + recall) ``` So a 10-class matrix silently contains ten different binary problems, each with its own positive rate. Class 1 might be 11% of the data and class 5 about 9%; the one-vs-rest views are therefore all imbalanced, and each one's TN cell is enormous. ## Reading the digit matrix Take a model on the ten handwritten-digit classes. A typical matrix shows eight rows almost entirely on the diagonal and two rows — 5 and 8 — carrying most of the mass off it. Row 5 might read: 5 predicted as 5 in 71% of true 5s, as 3 in 12%, as 8 in 11%. Column 5, meanwhile, is nearly clean: few other digits get called 5. That is the signature of **low recall, high precision** for class 5. The model is reluctant to say 5, so when it does it is usually right, but many real 5s escape into neighbouring shapes. The opposite signature — high recall, low precision — means the model over-uses that label: it catches nearly every true instance but sweeps in others too, and the mess shows up in the class's *column*, not its row. Those two diagnoses call for different work. Leaking out of a row usually means the class is under-represented, under-weighted, or genuinely visually close to its neighbour: you look at the specific confusion pair. Contaminating a column means the class is being used as a dumping ground: you look at what makes those other classes indistinguishable from it. ## Why the pair structure matters Recall alone tells you *how much* of a class is lost; the row tells you *where it went*. If 23% of true 5s are split between exactly two neighbours, that is a targeted problem — a discriminating feature, more examples of that pair, or in the extreme a labelling policy that never distinguished them cleanly. If a class's errors are smeared evenly across all nine other columns, that is generic underfitting, and no amount of pair-specific work will help. Note also that confusion is rarely symmetric. Many true 5s may be called 8 while few true 8s are called 5; the (5,8) and (8,5) cells are independent counts. Anyone who reports "5 and 8 get confused" without saying in which direction has not read the matrix. ## Practical cautions - **Always show support.** A recall of 0.50 on a class with four test samples is two coin flips, not a finding. Print the row total next to every per-class score. - **Normalising changes the question.** Dividing each cell by its row total gives you recall-style rates and makes rows comparable; dividing by column totals gives precision-style rates. Say which you did — a normalised matrix with no stated direction is unreadable. - **Do not compare a class's recall against another class's precision.** They are different denominators from different one-vs-rest tables. - **Small classes drive the eventual averaging.** Whatever aggregation you apply later starts from these per-class numbers, so mis-reading them corrupts everything downstream.
- For digit 5 you see recall 0.71 but precision 0.93 — what does that pattern tell you?The model rarely says 5, and is usually right when it does, but loses almost a third of the real 5s to other labels. The damage is inside row 5, so look at which columns absorbed them; the fix is a discriminating feature or more weight on that class, not tightening precision further.
- What do off-diagonal pairs like 5 into 8 tell you that per-class recall alone does not?Recall says how much of the class was lost; the pair says where it went. Errors concentrated in one neighbour point to a specific visual or definitional overlap you can attack directly. Errors spread evenly across all other classes point to general underfitting instead.
- Why must you report class support alongside per-class recall?Recall is a ratio whose denominator is the class's row total. With a handful of test samples the ratio jumps in coarse steps and carries almost no information, so a low score may be noise and a perfect score may be luck. Support tells the reader how much to trust each number.
saying these in an interview costs you the question
- Divides the diagonal cell by the column total and calls it recall
- Reads only the overall score and never opens the matrix
- Assumes a class's precision and recall move together
- Ignores class support when quoting per-class rates
- Assumes confusion between two classes is symmetric
- Presents a normalised matrix without saying which direction