skip to content

A segmentation model reports 95% pixel accuracy but 42 mIoU — what explains the gap?

level: seniorimportance: should knowfreq 58%

answer

  1. one metric weights pixels, one weights classes
  2. IoU denominator holds both error types
  3. rare classes carry equal weight
  4. the mean hides the per-class table
  5. aggregate globally, not per image

basics

~20 s

Pixel accuracy is dominated by whichever classes own most of the pixels, while mIoU averages per-class intersection over union with every class weighted equally. Rare classes scoring near zero barely move accuracy but drag the mean down hard.

solid answer

~50 s

Pixel accuracy is the fraction of pixels labelled correctly across the whole image, so in a street scene where road and building cover most of the frame, getting those two right buys most of the 95 percent immediately. Mean IoU computes, per class, `TP / (TP + FP + FN)` over that class's pixels, then takes an unweighted mean over classes. A pole class covering 0.3 percent of pixels contributes 0.3 points to accuracy at most, but it is one of perhaps 19 terms in the mean, so scoring 0.05 on it costs roughly a twentieth of the whole mIoU. IoU also punishes both false positives and false negatives on that class, so smearing the pole a few pixels wide hurts twice. That gap is the metric working as intended: it is telling you the model has learned the easy majority classes and failed the small ones. Report per-class IoU alongside the mean, because the mean alone hides which classes failed.

go deeper

for a junior

Be able to write IoU from pixel counts as true positives over true positives plus false positives plus false negatives, and to say that the mean is taken over classes, not over pixels.

for a middle

Explain the weighting arithmetic concretely. Show why a class at a fraction of a percent of pixels barely touches accuracy yet carries a full one-over-class-count share of the mean.

for a senior

Demonstrate evaluation discipline: global accumulation over per-image averaging, explicit handling of ignore labels, per-class reporting, and a concrete plan for the failing classes rather than a claim that the metric is unfair.

for a principal

Own what the team optimises. Decide whether an unweighted mean matches the product's real cost of error, when a class-specific target should override the headline number, and how evaluation conventions are fixed so results stay comparable across teams and time.

## The two metrics **Pixel accuracy** is the simplest thing you can compute: count pixels whose predicted label equals the ground-truth label, divide by the number of labelled pixels. It is a single number over the whole dataset, and every pixel votes equally. That last property is the problem: pixels are not distributed equally across classes. **Intersection over union** for one class `c` is computed from that class's pixel counts: `IoU_c = TP_c / (TP_c + FP_c + FN_c)` where `TP_c` is pixels correctly labelled `c`, `FP_c` is pixels labelled `c` that are something else, and `FN_c` is pixels of `c` labelled as something else. Equivalently it is the area of the intersection of the predicted and true regions divided by the area of their union. **Mean IoU** averages `IoU_c` over the classes, unweighted. ## Why the numbers diverge so far Consider a street scene with roughly 19 classes. Road might be 35 percent of all pixels, building and vegetation another 30 percent between them, sky maybe 4 percent, and at the other end a pole class at 0.3 percent, a traffic-light class lower still. For **accuracy**, the arithmetic is a pixel-weighted average. Get road, building, vegetation, sky and car right and you have already accounted for well over 90 percent of pixels. Every rare class combined can contribute at most one or two points. Accuracy is therefore almost a measure of how well you segment the largest handful of classes, and 95 percent is unimpressive rather than good. For **mIoU**, the arithmetic is a class-weighted average: each class carries weight `1/19` regardless of its pixel share. If ten common classes score 0.75 and nine rare ones score 0.05, the mean is about 0.42. Every rare class the model gives up on costs roughly five points of mIoU, and none of them costs more than a fraction of a point of accuracy. On top of the weighting, IoU is a **stricter** per-class score than recall. Its denominator contains both false positives and false negatives, so a model that hedges by painting a rare class generously is penalised for the over-prediction, not rewarded for the extra recall. A thin two-pixel structure predicted three pixels wide and offset by one has poor IoU even though a human would call the prediction correct. ## Which to report Report **mIoU as the headline plus the full per-class IoU table**. The mean is what makes the metric comparable across models, but it is an average over things you care about differently. In practice the per-class column is where the decisions live: it tells you whether the failure is concentrated in two classes you could fix with data, or spread across every small class, which points at a resolution or architecture limit instead. Use pixel accuracy only as a sanity check. It is useful for spotting a catastrophic regression and almost useless for comparing two working models. ## Aggregation detail that changes the number There are two ways to compute mIoU over a dataset, and they do not agree. 1. **Accumulate globally.** Sum `TP_c`, `FP_c` and `FN_c` over every image in the dataset, then compute each class's IoU once from those totals and average over classes. This is the standard and the one you should use. 2. **Average per image.** Compute mIoU on each image, then average over images. This is unstable, because a class absent from an image and correctly not predicted gives `0/0`, which is undefined. Whatever convention you pick for that case — skip it, call it one, call it zero — changes the score substantially, and images containing only easy classes get the same weight as crowded ones. If two teams report different numbers on the same model, aggregation convention and the treatment of ignore or void labels are the first two things to check. ## What you do about a real gap - **Look at the per-class table first** and separate classes that are rare from classes that are thin. They have different fixes. - **Sample training crops** so rare classes appear more often than their natural pixel frequency, rather than accepting whatever a uniform random crop gives you. - **Reweight the per-pixel loss** by inverse class frequency so the gradient is not dominated by road and sky. - **Raise the effective output resolution** for thin classes, since a structure narrower than the decoder can resolve will never reach good IoU no matter how it is weighted. - **Check the annotation** on the failing classes. Thin-structure boundaries are exactly where human labels are least consistent, and a class can cap out below what the model is capable of. ## What weak answers get wrong - Treating the gap as a bug in the metric. It is a correct report of imbalanced performance. - Defining IoU with recall's denominator, `TP / (TP + FN)`, which ignores over-prediction entirely. - Assuming mIoU weights classes by pixel share. If it did, it would just be a variant of accuracy. - Quoting a mIoU without saying how it was aggregated or how ignore pixels were handled.

  • Do you accumulate IoU over the whole dataset or average per-image mIoU?
    Accumulate true positives, false positives and false negatives per class across the entire dataset, then compute each class IoU once and average. Per-image averaging is unstable because a class absent from an image yields an undefined zero over zero, and the convention chosen for that case moves the reported number noticeably.
  • Why can adding one more rare class to the label set lower mIoU sharply?
    The mean is over classes, so each class carries weight one over the class count. Adding a class the model handles badly both introduces a low term and shrinks the weight of the good ones. The model has not got worse; the metric is now asking about something it never learned.
  • The one class that matters commercially has the worst IoU. What now?
    Stop optimising the mean. Track that class's IoU as its own objective, oversample crops containing it, reweight it in the loss, and check whether its annotation quality even supports the target. Report the mean for comparability but make the decision on the class-specific number, and say so explicitly in the evaluation.

Pixel accuracy grades a spelling test by counting correct letters; mIoU grades it word by word. A page of correctly spelled common words with every rare word mangled scores brilliantly on the first and poorly on the second.

saying these in an interview costs you the question

  • Defines IoU as TP over TP plus FN
  • Thinks mIoU weights classes by pixel share
  • Calls the accuracy-mIoU gap a metric bug
  • Quotes mIoU without stating the aggregation method
  • Ignores per-class IoU and reports only the mean

context