skip to content

Why does a per-RoI mask branch predict one binary mask per class instead of a per-pixel softmax over classes?

level: middleimportance: should knowfreq 48%

answer

  1. another branch already decides the class
  2. sigmoid per pixel, not softmax across classes
  3. only one of the K maps sees the loss
  4. the mask answers foreground or not

basics

~20 s

Decoupling. The classification branch already names the class, so each mask channel only has to answer whether a pixel is inside the object. A per-pixel softmax would make classes compete for pixels and tie mask quality to classification confidence.

solid answer

~50 s

The mask branch is bolted onto a detection head that already produces a class and a box for each region of interest, so mask prediction and class prediction do not need to be the same decision. The branch emits `K` low-resolution binary maps per region, one per class, each with independent per-pixel sigmoid outputs, and the training loss is applied only to the channel of the ground-truth class — the other channels get no gradient from that region. At inference the classification branch picks which channel to read. A per-pixel softmax over classes instead forces the classes to compete at every pixel: mass given to `dog` is taken from `cat`, so a shaky class decision degrades the shape as well. Decoupling turns mask prediction into the much easier question "is this pixel foreground for the object in this box", which generalises better and is why mask heads are built this way.

go deeper

for a junior

Recall that the head predicts one binary mask per class for each region and that a separate branch decides the class. Knowing the mask answers foreground versus background is the level-appropriate takeaway.

for a middle

Explain the coupling a per-pixel softmax creates between class confidence and silhouette quality, and state precisely which channel the loss touches during training and which one is read at inference.

for a senior

Bring the operational consequences: mask detail is capped by the fixed low resolution and by the box, a misclassified object still yields a clean mask under the wrong label, and raising mask resolution is a cost decision you should be able to justify.

for a principal

Frame it as an interface decision between heads. Keeping mask prediction independent of classification means each head can be improved, retrained or swapped without regressing the other; argue when that separation is worth its extra parameters.

## Where the branch sits Instance masks are usually produced by adding a third head to a two-stage detector. The first stage proposes regions; for each region the head already runs a classification branch (which class, or background) and a box-regression branch (refine the rectangle). The mask branch is a small stack of convolutions on the same pooled region features, ending in a set of low-resolution maps — a common choice is 28x28 — that are then resized onto the predicted box to give a full-resolution mask. The design question is what those maps should represent. ## Two candidate parameterisations **Per-pixel multinomial.** Emit one map with `K` channels and take a softmax across channels at each pixel, exactly as semantic segmentation does. Each pixel is assigned to one class or background, and the classes compete: raising `dog` at a pixel necessarily lowers `cat` there. **Per-class binary.** Emit `K` independent maps, each with a per-pixel sigmoid. Map `k` answers, for every pixel in the region, "is this pixel part of the object, assuming the object is of class `k`". No cross-channel normalisation, so the channels do not compete. Mask heads use the second, and the training rule is the other half of the design: the loss is the average binary cross-entropy over the pixels of the ground-truth class's channel *only*. The remaining `K-1` channels receive no gradient from that region. ## Why decoupling helps **The class question is already answered elsewhere.** Making the mask branch also discriminate between classes duplicates work the classification branch does with better features and a cleaner objective — it sees the whole region, not one pixel at a time. **Competition couples two error modes.** Under a softmax, a region whose class is genuinely ambiguous spreads probability across several channels, so no channel confidently claims the object's pixels and the *shape* degrades even when the pixel set is obvious. Under independent sigmoids, an ambiguous class costs you the label, not the silhouette; the mask stays a clean foreground/background call. **The per-pixel task becomes trivial to state.** "Foreground or not, within this box" is a binary problem with strong local evidence and roughly balanced positives and negatives inside a tight region. That is a far easier learning problem than a `K`-way decision at every pixel with almost all pixels belonging to a handful of classes. **Class-specific shape priors survive.** Because each channel is trained only on regions of its own class, channel `k` can specialise in what a `k` looks like — the silhouette of a person is not the silhouette of a car — without the specialisation being expressed as a competition. ## Practical consequences *Inference reads one channel.* The classification branch's chosen class selects the channel; the score attached to the instance is the classification score, not a mask score. A consequence is that a badly classified object can still produce a good-looking mask that ends up filed under the wrong class. *Only one channel is supervised per region.* People are often surprised that the other channels are not pushed towards zero. They are simply not part of that region's loss, which keeps the objective from spending capacity on suppressing masks nobody will read. *The mask lives in the region's own frame.* It is predicted at fixed low resolution relative to the box, then resized. That makes the branch cheap and box-size-independent, but it caps boundary detail: thin structures and sharp corners are smoothed by the resize, and a mask can be no better than the box it is painted into. Raising the mask resolution recovers some detail at extra compute. *Overlap is allowed.* Because masks are predicted per region independently, two instances can claim the same pixel. That is fine for instance segmentation and is precisely the loose end that has to be tied off when the output is merged into a panoptic labelling. ## What a weak answer looks like Saying "binary is simpler" without naming the coupling that softmax introduces, or claiming that all `K` channels are trained on every region, or asserting that the mask branch classifies the object. The point to land is that the head has *already* decided the class, so making the mask decision independent of it is free accuracy.

  • Which of the K mask channels contributes to the loss for a given training region?
    Only the channel of that region's ground-truth class. Its average per-pixel binary cross-entropy is the mask loss; the other channels receive no gradient from that region at all. At inference the class predicted by the classification branch selects which channel is read out as the instance's mask.
  • What is lost by predicting the mask at a small fixed resolution in the region's own frame?
    Boundary detail. The low-resolution map is resized onto the box, so thin limbs, wires and sharp corners come back smoothed, and the mask can never be better than the box containing it. The upside is cost that does not scale with object size; raising the mask resolution buys back detail at more compute.
  • How does this design let two instances claim the same pixel?
    Each region's mask is predicted independently from its own pooled features, with no normalisation across regions, so two overlapping detections can both mark a contested pixel as foreground. Instance segmentation tolerates that; a panoptic output cannot, so the overlap is resolved at merge time by a fixed precedence rule.

saying these in an interview costs you the question

  • Says the mask branch must also classify the object
  • Claims all K mask channels get gradient from every region
  • Thinks a per-pixel softmax over classes is equivalent here
  • Believes the mask is predicted in full image coordinates
  • Assumes mask resolution is proportional to box size

context