skip to content

Detection and Segmentation

What gets built on top of a backbone: boxes with class labels, the IoU matching and mAP that score them, and per-pixel masks from an encoder-decoder with skip connections.

on this pageshow

explore

questions

17

How do semantic, instance, and panoptic segmentation differ in what they label each pixel?

level: juniorimportance: must knowfreq 74%

answer

  1. start from a counting task
  2. class label versus object identity
  3. things are countable, stuff is not
  4. which output covers every pixel once

basics

~20 s

Semantic segmentation labels each pixel with a class but no identity, so touching objects merge. Instance segmentation returns one mask per countable object. Panoptic gives every pixel exactly one class and, for countable classes, an instance id.

solid answer

~50 s

The three differ in whether identity and completeness are part of the output. **Semantic** segmentation assigns a class to each pixel and nothing else, so ten sheep standing shoulder to shoulder come back as one connected `sheep` blob. **Instance** segmentation returns a set of objects, each with its own mask and score, which is what a counting task actually needs; it normally covers only countable "thing" classes, so sky and grass pixels come back unlabelled, and two masks may overlap because each detection is predicted independently. **Panoptic** segmentation unifies both: every pixel gets exactly one segment, thing pixels carry a class plus an instance id, and "stuff" classes such as sky, road or grass carry a class with no id. Panoptic is the only one of the three that is both complete and non-overlapping by construction.

go deeper

for a junior

Be ready to state the three outputs in one sentence each and to say which one a counting task needs. Knowing that sky is stuff and a sheep is a thing is enough at this level.

for a middle

Explain why identity cannot be read off a per-pixel class map, and therefore why instance models individuate first. Be able to say which of the three outputs are complete and which are non-overlapping.

for a senior

Show you pick the task from the product question, not the other way round: counting and tracking need instances, area and coverage questions need stuff, and paying for panoptic when semantic answers the question is wasted budget and latency.

for a principal

Own the call about which annotation format the team commits to. Panoptic labels cost the most to collect and lock you into a stricter output contract; decide whether downstream consumers actually need per-pixel exclusivity or whether a cheaper format serves them.

## The three outputs All three tasks take an image and produce something denser than a box, but they disagree on two questions: *does the output distinguish individual objects*, and *does the output cover every pixel exactly once*. **Semantic segmentation** produces a label map the same size as the image: each pixel carries one class from a fixed set. That is all. There is no notion of "which sheep" — pixels of the same class that touch form one connected region, and pixels of the same class that are far apart are equally just `sheep`. The output is complete (every pixel is labelled) and non-overlapping (a pixel has one label), but it is identity-blind. **Instance segmentation** produces a *list*: for each detected object, a class, a confidence score, and a binary mask marking that object's pixels. Ten sheep produce ten masks. The output is identity-aware but neither complete nor guaranteed non-overlapping. It is not complete because instance segmentation is normally defined only over countable classes — there is no "the sky instance" — so background pixels belong to no mask at all. It is not guaranteed non-overlapping because each mask is predicted independently for its own region; two detections that both cover a contested pixel each mark it as theirs, and nothing in the formulation forbids that. **Panoptic segmentation** demands both properties at once. Its output is a single labelling in which every pixel is assigned exactly one pair: a class, plus an instance id when the class is countable. It splits the label set in two: - **Things** — countable classes with instances: person, car, sheep, bicycle. Each occurrence is its own segment with its own id. - **Stuff** — amorphous, uncountable regions: sky, road, grass, wall. All pixels of a stuff class in an image form a single segment with no id; asking "which sky" is meaningless. ## Why the distinction bites in practice The cleanest way to feel the difference is a counting task. Photograph a flock and ask how many sheep are in it. A semantic model that is *perfect* — every sheep pixel correctly labelled `sheep`, no false pixels — still cannot answer, because the flock is one merged region and the model never separated the animals. Counting connected components of that region is a hack that fails exactly when it matters: overlapping animals become one component, and one animal split by an occluding fence post becomes two. An instance model answers directly: the number of masks above the score threshold is the count. Run it the other way and the same logic shows why instance output is not always enough. If the question is "what fraction of this image is road", instance segmentation has nothing to say — road is stuff and gets no mask. Semantic or panoptic output answers it immediately. ## Consequences for the model and the loss The output shape drives the architecture. A semantic model can be a single dense predictor: one output map, per-pixel class scores, done in one pass. An instance model has to *individuate* first — typically by detecting objects and predicting a mask inside each detected region — because there is no way to read "which object" off a per-pixel class map. That is why instance segmentation is usually built by attaching a mask branch to a detection head rather than by adding channels to a semantic model. The evaluation shape follows too. Semantic output is scored by comparing label maps class by class. Instance output is scored like detection: each predicted mask is matched to a ground-truth mask and precision/recall is averaged across classes. Panoptic output is scored with panoptic quality, which has to reward both finding the right set of segments and drawing each one accurately, because both are part of the task definition. ## Common confusions - *Panoptic is not "instance segmentation with more classes."* Adding stuff classes to an instance model would still let two outputs claim a pixel; the panoptic constraint is that the final labelling is a function from pixels to segments, so contested pixels must be resolved. - *Not every panoptic pixel has an instance id.* Stuff pixels have a class only. A model that emits ids for sky has misunderstood the format. - *Semantic segmentation is not "bad" instance segmentation.* For stuff-heavy questions it is the right and cheaper tool; the failure is only when identity is required. - *Instance masks are not boxes.* A detection returns a rectangle; an instance mask returns the pixel set inside it, which for thin or curved objects is a small fraction of the rectangle.

  • Why can two instance masks overlap while a panoptic output forbids it?
    Instance segmentation is a list of independent detections, each carrying its own mask and score, so nothing stops two of them from marking the same pixel. Panoptic output is defined as a single assignment from pixels to segments, so each pixel belongs to exactly one; any contested pixel has to be resolved when the thing and stuff predictions are merged.
  • Where do stuff classes such as sky or road appear in an instance segmentation output?
    They normally do not appear at all. Instance segmentation is defined over countable thing classes, so sky and road pixels are simply left unlabelled and are treated as background by the metric. Panoptic segmentation adds them back as stuff segments carrying a class but no instance id.
  • Can you get instance-style counts by finding connected components in a semantic map?
    Only when objects never touch and are never split. Two adjacent sheep merge into one component and undercount; one sheep occluded by a post splits into two components and overcounts. The heuristic fails precisely in the crowded, occluded scenes where counting is hard, which is why the task is posed as instance segmentation.

Semantic segmentation is a colouring book where every sheep gets the same colour; instance segmentation numbers each sheep; panoptic colours the whole page and numbers only the things you could count.

saying these in an interview costs you the question

  • Claims panoptic is just instance segmentation with extra classes
  • Says a semantic map can count objects of a class directly
  • Thinks every panoptic pixel carries an instance id
  • Treats stuff classes as countable objects
  • Assumes instance masks can never overlap

context

open as a page

In object detection, how does IoU between a predicted and ground-truth box decide a detection is correct?

level: juniorimportance: must knowfreq 78%

basics

~20 s

IoU is the overlap area of a predicted and a ground-truth box divided by the area they jointly cover. A prediction counts as correct when its IoU with an unclaimed ground-truth box of the same class clears a threshold, commonly 0.5.

open as a page

In semantic segmentation, how does a CNN classifier become a dense per-pixel predictor?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Drop the flatten-and-dense head and make every layer convolutional. The dense head is rewritten as an equivalent convolution, the last layer emits one score map per class, and a decoder upsamples those maps back to the input resolution.

open as a page

Why does RoIAlign produce better instance masks than RoI pooling on a detection head?

level: middleimportance: must knowfreq 58%

basics

~20 s

RoI pooling rounds twice, the box onto the feature grid and then the bin edges, so features come from up to half a cell off. RoIAlign keeps float coordinates and samples bilinearly, preserving the alignment masks need.

open as a page

How does GIoU repair the 1 - IoU box loss when the predicted and ground-truth boxes do not overlap?

level: middleimportance: must knowfreq 60%

basics

~20 s

Two disjoint boxes have IoU 0 however far apart they sit, so a 1 - IoU loss is flat there and produces no gradient. GIoU subtracts the empty fraction of the smallest box enclosing both, which keeps shrinking as they approach, restoring a direction to move.

open as a page

In an object detector, what do anchor boxes do, and how does an anchor-free head replace them?

level: middleimportance: must knowfreq 72%

basics

~20 s

Anchors are fixed reference boxes tiled over every feature-map location; the head predicts a class score and offsets that nudge an anchor onto an object. Anchor-free heads drop them and regress the four side distances straight from each location.

open as a page

Why does U-Net concatenate encoder feature maps into its decoder instead of upsampling alone?

level: middleimportance: must knowfreq 70%

basics

~20 s

Downsampling in the encoder destroys the precise location of edges, and upsampling cannot invent it back. U-Net concatenates the matching high-resolution encoder maps into each decoder stage, so the decoder combines deep semantics with exact boundary detail.

open as a page

Why does class-wise non-maximum suppression delete correct boxes on a shelf of densely packed cartons?

level: seniorimportance: must knowfreq 66%

basics

~20 s

Suppression assumes heavily overlapping boxes of the same class are duplicates of one object. Identical cartons packed side by side genuinely overlap above the threshold, so the lower-scoring box is deleted even though it is a real, distinct object.

open as a page

Why does a per-RoI mask branch predict one binary mask per class instead of a per-pixel softmax over classes?

level: middleimportance: should knowfreq 48%

basics

~20 s

Decoupling. The classification branch already names the class, so each mask channel only has to answer whether a pixel is inside the object. A per-pixel softmax would make classes compete for pixels and tie mask quality to classification confidence.

open as a page

A detector scores about 100,000 candidate boxes per image against roughly 10 objects - what breaks in training?

level: seniorimportance: should knowfreq 42%

basics

~20 s

With a thousand background candidates per object, the summed training loss is dominated by easy background and the model drifts toward predicting nothing. Fixes constrain which candidates contribute: a fixed positive-to-negative sampling ratio, hard-negative mining, and an ignore band for ambiguous candidates.

open as a page

Why can a detector win on mAP at IoU 0.5 but lose on mAP averaged over 0.50 to 0.95?

level: seniorimportance: should knowfreq 55%

basics

~20 s

mAP at IoU 0.5 forgives loose boxes, while averaging over the ten thresholds from 0.50 to 0.95 in 0.05 steps scores localisation quality directly. A model that classifies confidently but draws sloppy boxes tops the first and collapses on the second.

open as a page

Your detector finds buses reliably but misses 16-pixel traffic signs. Why does a single-scale head fail here?

level: seniorimportance: should knowfreq 56%

basics

~20 s

A head on the backbone's deepest feature map sees a stride near 32 pixels, so a 16-pixel sign occupies under one cell. Downsampling destroys its detail, and often no location is assigned to it. A feature pyramid fixes this.

open as a page

A segmentation model reports 95% pixel accuracy but 42 mIoU — what explains the gap?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Pixel accuracy is dominated by whichever classes own most of the pixels, while mIoU averages per-class intersection over union with every class weighted equally. Rare classes scoring near zero barely move accuracy but drag the mean down hard.

open as a page

How would you choose between a two-stage and a one-stage detector under a 30 ms per-frame budget?

level: principalimportance: should knowfreq 47%

basics

~20 s

Budget the worst case, not the average. A one-stage head costs the same on every frame; a two-stage head's second pass scales with proposal count, so busy scenes are its slowest. Decide on tail latency first, then on the accuracy you need.

open as a page

When merging instance masks and stuff maps into a panoptic output, how do you resolve contested pixels?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

A panoptic output allows one label per pixel, so conflicts follow a fixed precedence: instance masks are painted in confidence order, each keeping only unclaimed pixels, and stuff fills the remainder. Fragments left too small are dropped.

open as a page

Why does a transposed-convolution decoder produce checkerboard artifacts in its output?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

A transposed convolution writes a scaled copy of its kernel into the output at stride spacing. When the kernel size is not divisible by the stride, some output positions receive more overlapping writes than others, and that uneven overlap prints a periodic checkerboard.

open as a page

Per-class AP on a 12-instance defect class swings 8 points on one missed box - how do you report it?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

With 12 instances, recall moves in steps of about 8 points, so that class's AP is a coarse, high-variance statistic. Publish per-class AP beside its instance count with an uncertainty estimate, and never let the headline mean carry a ship decision it cannot support.

open as a page