When merging instance masks and stuff maps into a panoptic output, how do you resolve contested pixels?
answer
- the output allows one label per pixel
- paint in some deterministic order
- countable objects outrank background regions
- what happens to a mostly-buried mask
- PQ splits into two named factors
basics
~20 sA panoptic output allows one label per pixel, so conflicts follow a fixed precedence: instance masks are painted in confidence order, each keeping only unclaimed pixels, and stuff fills the remainder. Fragments left too small are dropped.
solid answer
~50 sThe two prediction streams disagree in two ways: instance masks can overlap each other, and an instance mask can cover pixels that the stuff prediction also labels. The standard merge resolves both with a precedence rule rather than a learned arbiter. Sort the surviving instances by confidence and paint them in descending order onto an empty canvas; each instance keeps only the pixels no earlier instance took. If an instance's surviving area is a small fraction of its original mask — it was mostly buried under a more confident one — drop it entirely rather than keeping a fragment. Then fill unclaimed pixels from the stuff prediction, so things win contested pixels against stuff. Finally, stuff regions smaller than an area threshold are set to unlabelled. The result satisfies the panoptic contract: every pixel belongs to exactly one segment, and no two segments overlap.
go deeper
Recall that panoptic output permits only one label per pixel, so overlapping predictions must be resolved by a rule rather than left as they are. Knowing that countable objects take precedence over background regions is enough here.
Explain the merge order concretely: instances by descending confidence, each clipped by earlier ones, then stuff into the remainder. Be able to state what panoptic quality's two factors measure and how a segment is counted as matched.
Show you can read the two factors as a diagnosis and act on it: which failures move recognition quality, which move segmentation quality, and how the overlap and area thresholds in the merge trade dropped instances against kept fragments.
Own the question of whether a single scalar is the right target at all. The merge encodes a cost asymmetry between missing an object and mislabelling background pixels; decide whether that asymmetry matches the product's real costs, and say so before the team optimises the number.
## Why merging is needed at all Panoptic segmentation demands a single assignment from pixels to segments: every pixel gets exactly one class, and one instance id if that class is countable. Neither of the streams that feed it satisfies that on its own. - The **instance stream** is a list of detections, each predicted from its own region independently. Two detections can and do claim the same pixel — a person standing in front of another produces two masks that overlap around the boundary. - The **stuff stream** labels the image densely, and it does not know which pixels a detector will claim; it will happily call the pixels of a car `road` if the car sits on the road. So a deterministic merge step is part of the task, not an implementation detail. ## The precedence rules **Things beat stuff.** Contested pixels go to the instance. The asymmetry is justified by what the errors cost: a countable object silently absorbed into a stuff region disappears as an object entirely, whereas a stuff region losing a few pixels to a thing changes only its area. **Among things, confidence decides.** Instances are sorted by their detection score and painted in descending order. Each one writes only where the canvas is still empty, so a lower-scoring mask is clipped by every higher-scoring one it overlaps. This is a greedy rule, not an optimisation — it is cheap, deterministic and matches the intuition that the model's more confident claim should win. **Fragments are dropped.** After clipping, an instance whose surviving pixels are only a small fraction of the mask it originally predicted is discarded rather than kept as a sliver. Keeping such a fragment costs twice: it is almost certainly a duplicate detection of the object in front, and as a separate segment it will register as a spurious prediction under the metric. **Small stuff becomes void.** Stuff regions below an area threshold are marked unlabelled rather than committed to, since tiny stuff fragments are usually prediction noise. Panoptic evaluation ignores designated void regions rather than penalising them. ## Scoring the result: panoptic quality Panoptic quality (PQ) compares the merged prediction against the ground truth with one number: `PQ = (sum of IoU over matched pairs) / (|TP| + 0.5*|FP| + 0.5*|FN|)` where a predicted segment and a ground-truth segment of the same class are **matched** when their intersection over union exceeds 0.5. `TP` are the matched pairs, `FP` are predicted segments left unmatched, and `FN` are ground-truth segments left unmatched. Because the merged output is non-overlapping, the 0.5 threshold makes the matching **unique** — at most one prediction can overlap a given ground-truth segment by more than half of their union — so no greedy tie-breaking by score is needed at evaluation time. PQ factors exactly into two interpretable terms: `PQ = SQ * RQ`, where `SQ = (sum of IoU over matched pairs) / |TP|` and `RQ = |TP| / (|TP| + 0.5*|FP| + 0.5*|FN|)` - **SQ, segmentation quality**, is the mean overlap of the segments you *did* match. It punishes sloppy boundaries on objects you found correctly. - **RQ, recognition quality**, is an F1-style score over segments. It punishes objects you missed and segments you invented, counting each unmatched segment at half weight in the denominator, and it treats every segment equally regardless of how many pixels it covers. The split is what makes PQ diagnostic. A model that finds everything but traces boundaries loosely has high RQ and low SQ. A model that draws beautiful masks but misses half the objects, or emits duplicates, has high SQ and low RQ. There is a cliff between the two: boundaries that degrade past IoU 0.5 stop being a match at all, so the failure migrates from SQ into RQ and the score drops faster than the boundary error alone suggests. ## What the merge does to the metric The rules above are chosen with PQ in mind. Dropping clipped fragments protects RQ, because a fragment is a segment that will be unmatched and counted as a spurious prediction. Letting things win contested pixels protects both terms, since a swallowed object is a missed segment. Setting tiny stuff regions to void keeps guesses out of the denominator entirely. Tuning the overlap and area thresholds is a genuine trade: raise them and you drop marginal instances, losing matches; lower them and you keep slivers that count against you. ## Frequent errors Resolving overlaps by mask area rather than confidence, keeping a pixel in two segments "because both models were confident", describing PQ as an average IoU over all segments (it is not — the denominator counts unmatched segments), or claiming that a segment below the 0.5 overlap still counts as found.
- Why does an overlap threshold above 0.5 make panoptic matching unique?Because the merged prediction and the ground truth are both non-overlapping labellings, at most one predicted segment can exceed half the union with a given ground-truth segment — two such predictions would have to share more than half of it between them, which they cannot do without overlapping. So the matching is one-to-one by construction, with no score-based tie-breaking.
- A model finds every object but its boundaries are consistently loose. Which PQ term falls?Segmentation quality, the mean overlap over matched pairs, falls while recognition quality stays near its ceiling, since every ground-truth segment still gets a match. The caveat is the cliff: once a boundary degrades past the 0.5 overlap the pair stops matching, and the same failure starts hitting recognition quality as a missed segment plus a spurious one.
- Why drop an instance whose mask is mostly clipped away by more confident ones?A heavily buried mask is usually a duplicate detection of the object in front rather than a genuinely occluded second object. Kept as a segment it will not match anything in the ground truth and will count as a spurious prediction, dragging recognition quality down; discarding it costs a match only in the rarer case where the occluded object was real.
saying these in an interview costs you the question
- Leaves a contested pixel in two segments at once
- Resolves instance overlaps by mask area rather than confidence
- Lets a stuff label override an instance mask by default
- Calls PQ an average IoU over all segments
- Counts a pair below 0.5 overlap as a match
- Keeps heavily clipped mask fragments as separate segments