Why does RoIAlign produce better instance masks than RoI pooling on a detection head?
answer
- the box lands between feature cells
- count the rounding steps, not one
- half a cell times the stride
- interpolate instead of snapping
- small objects pay the most
basics
~20 sRoI pooling rounds twice, the box onto the feature grid and then the bin edges, so features come from up to half a cell off. RoIAlign keeps float coordinates and samples bilinearly, preserving the alignment masks need.
solid answer
~50 sA proposal box lives in input-pixel coordinates but has to be read off a feature map that has been downsampled, so its corners land on fractional feature-grid positions. RoI pooling rounds twice: first it snaps the box to whole feature cells, then it splits the box into a fixed grid of bins and rounds each bin boundary as well. Both roundings shift the sampled region by up to half a feature cell, which at stride 16 is around eight input pixels. Classification barely notices — pooling a whole region to a class score is fairly translation-tolerant — but a mask is a per-pixel spatial output, so every predicted boundary inherits the shift. RoIAlign removes both roundings: bin centres stay fractional and each sample point is computed by bilinear interpolation from the four neighbouring feature cells, then the samples in a bin are aggregated. The gain is largest on small objects, whose whole extent is only a few feature cells wide.
code
python · 19 linesfeat = [[0.0, 0.0, 0.0, 0.0],
[0.0, 1.0, 5.0, 0.0],
[0.0, 3.0, 9.0, 0.0],
[0.0, 0.0, 0.0, 0.0]]
def bilinear(f, y, x):
y0, x0 = int(y), int(x)
dy, dx = y - y0, x - x0
return (f[y0][x0] * (1 - dy) * (1 - dx)
+ f[y0][x0 + 1] * (1 - dy) * dx
+ f[y0 + 1][x0] * dy * (1 - dx)
+ f[y0 + 1][x0 + 1] * dy * dx)
y, x = 1.7, 2.3 # where the float box actually falls on the feature grid
stride = 16 # input pixels covered by one feature cell
print(round(bilinear(feat, y, x), 2)) # 5.46 sampled at the true location
print(feat[round(y)][round(x)]) # 9.0 snapped to the nearest cell
print(round(abs(round(y) - y) * stride, 1)) # 4.8 input pixels of shift, one axisgo deeper
Recall that a proposal box rarely lines up with the downsampled feature grid and that RoIAlign interpolates instead of rounding. Knowing that this matters for masks more than for class scores is enough here.
Name both quantisation steps in RoI pooling, convert a half-cell error into input pixels using the stride, and explain why a per-pixel output inherits a shift that a whole-region class decision shrugs off.
Diagnose from symptoms: boxes fine, masks systematically displaced, worst on small objects and on coarse pyramid levels. Be ready to say what else you would rule out first, such as mask resolution or a wrong stride assumption in the region-to-feature mapping.
Treat exact coordinate handling as a property the whole pipeline must agree on. Any place that maps between image and feature coordinates — proposal assignment, pyramid level selection, mask paste-back — can reintroduce the same half-pixel class of bug, so decide where that convention is defined once.
## The coordinate problem A detector proposes boxes in input-image coordinates. The features those boxes must be read from have been downsampled by the backbone — a stride of 16 means one feature cell covers a 16x16 input patch. Dividing a box's coordinates by the stride almost always lands on fractional feature positions: a box at `x = 100` with stride 16 sits at feature coordinate `6.25`. Both pooling schemes must turn that fractional region into a fixed-size grid of features, because the downstream head expects a fixed input shape. They differ in how honest they are about the fraction. ## What RoI pooling rounds RoI pooling quantises **twice**: 1. **Proposal to grid.** The box's fractional feature coordinates are rounded to integer cell indices, so the region actually pooled is not the region proposed. Error: up to half a cell per edge. 2. **Grid to bins.** The snapped region is divided into the output grid — say 7x7 — and each bin boundary is rounded to integers too, so bins end up with unequal sizes and, again, cover slightly the wrong feature cells. Error: up to another half cell. Each half-cell of error is `stride/2` input pixels: eight pixels at stride 16, and more on the coarser levels of a feature pyramid. The features handed to the head therefore describe a region offset from the box the head believes it is looking at. ## Why classification survives it and masks do not Classification asks one question about the whole region: which class. That question is largely invariant to a small translation — a cat shifted eight pixels is still a cat, and max-pooling over a bin discards much of the spatial detail anyway. The argmax rarely flips. A mask asks a question **at every pixel**: is this one inside the object. The predicted low-resolution mask is placed back onto the box, so a systematic misalignment between features and box coordinates translates into a systematic displacement of the whole predicted silhouette. Nothing downstream can undo it, because the head was never told the features came from somewhere else. The damage scales inversely with object size. An object 300 pixels across shifted by eight pixels loses a sliver of its border. An object 40 pixels across shifted by eight pixels loses a fifth of its extent, and the per-bin evidence for its boundary was thin to begin with. This is why alignment shows up first and worst in small-object mask scores. ## What RoIAlign does instead RoIAlign performs **no rounding at all** on the coordinates: - The box keeps its fractional feature coordinates. - The region is divided into the output grid with fractional bin boundaries, so every bin is exactly the same size. - Inside each bin, a small fixed number of sample points (four is a common choice, placed at regular positions in the bin) is chosen, again at fractional coordinates. - Each sample's value is computed by **bilinear interpolation** from the four feature cells surrounding it, weighting each by the overlapping area implied by the fractional offsets. - The samples in a bin are aggregated — averaged or maxed — to give that bin's output. Bilinear interpolation is differentiable with respect to the feature values, so gradients flow back to the backbone exactly as they did through pooling; nothing about training changes. The exact number of sample points per bin turns out to be relatively unimportant — what matters is that the coordinates are never snapped. ## Reading the symptom in practice The fingerprint of a misalignment problem is a **gap between box quality and mask quality**: boxes match ground truth well, masks do not, and the mask errors look like a consistent shift or a border that is systematically off rather than random noise. Note that mask overlap is a much harsher test than box overlap for slender shapes — a perfectly placed rectangle around a bicycle encloses mostly background, so the same object can score high on box overlap and low on mask overlap even before alignment enters the picture. That makes mask-side metrics the sensitive place to look when you are hunting for feature misalignment. ## Common misconceptions - *"RoIAlign is a bigger pooling window."* The window is the same; only the sampling is exact. - *"RoI pooling rounds once."* Two roundings, and the second is easy to forget. - *"It matters most for large objects because they cover more bins."* The opposite: a fixed pixel offset is a larger fraction of a small object. - *"Interpolation blocks gradients."* Bilinear sampling is a differentiable weighted sum.
- Why does box classification tolerate rounding that ruins masks?Classification produces one decision for the whole region and is fairly translation-tolerant, so a half-cell shift rarely changes which class wins. A mask is a spatial output evaluated per pixel, so the same shift displaces every predicted boundary and the error is preserved all the way to the output rather than being averaged away.
- Mask overlap with ground truth is far below box overlap for a predicted bicycle. What does that tell you?Not necessarily that anything is broken. A tight rectangle around a bicycle encloses mostly background, so even a perfect mask covers a small fraction of the box and mask overlap is a much stricter measure for slender objects. Look at whether the mask error is a consistent displacement, which points at feature misalignment, or a diffuse boundary, which points at mask resolution.
- Does removing the rounding change how gradients reach the backbone?No. Bilinear interpolation is a weighted sum of four feature values with weights fixed by the fractional coordinates, so it is differentiable and the gradient is simply distributed back to those four cells in the same proportions. Training proceeds unchanged; only the forward sampling is more accurate.
RoI pooling is reading a ruler to the nearest centimetre while being asked to cut to the millimetre; RoIAlign reads the fraction between the marks instead of snapping to one.
saying these in an interview costs you the question
- Describes RoIAlign as just a larger pooling window
- Claims RoI pooling quantises only once
- Says the misalignment hurts large objects most
- Thinks bilinear sampling blocks gradient flow
- Ignores the stride when estimating the pixel-level error