skip to content

Your detector finds buses reliably but misses 16-pixel traffic signs. Why does a single-scale head fail here?

level: seniorimportance: should knowfreq 56%

answer

  1. count cells per object, not pixels
  2. stride 32 versus a 16-pixel object
  3. shallow layers have resolution, not meaning
  4. top-down path plus lateral connections
  5. objects routed to levels by size

basics

~20 s

A head on the backbone's deepest feature map sees a stride near 32 pixels, so a 16-pixel sign occupies under one cell. Downsampling destroys its detail, and often no location is assigned to it. A feature pyramid fixes this.

solid answer

~50 s

Backbones downsample. By the last stage the feature map has a stride around 32, meaning one cell summarises a 32x32 patch of input. A bus spans many cells and survives; a 16-pixel sign is smaller than a single cell, so after downsampling there is almost no spatial evidence left, and under the usual assignment rules it may not claim any positive location or anchor at all. The fix is multi-scale prediction. A feature pyramid takes the backbone's stages, sends a top-down path from the deep semantic levels back down, and adds lateral connections from the matching shallow stages, so every level has both fine spatial resolution and strong semantics. Objects are then assigned to levels by size: the sign predicts on a stride-4 or stride-8 level, the bus on a coarse one. Using shallow features alone would give resolution without semantics, which is why the top-down path matters.

go deeper

for a junior

Know that a backbone shrinks the feature map as it goes deeper, and that very small objects can vanish by the last stage. Be able to say a detector predicts from several resolutions rather than only one.

for a middle

Do the arithmetic out loud: stride 32 means one cell per 32x32 patch, so a 16-pixel object is sub-cell. Explain the top-down path plus lateral connections, and why objects are routed to pyramid levels by size.

for a senior

Diagnose before you rebuild. Bin recall by ground-truth box area, confirm the failure is size-dependent, check whether small objects receive positive assignments at all, and rank the fixes by cost rather than reaching for the architecture first.

for a principal

Own the resolution budget. Decide whether the product's minimum detectable object size is met by pyramid levels, larger inputs, or tiling, and what each does to per-frame latency and hardware cost across the fleet before committing the team to a retrain.

## Where small objects disappear A convolutional backbone reduces spatial resolution in stages, typically to strides of 4, 8, 16 and 32 relative to the input. Stride is the key number: at stride 32, one feature-map cell corresponds to a 32x32 pixel patch of the original image. Now put a 16-pixel traffic sign into that. It fits entirely inside one cell, and it does not even fill it — most of the cell's evidence is background. Everything that distinguished the sign, the few pixels of shape and colour, has been averaged and strided away through four downsampling stages. A bus 300 pixels wide covers roughly nine by nine cells at the same stride and loses nothing that matters. A second, quieter failure runs alongside the first. Assignment is geometric. On an anchor-based head, the smallest anchor at stride 32 may be, say, 32 pixels on a side; a 16-pixel ground-truth box overlaps it at well under the positive threshold, so no anchor is labelled positive for it and the object contributes almost no learning signal beyond a best-match rescue. On an anchor-free head, the object may cover so few locations that it is represented by a single noisy one. Either way the model is not merely bad at small objects — it was barely trained on them. ## Why not just use the shallow layers The obvious response is to attach the head to a stride-4 feature map, where the sign covers sixteen cells. The problem is that early features are generic: edges, colours, simple textures. They have the resolution to localise but not the semantics to classify. A head there will fire on every high-contrast blob. The deep features have the opposite profile — rich semantics, no resolution. A pyramid is the construction that refuses to choose. ## The feature pyramid A feature pyramid builds one prediction level per backbone stage: 1. Take the deepest stage as the top pyramid level. 2. Upsample it by two and add a *lateral connection*: the corresponding backbone stage, passed through a small projection so the channel counts match. 3. Repeat down the pyramid. Each resulting level therefore carries semantics that flowed down from the deep stages and spatial detail that came in sideways from the shallow stage at that resolution. The detection head is then applied to every level, usually with shared weights, which keeps the parameter cost close to that of a single-scale head. Assignment becomes size-based: an object is routed to the level whose stride suits it, commonly by a rule on the square root of its box area, so small objects train and predict on the fine levels and large objects on the coarse ones. Two consequences follow. Each level sees a narrow range of object sizes, which makes the regression easier. And a head with shared weights across levels sees a roughly scale-normalised problem, which is part of why it generalises across sizes. ## The alternatives and what they cost **Raise the input resolution.** Feeding a larger image gives the sign more pixels and pushes it up into a size range the strides already handle. It is effective and it is expensive: compute and memory grow roughly with the pixel count, and everything else in the frame gets slower too. **Keep resolution in the backbone.** Removing a stride in the last stage and using dilated convolutions there preserves the feature-map resolution while keeping the receptive field large, since dilation spreads the kernel's sampling points apart without adding parameters. The cost is that the deep stage now runs at four times the spatial size, which is a substantial compute and memory increase. **Tile the image.** Splitting a large frame into overlapping crops and running the detector on each raises the effective resolution per object without changing the model, at the cost of several forward passes per frame and a stitching step that must reconcile detections across crop boundaries. **Fix the anchors or the assignment.** If the strides are fine but the smallest anchor is too large, adding smaller anchors or relaxing the positive threshold for tiny objects can recover a surprising amount, and it is nearly free. Always check this before rebuilding the architecture. ## Diagnosing it properly The symptom to establish first is that the failure is size-dependent, not class-dependent. Bin the validation set by ground-truth box area and look at recall per bin: a clean monotonic collapse below some pixel size points at resolution and assignment; a failure spread evenly across sizes for one class points at something else, such as appearance or label noise. Then check whether the small objects even receive positive assignments during training — if they do not, no architectural change downstream will help until assignment is fixed. ## What an interviewer is listening for The word *stride*, and the arithmetic that turns it into a cells-per-object number. The distinction between losing the signal and never assigning the object. The reason shallow features alone are not the answer, which is where the top-down path earns its place. And a ranking of fixes by cost, with the cheap assignment check before the expensive resolution change.

  • Why are the lateral connections necessary — why not just upsample the deepest feature map several times?
    Upsampling cannot invent detail that downsampling destroyed; you would get a blurry map with the same information content as the coarse one. The lateral connection injects the backbone's own stage at that resolution, which still holds the fine spatial structure. The addition combines semantics that flowed down with detail that came sideways, and only the combination supports both classifying and localising a small object.
  • Before rebuilding the architecture, what cheap check would you run on the small-object failure?
    Check whether small objects receive positive assignments at all. Bin the training boxes by area and count positives per bin. If the smallest bin gets almost none, the smallest anchor or the positive overlap threshold is the cause, and adding smaller anchors or relaxing that threshold costs almost nothing. Rebuilding a pyramid before ruling this out risks spending weeks on the wrong layer.
  • What does raising the input resolution buy you that a pyramid does not, and what does it cost?
    It changes the object's absolute pixel size, so a 16-pixel sign genuinely becomes a 32-pixel one with more real detail to work from; a pyramid can only make better use of the detail already captured. The cost is that compute and memory scale roughly with pixel count and every object in the frame becomes more expensive, so it is usually the second lever, not the first.

Reading a page through progressively coarser pixelation: the headline stays legible far longer than the footnote. The pyramid is like keeping both the pixelated overview and the sharp scan, and reading each item at whichever zoom shows it best.

saying these in an interview costs you the question

  • Blames the loss function rather than resolution and assignment
  • Proposes attaching the head to shallow features alone
  • Thinks upsampling a coarse map restores lost spatial detail
  • Cannot relate feature-map stride to object size in pixels
  • Ignores that small objects may get no positive assignment

context