In an object detector, what do anchor boxes do, and how does an anchor-free head replace them?
answer
- the head predicts a correction, not coordinates
- reference rectangles tiled at every location
- a few scales times a few aspect ratios
- positives chosen by an overlap threshold
- the alternative regresses four side distances
basics
~20 sAnchors are fixed reference boxes tiled over every feature-map location; the head predicts a class score and offsets that nudge an anchor onto an object. Anchor-free heads drop them and regress the four side distances straight from each location.
solid answer
~50 sA detection head is dense: at every location of a feature map it emits predictions. An anchor-based head attaches a fixed set of reference boxes there — a few scales crossed with a few aspect ratios — and for each one predicts a class or objectness score plus four numbers that turn the anchor into a box. Centre offsets are divided by the anchor's width and height and sizes are log ratios, so the network learns a small scale-free correction. Anchors also give a discrete training rule: high overlap with a ground-truth box makes an anchor positive, low overlap negative, and the middle band is ignored. An anchor-free head keeps the dense grid but drops the reference boxes: locations inside a ground-truth box are positive and each predicts its distances to the four edges. What really changes is the assignment rule and its hyperparameters, not the shape of the output.
go deeper
Be able to say that a detector predicts at every location of a feature map, and that anchors are fixed reference rectangles the network adjusts rather than boxes it invents from scratch.
Explain the mechanics: scales times aspect ratios per location, offsets normalised by anchor size, log size ratios, and the positive/ignore/negative overlap thresholds used to build the training labels. Then contrast with predicting four side distances from a location.
Show you have tuned this. Talk about clustering box shapes from the training set, recognising a dataset whose shape distribution the anchor set does not cover, and reading the symptom as low recall on a specific object class rather than a general accuracy problem.
Own the choice for a programme: anchor-free removes hyperparameters that quietly encode one dataset's statistics, which matters when many teams retrain on shifting data, but it moves the tuning burden into assignment and quality scoring. Decide which knobs your organisation can maintain.
## What a detection head has to produce An image can contain zero objects or two hundred, but a convolutional network produces a fixed-shape tensor. Detection heads resolve this by predicting densely and filtering afterwards: the head runs over a feature map and, at every spatial location, emits a fixed number of candidate boxes with scores. Most of those candidates are background and are thrown away at inference. The design question this leaf is about is how a *location* becomes a *box*. ## Anchors as priors An anchor (also called a prior or a default box) is a rectangle of predefined size and shape, notionally centred on the feature-map cell. A typical head defines, say, three scales times three aspect ratios, giving nine anchors per location; a feature map of 50x50 then carries 22,500 anchors, and a multi-level head carries far more. For each anchor the head predicts: - a score — either one objectness value plus class scores, or a score per class directly; - four regression numbers `t_x, t_y, t_w, t_h`. The regression numbers are not pixel coordinates. The standard parameterisation is ``` b_x = a_x + a_w * t_x b_w = a_w * exp(t_w) b_y = a_y + a_h * t_y b_h = a_h * exp(t_h) ``` where `a_*` is the anchor and `b_*` the predicted box. Two properties matter. Dividing the centre shift by the anchor's own size makes the target scale-invariant: moving a small box by five pixels and a large box by fifty are the same target magnitude. Predicting a *log* size ratio makes the target symmetric around zero and guarantees a positive width and height whatever the network outputs. The result is that the network never learns to produce a box from nothing — it learns a small correction to a box that is already roughly right. ## The assignment rule Anchors also solve the labelling problem. Before you can compute a loss you must decide which of the tens of thousands of predictions is responsible for which ground-truth object. The classical rule compares each anchor to each ground-truth box by overlap: - overlap above a high threshold (commonly around 0.5 to 0.7): positive, trained to classify as that object and to regress toward its box; - overlap below a low threshold (commonly 0.3 to 0.4): negative, trained toward background; - in between: ignored, contributing no loss, so borderline cases do not push the classifier in a random direction. A rescue rule is normally added: for every ground-truth box, its single best-matching anchor is forced positive even if no anchor clears the threshold. Without it, objects whose shape is unlike any anchor would receive no supervision at all. ## The cost of anchors Anchors are hyperparameters, and they are hyperparameters coupled to the dataset. Scales and aspect ratios chosen for upright pedestrians describe tall thin boxes; the same set applied to overhead imagery, where an object's axis-aligned box changes aspect ratio continuously as the object rotates, may leave whole shape ranges without a well-matching anchor. Those objects get few positives, are under-trained and are missed at inference. Practitioners respond by clustering box shapes in the training set to pick anchor sizes, by tuning the thresholds, or by adding more anchors — which raises memory and the already extreme background-to-foreground ratio without necessarily helping. ## Anchor-free heads Anchor-free heads keep the dense grid and delete the reference boxes. The two common families are: **Centre-based.** A location is positive if it falls inside a ground-truth box (often restricted to a small region around the box centre). It predicts four non-negative distances — to the left, top, right and bottom edges — so the box is recovered as `(x - l, y - t, x + r, y + b)` from the location's own coordinates. Because locations near the border of a large object produce poor boxes, such heads usually add a per-location quality score, sometimes called centre-ness, that downweights off-centre predictions when boxes are ranked. **Keypoint-based.** The head predicts heatmaps of box corners or centres and then pairs or sizes them, treating detection as keypoint localisation rather than box regression. The honest summary is that anchor-free is not assignment-free. You still choose which locations are positive, how objects of different sizes map onto pyramid levels, and how ambiguous overlapping objects are resolved. What you gain is fewer geometric hyperparameters, fewer predictions per location, and no dependence on the box-shape statistics of the training set; what you give up is the built-in prior that made the regression target small and easy, which is why anchor-free heads care more about their quality-score and assignment design. ## What an interviewer is listening for That anchors are constants, not learned weights. That the head predicts a *correction*, not coordinates. That the overlap thresholds are a labelling device, distinct from anything done at inference. And that switching to anchor-free moves the difficulty rather than removing it.
- Anchor shapes tuned for upright pedestrians are reused on overhead drone imagery where vehicles appear at arbitrary rotation. What breaks?Coverage. An axis-aligned box around a rotated vehicle changes aspect ratio with the rotation angle, so for many objects no anchor clears the positive overlap threshold. Those objects fall back on the best-match rescue rule or get no positive at all, are under-trained, and are missed at inference. Fixes: re-cluster anchor shapes on the actual data, lower the positive threshold, or move to an anchor-free head whose positives depend only on being inside the box.
- Why regress a log size ratio instead of the box width in pixels?It makes the target scale-free and the output valid by construction. A ten-percent size error costs the same whether the object is 20 or 200 pixels wide, so one head can serve both. And since width is recovered as anchor width times the exponential of the prediction, the network can output any real number and still yield a positive width — no clamping needed.
- Why are anchors with middling overlap ignored rather than labelled negative?An anchor that half-covers an object is neither a clean example of that object nor a clean example of background. Training it as negative teaches the classifier to suppress evidence that is genuinely present, which hurts recall and makes scores noisy near objects. The ignore band simply removes those ambiguous cases from the loss so the decision boundary is learned from unambiguous ones.
Anchors are like the printed guide rectangles on a form: the network only has to say which box you are in and nudge the edges a little. An anchor-free head hands you a blank sheet and a pen at each point, and asks how far each edge is.
saying these in an interview costs you the question
- Says anchors are learned parameters updated by gradient descent
- Thinks the head regresses absolute pixel coordinates per anchor
- Claims anchor-free heads need no assignment rule at all
- Assumes more anchors always improves recall for free
- Confuses anchors with the boxes surviving suppression at inference