A detector scores about 100,000 candidate boxes per image against roughly 10 objects - what breaks in training?
answer
- structural, not a dataset flaw
- many tiny terms outweigh a few large
- predicting nothing scores very well
- choose which candidates contribute
- keep the worst negatives, ignore the ambiguous
basics
~20 sWith a thousand background candidates per object, the summed training loss is dominated by easy background and the model drifts toward predicting nothing. Fixes constrain which candidates contribute: a fixed positive-to-negative sampling ratio, hard-negative mining, and an ignore band for ambiguous candidates.
solid answer
~50 sCandidates are labelled by IoU against ground truth, so with around 100,000 candidate boxes and about 10 objects, roughly a thousand negatives exist per positive - and nearly all of them are trivial, covering empty road or sky. If every candidate contributes to the loss, the total is a sum of many tiny easy-negative terms that still outweighs the handful of positive terms, so the fastest way down is to predict background everywhere; the box branch, which only has targets on positives, barely learns. The classic sampling answers are to draw a fixed positive-to-negative ratio such as 1:3 into each training batch of candidates rather than using all of them, to mine hard negatives by ranking negatives on their current loss and keeping the worst, and to mark candidates whose overlap falls between the positive and negative cut-offs as ignored so they are not taught to be background.
go deeper
Know that a detector asks a yes-or-no question at a huge number of image locations, that almost all of them are background, and that this imbalance is the norm rather than a broken dataset.
Explain why many small easy-negative losses outweigh a few positive ones, why predicting background everywhere is a strong local optimum, and what a fixed positive-to-negative sampling ratio changes about the loss the model sees.
Demonstrate you have tuned this: mine hard negatives by current loss, set two IoU cut-offs with an ignore band between them, keep regression on positives only, and reason about which false positives actually cost average precision.
Own the tradeoff between recall and confident false alarms as a product setting, not a hyperparameter. Decide what an ignore region means in your annotation contract and what the cost of a confident miss versus a false alarm is to the business.
## Where the ratio comes from A dense detector evaluates a fixed grid of candidate boxes across the image and across scales - on the order of 100,000 per image is a realistic figure. Each candidate is labelled by its IoU with the ground-truth boxes: high overlap makes it a positive with a regression target, low overlap makes it a negative that should be classified as background. A typical image holds around ten objects, and only candidates that land on them qualify, so the label distribution is roughly 1000 negatives to 1 positive. This is a *structural* imbalance, not a dataset flaw. It is not caused by rare classes or a biased collection process; it comes from the fact that objects occupy a small fraction of the image and the detector must ask the question everywhere. ## What it does to training Three distinct failure modes, worth separating in an answer: **The loss is dominated by easy negatives.** A patch of blank sky is classified as background almost immediately, with a tiny per-candidate loss. But there are tens of thousands of them, and many small numbers sum to a large one - larger than the contribution of ten positives. The gradient the optimiser follows is therefore mostly a request to be slightly more confident about background it already gets right. **The degenerate solution is attractive.** Predicting background everywhere scores extremely well on a metric that is 99.9% background. Early training collapses to it, and escaping requires that the rare positive signal survives being averaged against the flood. **The box branch starves.** The regression loss is only defined on positives - a background candidate has no target box to regress toward. So the localisation head sees around ten supervised examples per image while the classification head sees a hundred thousand. Any scheme that further dilutes the positives dilutes box learning too. ## The sampling and mining answers **Fixed ratio sampling.** Do not train on all candidates. Build each image's training set from all (or most) positives plus enough negatives to hit a chosen ratio - 1:3 positive to negative is a long-standing default - and drop the rest. The imbalance the loss sees is then a design parameter rather than a property of the image geometry. **Hard-negative mining.** Random negatives are almost all trivial and teach little. Instead score every negative with the current model, sort by loss, and keep the worst offenders - the background patches the model is currently getting wrong, which are the ones carrying information. Note the name: hard-negative mining *keeps* the hard ones. Candidates commonly mined this way are the near-misses: a box on part of an object, or on an object of a confusable class. **An ignore band.** A candidate overlapping an object at, say, IoU 0.4 is neither a clean positive nor honestly background. Labelling it background teaches the model to suppress a nearly correct detection and injects contradictory targets, which shows up as unstable confidence. The usual treatment is two cut-offs with a dead zone between them: above the upper one is positive, below the lower one is negative, and anything in between contributes nothing to the loss. Regions annotated as crowds or as ignore-regions get the same treatment. **Keep the box loss on positives.** State explicitly that regression is computed only over matched candidates, and that the classification and regression terms are weighted so the handful of positives is not drowned. ## The metric side of the same coin The imbalance also shapes evaluation, and this is the part senior candidates get to. Average precision is computed by sweeping the confidence cut-off from high to low, so what matters is *where the false positives rank*. Thousands of extremely low-confidence false positives sitting below every true positive barely move AP - they only extend the tail of the precision-recall curve past the recall the model can actually reach. A handful of *high-confidence* false positives, by contrast, sit above real detections and depress precision at every recall level after them. So the sampling and mining choices that matter most for the metric are the ones that fix confidently wrong background, not the ones that reduce the raw count of emitted boxes. That also disposes of the naive fix: raising the confidence threshold at inference removes low-confidence noise the metric was mostly ignoring anyway, and does nothing about the training dynamics that produced the confident errors.
- Why is the box-regression loss computed only over positive candidates?A background candidate has no ground-truth box to regress toward, so there is no target to define a regression error against. Including negatives would either require an arbitrary target or force the head to fit noise, and it would swamp the ten genuine localisation examples per image with thousands of meaningless ones. Localisation is supervised only where an object actually is.
- What goes wrong if every candidate is labelled either positive or negative, with no ignore band?Candidates that overlap an object substantially but fall short of the positive cut-off get taught to report background. Those are the near-correct detections you want the model to make, so training pushes down the very predictions that would have been useful, and confidence around object boundaries becomes unstable. A dead zone between two cut-offs simply excludes them from the loss.
- Does raising the inference confidence threshold solve the imbalance?No. It suppresses low-confidence false positives, which is the part average precision was already largely insensitive to because they rank below the true detections, and it costs recall. The damaging errors are confident false positives that outrank real objects, and those come from training dynamics. Thresholding at inference cannot undo a model that learned background is always the safe answer.
saying these in an interview costs you the question
- Says train on every candidate, the model will sort it out
- Claims a higher inference confidence threshold fixes the imbalance
- Computes the box-regression loss on background candidates
- Thinks hard-negative mining means discarding hard examples
- Labels ambiguous partial-overlap candidates as clean background