skip to content

How would you choose between a two-stage and a one-stage detector under a 30 ms per-frame budget?

level: principalimportance: should knowfreq 47%

answer

  1. budget the tail, not the mean
  2. stage two scales with proposal count
  3. slowest exactly on the busiest frames
  4. 30 ms is the frame's, not the head's
  5. resolution is another way to spend it

basics

~20 s

Budget the worst case, not the average. A one-stage head costs the same on every frame; a two-stage head's second pass scales with proposal count, so busy scenes are its slowest. Decide on tail latency first, then on the accuracy you need.

solid answer

~50 s

A two-stage detector runs a proposal head that emits class-agnostic candidate regions, then crops features inside each surviving proposal and runs a per-region head that classifies and refines the box. That second pass is where the accuracy comes from, since the features are already aligned to the object, and it is also where the latency risk lives: its cost grows with proposal count. On a driving camera the crowded frames are exactly the ones you cannot afford to be slow on, and a fixed-cadence pipeline is governed by p99, not the mean. A one-stage head predicts densely in a single pass, so its cost is essentially constant and easy to budget. My decision order: establish the p99 budget left after the backbone and the rest of the stack; check whether the accuracy gap survives at the ranges that matter; then ask whether the same milliseconds buy more as input resolution than as a second stage.

go deeper

for a junior

Know the shapes: a two-stage detector proposes regions and then classifies each one, a one-stage detector predicts classes and boxes directly in a single dense pass. Be able to say which is generally faster.

for a middle

Explain the mechanics of each stage — what a proposal actually is, what the per-region pass adds, and why the second stage improves localisation. Then state where each architecture's compute goes.

for a senior

Bring measurement. Decompose the frame budget, measure p99 on crowded frames on the target hardware, and identify the proposal cap as the knob that trades tail latency against recall on exactly the hardest scenes.

for a principal

Own the tradeoff end to end: what accuracy the product actually requires at the ranges that matter, whether the same milliseconds buy more as resolution or backbone capacity, what the maintenance cost of two heads is across retraining cycles, and what evidence would reverse your call.

## The two architectures **Two-stage.** Stage one is a small dense head, usually run over pyramid levels, that predicts a class-agnostic objectness score and a box refinement for each candidate location. Its output is filtered by score, reduced by suppression and truncated to a top-N list of proposals. Stage two takes each surviving proposal, extracts a fixed-size crop of the backbone features inside it, and runs a small per-region network that outputs class scores and a class-specific box refinement. **One-stage.** A single dense head over the same pyramid predicts class scores and box regression at every location directly. There is no proposal list and no per-region pass; the output goes straight to score thresholding and suppression. ## Where the accuracy difference comes from Three things favour two stages. The second pass sees features cropped from a region that is already roughly on the object, so classification is made from evidence centred on the thing being classified rather than from a location that merely happens to be inside it. The box is refined twice, so localisation errors compound less. And the proposal stage acts as a filter: the per-region head is trained on a sampled, far more balanced set of regions rather than on a dense grid where the overwhelming majority of locations are background, which makes its classification problem an ordinary one. One-stage heads have to solve that last problem inside the dense objective itself, and the design of that objective is what determines how much of the gap remains. Treat the gap as an empirical quantity on your data, not a constant: it narrows substantially for large, well-separated objects and is widest for small or crowded ones. ## Where the latency difference comes from This is the part that decides the question. A one-stage head's cost is a function of the input size and the head's width. It does not depend on image content. Its p50 and p99 are nearly the same number, which is exactly what a fixed-cadence perception pipeline wants. A two-stage head's stage-two cost is proportional to the number of proposals it keeps. That number is capped, but the cap is a tuning decision with an accuracy cost: lower it and crowded frames lose objects; keep it high and every frame pays for the worst case. Two additional content-dependent steps sit in the same path — suppression over the proposal list, and suppression over the final detections — and both slow down on busy scenes. So the pathology is specific and it is nasty: the two-stage detector is slowest precisely on the frames that matter most, a junction full of vehicles and pedestrians. A budget met at the mean is not met at all. ## Doing the budgeting honestly Thirty milliseconds is not the head's budget; it is the frame's. Subtract capture and preprocessing, the backbone forward pass, post-processing, and whatever else consumes the same accelerator — tracking, other perception tasks, planning. What is left for the head may be a small fraction of the total, and if the backbone alone eats most of it, the head choice is moot and the real decision is the backbone. Measure p99 on representative crowded frames on the actual target hardware, not the mean on a benchmark clip on a workstation. Content-dependent latency only shows up when you feed content that varies. ## The question worth asking instead Given a fixed millisecond budget, a second stage is one way to spend it, and rarely the first one to try. The alternatives, roughly in order of how often they win: - **Higher input resolution with a one-stage head.** If the accuracy gap on your data is concentrated in small objects — and it usually is — more pixels attack the cause directly and keep the constant-cost property. - **A larger backbone.** Detection accuracy is often backbone-limited before it is head-limited, and backbone cost is content-independent. - **Better assignment and post-processing tuning.** Free, and frequently worth more than either. - **A second stage.** Correct when localisation precision is genuinely the binding constraint, the scene density is bounded, and you can tolerate content-dependent latency. ## The organisational side A two-stage pipeline has more moving parts to maintain: a proposal count, two suppression stages, two sets of assignment rules, two heads to retrain when the data shifts. If several teams retrain on evolving data, the simpler head has a real, recurring cost advantage that does not appear in any accuracy table. Conversely, if you already operate a two-stage system well and its p99 fits, switching for elegance is not a decision, it is churn. ## What an interviewer is listening for At principal level nobody is looking for "two-stage is more accurate, one-stage is faster". They want the budget decomposed, the tail-versus-mean distinction named, the content-dependence identified as the real risk, alternative uses of the same milliseconds considered, and an explicit statement of what evidence would change the decision.

  • You cap proposals to bound the two-stage latency. What does that cost you?
    Recall on exactly the frames the cap binds. The cap is applied after score-ranking the proposals, so on a crowded junction the objects dropped are the lower-scoring ones — typically small, distant or occluded, which are often the ones that matter. You have converted a latency problem into a scene-dependent accuracy problem, and it will not show up in an aggregate number that averages over mostly empty frames.
  • What evidence would make you accept the two-stage detector despite its variable latency?
    Measured p99 on the target hardware over the busiest real frames, sitting comfortably inside the budget left after the backbone and the rest of the stack — plus an accuracy gap that survives on the object sizes and ranges the product actually depends on, and that a resolution increase on the one-stage head does not close for the same milliseconds.
  • Why is a benchmark's mean latency a poor basis for this decision?
    Because the two-stage cost is content-dependent and the mean hides that. A clip of mostly sparse frames produces a mean close to the empty-frame cost, while the frames that determine whether the system misses its cadence are the crowded ones in the tail. A fixed-rate pipeline is governed by the worst frame it must still deliver, so p99 on representative busy content is the only number that answers the question.

saying these in an interview costs you the question

  • Answers only two-stage is accurate, one-stage is fast
  • Budgets the mean latency instead of the tail
  • Misses that stage-two cost depends on scene content
  • Treats 30 ms as available entirely to the detection head
  • Never considers spending the budget on resolution instead

context