When is running an architecture search worth its compute versus scaling a known design?
answer
- an amortization question, not a technical one
- divide the cost by models shipped
- the space may be doing the work
- random search in the same space is the control
- device-specific unless re-measured
basics
~20 sAn architecture search pays when its cost amortizes: unusual target hardware, or one space reused across many deployment targets. For a single model on ordinary hardware, scaling a well-tuned published design is cheaper and often as good.
solid answer
~50 sI treat it as an amortization question. Early controller-driven searches cost thousands of accelerator-days for one dataset and one device, which rarely repays itself for a single model, and the honest baseline is brutal: random sampling within the same space, at the same budget and training recipe, has repeatedly come close to the searched result — evidence that much of the gain came from the hand-designed space rather than from searching it. A search earns its place when the cost spreads: a supernet trained once and used for several latency tiers, or unusual accelerator hardware no published family targets. Otherwise I take a published design, scale it to the compute budget, and spend the compute on data and the training recipe. If we do search, I insist on the random-search control and on retraining finalists from scratch.
go deeper
Be ready to say that automated architecture search is expensive and that starting from a published, well-tested design is the normal choice for a first model.
Explain what makes a search costly — thousands of candidates each needing training — and how weight sharing collapses that into a single supernet training run.
Demonstrate that you would run the control: random search in the same space, equal training recipes, finalists retrained from scratch and re-measured on the real device before anything is believed.
Own the amortization argument. Justify the compute by the number of deployment targets it serves and the strangeness of the hardware, and be willing to say no and scale a published design instead.
## The cost side The first generation of architecture searches trained thousands of candidate networks essentially from scratch. The resulting bill — thousands of accelerator-days for one search, on one dataset, targeting one device — is the number that made the field uncomfortable, and it prompted two responses. **Cheaper evaluation.** Weight sharing is the main lever. Train one over-parameterized *supernet* whose subnetworks are the candidates, then score a candidate by inheriting the relevant supernet weights instead of training it. One training run replaces thousands. Differentiable relaxations go further by folding the architecture choice into the same gradient loop as the weights. Costs fall from accelerator-*centuries* to something a team can actually run. **Amortization by construction.** A once-for-all style supernet is trained so that many of its subnetworks are individually deployable without retraining. One expensive training run then serves a whole ladder of latency targets — phone tiers, an embedded board, a server-side variant. Here the search cost is divided by the number of models shipped, which is the only arithmetic that reliably makes it worth it. ## The baseline that decides the argument The question a lead must ask is not "did the search find a good model" but "did the *searching* add anything over the space it searched". Several careful re-evaluations found that **random sampling inside the same search space, given a comparable budget, produces architectures close to the searched ones**, and that reported gains often shrank once training recipes were equalized. Two mechanisms explain most of it: - **The space carries the prior.** If the operation menu, the block family and the macro skeleton were all chosen by humans who already knew what works, most points in the space are good and the strategy is picking among near-equivalent options. - **Recipe confounding.** Searched models are frequently published with a stronger training recipe — longer schedules, better augmentation and regularization — than the baseline they are compared against. Equalize the recipe and the gap narrows. So the control experiment is mandatory: random search in the same space, same budget, same training recipe. If the strategy cannot beat that, you bought a hand-designed space at the price of a search. ## When a search genuinely pays - **Unusual or new target hardware.** Published families are tuned around common accelerators. When your operator speed profile is different — an accelerator where some operations are disproportionately slow or unsupported — no published family encodes that, and a latency-aware search over your own measurements finds things a human would not guess. This is the strongest case. - **Many deployment targets from one investment.** Several device tiers, or a product line, over which one supernet amortizes. - **A space you will reuse.** If the block family and macro structure will be reused across projects, the up-front cost is shared over all of them. - **A constraint no existing family targets.** An unusual peak-memory ceiling or an accuracy target under a hard frame deadline that published sizes bracket badly. ## When it does not - **A single model on ordinary hardware.** Take a published design, apply a scaling rule to land on the compute budget, and stop. The compute is almost always better spent on data quality, the training recipe, and evaluation. - **When the team cannot run the control.** A search you cannot compare against random search in the same space produces a result nobody can defend. - **When the real problem is elsewhere.** Label noise, distribution shift and a weak evaluation set are not fixed by a better cell. ## Transfer is not free A latency-optimized architecture is optimized *for the device it was measured on*. Operator speed profiles differ between accelerators, so the winning candidate on one device can be mediocre on another; the same is true across compilers and execution backends. Treat a searched architecture as a device-specific artifact unless you have re-measured it elsewhere. ## How I would frame the decision to a team Estimate three numbers: the search cost including the supernet training, the number of models it will produce, and the plausible accuracy-or-latency delta over scaling a published design. Then require the random-search control and a from-scratch retrain of the finalists in the plan, before the compute is approved. If the delta only justifies itself when you assume the search wins, it does not justify itself.
- What exactly does a weight-sharing supernet amortize?The per-candidate training cost. Every candidate is a subnetwork of one over-parameterized network, so a single training run gives every candidate weights to be scored with, replacing thousands of independent trainings. What it does not guarantee is fidelity: inherited weights are trained under interference from all the other subnetworks, so their ranking of candidates only approximates the ranking full independent training would give.
- How would you prove a search strategy, not the space, produced the win?Run random sampling inside the identical space with the same evaluation budget, train both winners from scratch with the identical recipe, and compare under the same measurement protocol. If the searched model beats the random one by a margin larger than seed-to-seed variance, the strategy earned its cost. Anything else means you paid for a hand-designed space.
- Why can an architecture found under a latency-aware search disappoint on a different accelerator?The search optimized against one device's operator speed profile. Another accelerator may execute those operations at very different relative speeds, may fuse different patterns, or may handle some of them poorly, so the chosen shape loses its advantage. A latency-optimized architecture should be treated as a device-specific artifact and re-measured before it is reused elsewhere.
saying these in an interview costs you the question
- Assumes a searched architecture always beats a hand design
- Compares the searched model against a weaker training recipe
- Ignores that the search space encodes most of the prior
- Treats a latency-optimized architecture as hardware independent
- Counts only the strategy cost, not candidate evaluation