How would you specify a neural architecture search for an on-device image classifier?
answer
- three things must be written down
- what it can build, how it is scored, how it explores
- the space carries your prior
- score the constraint, not a stand-in for it
- evaluation cost dominates, not strategy choice
basics
~20 sA neural architecture search needs three pieces: the space of operations and connections a candidate may use, an objective that scores it, and a strategy that explores. For an on-device target, score latency measured on that device.
solid answer
~50 sI would write down three things explicitly. **The space**: what a candidate may be built from — a repeated cell over a fixed operation menu, or a macro chain with per-stage kernel size, expansion and channel choices. The space carries most of your prior, so designing it is a modelling decision. **The objective**: score validation accuracy together with latency measured on the target phone, either as a hard budget or a weighted trade-off, because a search will exploit any proxy you hand it instead of the constraint you care about. **The strategy**: a reinforcement-learning controller, an evolutionary population, or a differentiable relaxation that makes the operation choice continuous. I would also state the evaluation shortcut and total budget up front — proxy runs or a weight-sharing supernet — since candidate evaluation, not the strategy, is what a search actually spends.
go deeper
Be ready to say what an architecture search does at all: it automates the choice of layer types, sizes and connections instead of a human picking them by hand.
Explain the three components and give one strategy from each family, including how a differentiable relaxation turns a discrete operation choice into something a gradient can optimize.
Demonstrate the deployment judgement: score the real constraint measured on the real device, budget the evaluation rather than the strategy, and verify the cheap ranking by retraining the finalists.
Own the framing that the space is a modelling commitment. Be ready to argue how much of a reported search win belongs to the space, and what evidence would separate the two.
## Why a specification, not an algorithm Asking "which architecture search should I use" is the wrong first question. A search is defined by three choices, and two of them are yours regardless of which published strategy you pick. ### 1. The search space The space is the set of architectures the search can ever return. Two common shapes: - **Cell-based**: search a small directed graph — a *cell* — over a fixed menu of operations, then stack copies of the cell in a hand-designed macro skeleton (stem, stages, downsampling points, classifier head). Cheap, transferable, but the skeleton is hand-designed and does a lot of the work. - **Macro / chain-structured**: fix the block *type* and search per-stage hyperparameters — kernel size, expansion ratio, channel count, number of repeats, where to downsample. This is the shape most on-device searches take, because those are the knobs that actually move latency. The critical property: **the space encodes your prior**. A space in which every reachable point is already a decent model makes the search look brilliant; a space with no good points makes the best strategy useless. When you report a result, the space is part of the result. ### 2. The objective For a deployment-constrained model this is multi-objective. You have accuracy on a validation split, and you have a cost your deployment constraint actually names. There are three standard ways to combine them: - **Hard constraint**: maximize accuracy subject to cost <= budget; candidates over budget are rejected or heavily penalized. Simple and honest when the budget is genuinely hard (a frame deadline). - **Scalarized trade-off**: a weighted product such as accuracy multiplied by `(latency / target)^w` with a negative exponent, so exceeding the target is penalized smoothly. Tune `w` to control how badly you want to sit near the target. - **Pareto front**: keep the non-dominated set and choose afterwards. Most useful when you will ship several models to several device tiers. The operational point is that a search optimizes *exactly* what you write down, with none of a human's tacit judgement. If you write down a proxy — a count of operations — you will get an architecture that is excellent at that proxy, which is not what the product needs. Hardware-aware searches therefore put **latency measured on the target device** into the reward (measuring on the phone, or via a lookup table of per-operator measurements built from that phone), rather than a device-independent stand-in. MnasNet is the canonical example of exactly this multi-objective, on-device-measured formulation. ### 3. The search strategy Three families dominate: - **Reinforcement learning**: a controller emits an architecture description, the child is evaluated, and the resulting score is the reward used to update the controller. Conceptually clean, sample-hungry. - **Evolutionary**: maintain a population, mutate and recombine architectures, discard the worst (or, in regularized evolution, the oldest). Robust, easy to parallelize, easy to add constraints to. - **Differentiable relaxation**: replace the discrete choice among candidate operations with a continuous mixture weighted by learned coefficients, optimize architecture coefficients and weights jointly, then discretize by taking the strongest operation per edge. DARTS is the reference method. Enormously cheaper; the discretization step is the known weak point, since the mixture that trained well is not always the argmax that gets shipped. ### The fourth thing everyone forgets: how a candidate gets scored Strategy choice is nearly free; **candidate evaluation is the whole bill**. Training every candidate to convergence is what made early searches cost thousands of accelerator-days. The practical options are proxy evaluation (fewer epochs, a subset of the data, a smaller image size, a shrunken model) and **weight sharing**: train one over-parameterized supernet containing every candidate as a subnetwork, then score candidates by inheriting supernet weights. Both trade fidelity for speed, and both should be checked: does the cheap score actually rank candidates the same way full training would? ### Putting it together for an on-device classifier A defensible specification: a macro space over per-stage kernel size, expansion and repeats within a fixed block family; an objective of validation accuracy under a hard latency ceiling from the product's frame budget, measured on the target device via a per-operator latency table; a weight-sharing supernet trained once, searched with evolution under the latency constraint; then the top few candidates retrained from scratch with the full recipe and re-measured on the device before anyone believes the ranking. State the total budget before starting, and state what you will compare against — otherwise you cannot tell whether the search or the space produced the win.
- Why does a hardware-aware search put measured device latency in the reward instead of an operation count?Because a search optimizes precisely what it is scored on. If the reward is a stand-in for the real constraint, the winner will be the architecture that best exploits the gap between the stand-in and the deployment reality. Measuring on the target device, or using a per-operator table built from it, closes that gap so the search is optimizing the thing the product actually promises.
- How do you turn accuracy and latency into a single number a search can maximize?Three usual forms: a hard constraint that rejects candidates over budget, a scalarized reward such as accuracy multiplied by a penalizing power of the latency-to-target ratio, or keeping the Pareto front and choosing later. Hard constraints suit a real frame deadline, scalarization suits soft budgets, and a Pareto front suits shipping several device tiers from one search.
- What is the main risk of a hand-crafted cell search space?It can encode the answer. If the operation menu and the surrounding skeleton were chosen because they work, every sampled candidate is already good and the strategy gets credit for the space's quality. That also limits transfer: a space tuned around one block family and one device may contain nothing good for a different accelerator or task.
- Why retrain the winning architecture from scratch instead of shipping the supernet subnetwork?Inherited supernet weights were trained under interference from every other subnetwork sharing them, so they usually understate what the architecture can reach and may rank candidates differently from independent training. Retraining the top few with the full recipe, then re-measuring latency on the device, is how you check that the cheap ranking survived contact with real training.
saying these in an interview costs you the question
- Names a search strategy but never defines the space
- Optimizes an operation count when the constraint is device latency
- Assumes a cell tuned for one device is hardware independent
- Ignores that evaluating each candidate is the real cost
- Believes supernet-inherited scores equal fully trained accuracy