Would you choose Bayesian search or Hyperband for a four-hour run on 16 parallel workers?
answer
- start from the problem's properties, not the tool
- is there an honest cheap fidelity
- how many full evaluations fit in the budget
- one method is sequential by construction
- the space matters more than the optimiser
basics
~20 sPrefer Hyperband when a cheap fidelity ranks configurations like full training and workers are plentiful; prefer Bayesian search when evaluations are expensive and few. Hybrids sample configurations with a surrogate and schedule them with brackets.
solid answer
~50 sI would answer with the two questions that actually decide it. First: is there an honest cheap fidelity? If the learner trains iteratively — boosting rounds, epochs — and short runs rank candidates roughly like long ones, an early-stopping schedule turns four hours into far more effective trials, and it parallelises well because a whole rung runs at once. Second: how many full evaluations fit? If one training run eats an hour, sixteen workers give me about sixty-four evaluations, which is the regime where a surrogate's reuse of history genuinely pays. The awkwardness is that a surrogate is sequential by construction: sixteen workers asking one acquisition function at the same moment get sixteen near-identical proposals unless pending trials are imputed or the acquisition is penalised around them. So my default is asynchronous early stopping with model-based sampling of the configurations — brackets keep the workers busy, the surrogate decides what enters them — and plain Bayesian search when no honest fidelity exists.
go deeper
Know that these are two different ways to spend a tuning budget: one learns from past trials to choose the next setting, the other kills weak runs early to give survivors more time.
Explain the mechanics behind the choice: what a surrogate needs to be useful, what a cheap fidelity is, and why a rung of trials can run all at once while a surrogate proposes one point at a time.
Turn the budget into numbers — worker-hours divided by evaluation cost — and pick accordingly, naming the parallel-proposal problem and asynchronous promotion as the practical fixes you would apply.
Own the framing that space design and a stopping rule outrank the optimiser choice, and justify how much tuning machinery the team should operate given evaluation cost and headcount.
## Reframe the question "Which search method" is rarely the real decision. The real decision has four inputs: whether a cheap fidelity exists and is honest, how many full evaluations the budget buys, how wide and how structured the space is, and how much operational complexity the team should carry. Answer those and the method falls out. ## Input 1: is there an honest cheap fidelity? Early-stopping schedules exist only because you can give a configuration less of something and still learn something. Iterative learners provide this naturally — a boosted ensemble at 50 rounds tells you a lot about the same configuration at 500. Non-iterative fits do not, and the fallback fidelity of a training subsample is biased: less data systematically favours smaller, more regularised configurations, so the race may prune exactly the candidates that need the full dataset. No honest fidelity means the bandit machinery has nothing to allocate, and the budget should go to a model-based search over full evaluations. ## Input 2: how many full evaluations does the budget buy? With 16 workers and four hours you have 64 worker-hours. If one full training run takes an hour, that is about 64 full evaluations — small enough that reusing history matters, which favours a surrogate. If a full run takes four minutes, you can afford roughly a thousand, and the sequential surrogate refit plus the difficulty of keeping 16 workers usefully busy start to eat the advantage. If a full run takes six hours, you cannot complete even one per worker, and the conversation becomes about fidelity and space reduction, not about search algorithms. ## Input 3: parallelism, which is where the two methods differ most A surrogate-based search is **inherently sequential**: the acquisition function's value at a point depends on everything observed so far, and the whole point is to choose the next point given the last result. Ask it for 16 points simultaneously and, naively, you get 16 copies of the same arg-max. The standard remedies are to impute optimistic or pessimistic outcomes for in-flight trials before proposing the next one, or to locally penalise the acquisition surface around pending points. They work, but each parallel proposal is made with less information than a sequential one would have, so speedup with worker count is sublinear and flattens. Bracket-based early stopping is the opposite. Within a rung, every configuration trains independently — embarrassingly parallel. The friction is the rung boundary, a synchronisation barrier where fast workers idle until the slowest finishes before the top fraction can be selected. Asynchronous promotion removes the barrier by promoting a configuration as soon as it ranks in the top fraction of what has completed at its rung; it wastes a little accuracy in the promotion decision and recovers most of the utilisation. On 16 workers, that asymmetry is a real argument, not a footnote. ## Input 4: the shape of the space A handful of continuous knobs with smooth behaviour is where a Gaussian-process surrogate shines. A space that is mostly categorical, conditional (a knob that only exists when another takes a certain value) or high-dimensional suits a density-ratio surrogate such as a tree-Parzen estimator, or suits letting a schedule do the work. And a very wide space with a tight budget is a signal to shrink the space first — fix knobs that do not matter, and tune the three that do — which beats any algorithm choice. ## The answer most senior practitioners give Combine them. Use bracketed early stopping for the budget schedule, so the cluster stays busy and obvious losers die young, and sample each bracket's configurations from a surrogate fitted to everything observed so far rather than independently. That keeps Hyperband's parallelism and adds the information reuse that plain bracket search lacks. It is more machinery, so it needs to be justified by evaluation cost. ## What you should also say Two things separate a strategy answer from a menu recital. **Search space beats search algorithm.** Most tuning wins come from choosing the right knobs and the right ranges, not from the optimiser. A well-scoped space searched simply beats a badly scoped space searched cleverly, and it costs less to maintain. **Diminishing returns are real.** Have a stopping rule before you start: stop when the best score has not improved by more than the fold-to-fold noise for some number of trials. The four-hour budget is a ceiling, not a target, and continued tuning against a fixed validation signal buys smaller and smaller genuine gains. ## What interviewers listen for A decision framed on the properties of the problem rather than a favourite tool; the specific observation that surrogate-based search is sequential and bracket search is not; and the honesty to say that the space and the stopping rule usually matter more than which of the two you pick.
- How do you get a sequential acquisition rule to propose 16 points at once?Impute a provisional outcome for each in-flight trial before asking for the next proposal — a pessimistic or mean value is common — or penalise the acquisition surface in a neighbourhood around pending points. Both keep proposals diverse. Each parallel proposal is still made with less information than a sequential one, so scaling is sublinear.
- What is your stopping rule when the budget is a ceiling rather than a target?Stop when the best observed score has not improved by more than the fold-to-fold noise for a fixed number of consecutive trials, or when the incumbent's margin over the deployed configuration is within that noise. Spending a full budget because it exists is how teams accumulate tuning gains that do not survive contact with fresh data.
- When would you spend the four hours on neither search method?When the space itself is wrong. If the knobs on offer barely move the metric, or ranges were guessed, the highest-return use of the time is a small screening run to find which few hyperparameters matter and what ranges are plausible, then tune those properly next time. No optimiser rescues a badly scoped space.
saying these in an interview costs you the question
- Names a favourite method without asking about evaluation cost
- Assumes a surrogate search parallelises linearly across workers
- Proposes early stopping with no honest cheap fidelity available
- Ignores that a subsample fidelity biases against high-capacity settings
- Has no stopping rule and simply spends the whole budget