skip to content

Why sample a learning rate log-uniformly on [1e-4, 1e-1] instead of uniformly?

level: middleimportance: should knowfreq 45%

answer

  1. the range spans decades, not units
  2. length lives in the top decade
  3. ratios matter more than differences
  4. sample the exponent, not the value
  5. each decade gets an equal share

basics

~20 s

Because the interesting values span decades, not units. A uniform draw on that interval puts about 90 percent of the trials above 0.01 and almost none near 0.0001, whereas a log-uniform draw gives each decade an equal share of the budget.

solid answer

~50 s

The interval [1e-4, 1e-1] is three decades wide, but almost all of its *length* lies in the top decade. Sampling uniformly, the share of draws below 0.01 is about (0.01 - 0.0001)/0.0999, roughly 10 percent — so nine trials in ten land in the coarse region and the two smaller decades are barely explored. A log-uniform draw is equivalent to sampling the exponent uniformly on [-4, -1] and taking ten to that power, so each of the three decades receives about a third of the trials. Use a log scale for any hyperparameter whose effect is multiplicative and whose plausible range spans orders of magnitude: learning rates, regularisation strengths, kernel widths, class-weight ratios. Keep a linear scale for bounded proportions such as the fraction of rows or features sampled per tree, and for small integer ranges like tree depth.

code

python · 15 lines
python
import random

random.seed(0)
n = 1000

uniform_draws = [random.uniform(1e-4, 1e-1) for _ in range(n)]
log_draws = [10 ** random.uniform(-4, -1) for _ in range(n)]

def below(values, cutoff):
    return sum(1 for v in values if v < cutoff)

print("uniform, below 0.01:", below(uniform_draws, 1e-2))
print("log-uniform, below 0.01:", below(log_draws, 1e-2))
print("uniform, below 0.001:", below(uniform_draws, 1e-3))
print("log-uniform, below 0.001:", below(log_draws, 1e-3))

go deeper

for a junior

Know that a search range has a shape as well as bounds, and that learning rates and penalty strengths are conventionally searched across powers of ten rather than in equal steps.

for a middle

Be able to compute the wasted share: on [1e-4, 1e-1] a uniform draw lands below 0.01 only about a tenth of the time, while a log-uniform draw gives each decade a third of the trials. Say which hyperparameters get which scale.

for a senior

Demonstrate that you read the search output, not just its best score: boundary hits mean the range was wrong, flatness across decades means that axis is not worth its budget, and rounded integer draws must be logged as trained.

for a principal

Own the convention. Decide the default ranges and scales your teams start from for each model family, so searches are comparable across projects and nobody rediscovers that a uniform learning-rate range wastes nine trials in ten.

## The question a search range answers When you hand a hyperparameter to a random search you supply two things: the bounds, and the *shape* of the distribution between them. People agonise over the bounds and default the shape to uniform, which is the wrong default for a large class of hyperparameters. ## Why uniform wastes a wide range The interval from 0.0001 to 0.1 has length about 0.0999. The sub-interval below 0.01 has length 0.0099. Under a uniform draw the probability of landing there is the length ratio, about 0.099 — roughly one trial in ten. Push the question further: the probability of a draw below 0.001 is about 0.009, so with 60 trials you expect *half a sample* in that entire decade. But the decades are not equally uninteresting. For a learning rate, 0.0005 and 0.05 are two genuinely different training regimes, and 0.09 versus 0.095 is the same regime measured twice. A uniform draw spends its budget distinguishing values that behave identically and skips the ones that behave differently. ## What log-uniform does A log-uniform (also called reciprocal or log-scale) draw on `[a, b]` is: draw `u` uniformly on `[log10(a), log10(b)]`, return `10**u`. For `[1e-4, 1e-1]` the exponent is uniform on `[-4, -1]`, so each unit of exponent — each decade — gets a third of the draws. The density in the original units is proportional to `1/x`, which is exactly the statement that *ratios*, not differences, are what matter: the chance of landing in [0.001, 0.002] equals the chance of landing in [0.01, 0.02]. That matches how these hyperparameters actually behave. Halving a learning rate has a similar qualitative effect wherever you start; subtracting 0.005 does not. ## Which hyperparameters want which scale **Log scale** — the effect is multiplicative and the plausible range spans orders of magnitude: - learning rate / shrinkage in gradient boosting - L1 or L2 penalty strength (and its inverse parameterisation, where a large value means weak regularisation) - kernel width for a radial basis function - minimum leaf weight, or a class-weight ratio for imbalanced data **Linear scale** — the parameter is a bounded proportion or a small count: - the fraction of rows or of features sampled per tree, on [0.5, 1.0] - a proportion-shaped threshold - integer axes such as tree depth from 2 to 12, or the number of neighbours in a nearest-neighbour model Integer axes need one extra decision: either draw from the discrete range directly, or draw continuously and round — but round *before* the fit, and log the rounded value, or your trial log will not match what was actually trained. ## Reading the outcome Two diagnostics fall straight out of a log-scaled range: 1. **Boundary hits.** If the best trials cluster at the smallest sampled value, the range did not go low enough — extend it by a decade rather than adding trials inside the old range. A best value pinned at either endpoint is a statement about your bounds, not about the model. 2. **Flatness.** If the score is indistinguishable across three decades, that hyperparameter is not the live one, and its budget is better spent widening another axis. ## A common trap Defining the bounds on the log scale but then sampling uniformly *between the exponents' antilogs* — that is, writing the range as [1e-4, 1e-1] and drawing uniformly — is precisely the mistake this question is about, and it is easy to make because the notation looks logarithmic. The scale lives in the sampler, not in how you wrote the numbers. The same reasoning applies to a grid: a grid over a log-spanning hyperparameter should list `[1e-4, 1e-3, 1e-2, 1e-1]`, one value per decade, not `[0.025, 0.05, 0.075, 0.1]`.

  • Which hyperparameters would you deliberately keep on a linear scale?
    Bounded proportions and small counts. The fraction of rows or features sampled per tree lives on [0.5, 1.0], where 0.6 and 0.7 differ about as much as 0.9 and 1.0, so uniform is right. Small integer ranges like tree depth 2 to 12 are the same: draw them linearly, or enumerate them.
  • The best trial sits at the smallest learning rate you sampled. What do you do?
    Treat it as a statement about the bounds, not the budget. Extend the range downward by a decade and re-run rather than adding trials inside the old range. A best value pinned to an endpoint means the optimum is probably outside what you searched.
  • How would you space a grid over a regularisation strength that could be anywhere from 0.0001 to 1000?
    One value per decade or half-decade — 1e-4, 1e-3, ... , 1e3 — giving eight points that actually differ in behaviour. Equally spaced linear points would put seven of eight above 100 and never test any meaningful amount of regularisation.

Searching a learning rate uniformly is like looking for a house number between 1 and 1000 by walking equal distances down the street when the addresses are spaced logarithmically — you cover the far end thoroughly and never see the first block.

saying these in an interview costs you the question

  • Defaults every numeric range to uniform sampling
  • Thinks writing bounds as 1e-4 makes the sampling logarithmic
  • Samples a row-subsample fraction on a log scale
  • Ignores a best value pinned at a range endpoint
  • Believes narrow ranges are always safer than wide ones

context