skip to content

When should you use RandomizedSearchCV instead of GridSearchCV in scikit-learn?

level: middleimportance: must knowfreq 74%

answer

  1. cost is a product versus a chosen budget
  2. n_iter caps the number of fits
  3. param_distributions accepts rvs objects
  4. log-uniform for scale parameters
  5. both refit the winner by default

basics

~20 s

GridSearchCV fits every combination in param_grid, so its cost multiplies with each added parameter. RandomizedSearchCV draws a fixed n_iter samples from param_distributions, letting you cap the budget and sample continuous ranges instead of a hand-picked ladder.

solid answer

~50 s

`GridSearchCV(estimator, param_grid, cv=...)` evaluates the full Cartesian product: three parameters with five values each is 125 candidates times the number of folds. `RandomizedSearchCV(estimator, param_distributions, n_iter=..., random_state=...)` instead samples `n_iter` settings (default 10) — so the cost is something you choose rather than something the grid dictates. The second difference matters as much: `param_distributions` accepts either a list, which is sampled uniformly, or any object with an `rvs` method, such as `scipy.stats.loguniform(1e-4, 1e2)` or `randint(2, 50)`. That lets you sample a continuous range rather than committing to a ladder of values you guessed. When only two or three parameters actually influence the score, random sampling explores each one at many distinct values for the same budget, which is why it usually beats a grid of equal cost. Both expose the same API — `best_params_`, `best_score_`, `cv_results_` — and both refit the winner on the full data by default, so the search object itself is a usable estimator.

code

python · 24 lines
python
from scipy.stats import loguniform, randint
from sklearn.datasets import load_digits
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import RandomizedSearchCV, StratifiedKFold

X, y = load_digits(return_X_y=True)

search = RandomizedSearchCV(
    RandomForestClassifier(random_state=0),
    param_distributions={
        "n_estimators": randint(50, 400),
        "max_depth": randint(2, 20),
        "max_features": loguniform(0.05, 1.0),
        "criterion": ["gini", "entropy"],
    },
    n_iter=25,
    cv=StratifiedKFold(5, shuffle=True, random_state=0),
    scoring="accuracy",
    random_state=0,
    n_jobs=-1,
)
search.fit(X, y)
print(search.best_params_, search.best_score_)
print(search.predict(X[:5]))  # works because refit=True

go deeper

for a junior

Know that GridSearchCV tries every combination in param_grid while RandomizedSearchCV tries n_iter sampled ones, and that both leave you best_params_ and a refitted best_estimator_.

for a middle

Explain the cost arithmetic and that param_distributions accepts scipy distributions, so continuous parameters get sampled rather than laddered. Name log-uniform sampling for scale parameters.

for a senior

Show how you budget a search on real hardware — n_jobs and nested-parallelism, error_score while debugging, reading cv_results_ for the score spread rather than trusting the single winner.

for a principal

Frame tuning as a cost decision: how much compute a percentage point is worth, when random-then-local-grid beats an exhaustive sweep, and how search results are recorded so the team is not re-running the same space.

## The two searchers Both classes are meta-estimators: you hand them an unfitted estimator plus a description of the parameter space, they cross-validate every candidate they generate, and they keep the best. `GridSearchCV(estimator, param_grid, *, scoring=None, n_jobs=None, refit=True, cv=None, verbose=0, error_score=nan, return_train_score=False)` takes `param_grid` as a dict of lists — or a *list of dicts*, which is how you express mutually exclusive regions (one dict for a linear kernel, another for an RBF kernel with its own `gamma` values). It fits every combination. `RandomizedSearchCV(estimator, param_distributions, *, n_iter=10, ...)` takes the same keywords plus `n_iter` and `random_state`. Each of the `n_iter` candidates is drawn independently. ## Cost A grid's cost is the product of the list lengths, multiplied by folds, multiplied by fit time. Adding one more parameter with four values quadruples the run. This is the practical reason grid search collapses beyond three or four parameters — and the reason people shrink their grids until the search stops being informative. Random search decouples the budget from the dimensionality: `n_iter=60` is 60 fits per fold whether the space has two parameters or twelve. You spend what you have. There is a well-known geometric argument too. In a grid, a parameter with five listed values is only ever tried at those five points, no matter how many total fits you run — the other parameters vary while it repeats. With random sampling, `n_iter` draws give `n_iter` distinct values of *every* continuous parameter. When only a couple of parameters really matter — the common case — random search resolves those far more finely for the same money. ## Distributions, not just lists `param_distributions` values may be: - a list or array, sampled uniformly with replacement — use this for categorical choices like `kernel` or `class_weight`; - any frozen `scipy.stats` distribution exposing `rvs`, e.g. `loguniform(1e-5, 1e1)` for a regularization strength, `randint(2, 40)` for `max_depth`, `uniform(0.5, 0.5)` for a fraction in [0.5, 1.0]. Log-uniform sampling is the one to remember for scale parameters. A uniform draw on [1e-5, 10] puts almost every sample above 1; the log-uniform draw spreads them evenly across the orders of magnitude, which is how these parameters actually behave. ## What both give you afterwards After `fit`, both expose: - `best_params_` — the winning setting; - `best_score_` — its **mean cross-validated** score, not a hold-out score; - `best_estimator_` — the winner refitted on all the data passed to `fit`, present because `refit=True` by default; - `cv_results_` — a dict you can hand straight to `pandas.DataFrame`, with a column per parameter, per-fold scores, mean and std, ranks, and timings; - `best_index_`, `n_splits_`, `scorer_`, `refit_time_`. Because of the refit, the search object *is* an estimator: `search.predict(X_new)` and `search.score(X_test, y_test)` delegate to `best_estimator_`. Set `refit=False` when you only want the report and cannot afford the final fit — but then `predict` and `best_estimator_` are unavailable. `scoring` selects the criterion; with a list or dict of scorers you must set `refit` to the name of the one that decides the winner. Remember every scorer is maximized — regression error scorers are the negated forms (`'neg_mean_squared_error'`), so `best_score_` is legitimately negative there. ## Operational details worth naming - `n_jobs=-1` parallelizes candidate/fold combinations. Watch for nested parallelism: an estimator that already uses all cores plus `n_jobs=-1` oversubscribes the machine. - `error_score` defaults to `np.nan`, so a candidate whose fit raises becomes a NaN row instead of stopping the run. Use `error_score='raise'` while you are still finding out whether your space is even valid. - `verbose` prints progress; on a long search that is the difference between watching and guessing. - Both searchers `clone` the estimator for every candidate, so the object you passed in is never mutated. - Scikit-learn also ships halving variants, `HalvingGridSearchCV` and `HalvingRandomSearchCV`, which start many candidates on a small resource budget and promote survivors. They are still marked experimental and require `from sklearn.experimental import enable_halving_search_cv` before the import works. ## Choosing between them Use a grid when the space is small, discrete and genuinely worth exhausting — a handful of solvers, three depths — or when you must be able to say every listed combination was tried. Use random search when the space is large, when parameters are continuous, or when the budget is fixed by a deadline rather than by the space. A very common pattern is random search first to find the promising region, then a small grid around it.

  • How do you express "linear kernel with these C values, RBF kernel with these C and gamma values" in GridSearchCV?
    Pass a list of dicts as `param_grid`. Each dict is expanded into its own Cartesian product and the results are concatenated, so `[{'kernel': ['linear'], 'C': [...]}, {'kernel': ['rbf'], 'C': [...], 'gamma': [...]}]` avoids generating meaningless linear-plus-gamma combinations. `RandomizedSearchCV` accepts a list of distribution dicts the same way, sampling one dict per draw.
  • Why sample a regularization strength with loguniform rather than uniform?
    Scale parameters act multiplicatively: the interesting difference is between 0.001 and 0.01, not between 8 and 9. A uniform draw over [1e-5, 10] puts nearly every sample in the top order of magnitude and effectively never probes the small end. `loguniform` spreads draws evenly across the exponents, which matches how the parameter changes the fitted model.
  • What does best_score_ actually measure after a search finishes?
    The mean score across the cross-validation folds for the winning candidate — computed before any refit, on data the candidate's folds held out. It is not a hold-out estimate of the tuned pipeline, and because it is the maximum over many candidates it is optimistically biased. Report a score from data the search never saw instead.
  • Your search returns NaN for several candidates and no exception. Why?
    `error_score` defaults to `np.nan`, so a candidate whose fit raises — an invalid parameter combination, a solver that fails to converge into an error, a degenerate fold — is recorded as NaN with a warning rather than aborting the run. Set `error_score='raise'` to see the real traceback, then fix the space or guard the combination.

saying these in an interview costs you the question

  • Thinks RandomizedSearchCV samples a random subset of the grid points only
  • Reports best_score_ as the model's expected production performance
  • Uses a uniform range for a regularization strength spanning orders of magnitude
  • Assumes refit must be done manually after the search
  • Grows a grid by another parameter without noticing the cost multiplies

context