Why can two RandomForestClassifier fits on the same data give different predictions?
answer
- Randomness is the mechanism, not a defect
- Two draws per forest, one parameter
- Rows resampled, features subsampled
- Parallelism is not the culprit
- random_state as an integer pins both
basics
~20 sA random forest is randomized twice: each tree trains on a bootstrap resample of the rows, and each split considers a random subset of features. With random_state left at None those draws differ per fit. Pass an integer random_state for reproducible results.
solid answer
~40 sTwo sources of randomness are baked into the algorithm, and `RandomForestClassifier` exposes both through one parameter. Each tree is fitted on a bootstrap resample of the training rows (`bootstrap=True` by default), and at every split the tree considers only a random subset of features — `max_features='sqrt'` for the classifier, `1.0` for `RandomForestRegressor`. With `random_state=None`, both draws come from NumPy's global random state, so consecutive fits differ. Passing `random_state=0` (any fixed integer) makes the whole fit deterministic. Note what does *not* affect the outcome: `n_jobs` only distributes tree-building across cores, so `n_jobs=-1` and `n_jobs=1` with the same seed produce identical models. And a seed pinned on the estimator says nothing about how the data was split upstream — that needs its own seed.
code
python · 13 linesimport numpy as np
from sklearn.datasets import make_classification
from sklearn.ensemble import RandomForestClassifier
X, y = make_classification(n_samples=500, n_features=20, random_state=0)
a = RandomForestClassifier().fit(X, y).predict(X)
b = RandomForestClassifier().fit(X, y).predict(X)
print("unseeded identical:", bool(np.array_equal(a, b)))
c = RandomForestClassifier(random_state=42).fit(X, y).predict(X)
d = RandomForestClassifier(random_state=42, n_jobs=-1).fit(X, y).predict(X)
print("seeded, different n_jobs, identical:", bool(np.array_equal(c, d)))go deeper
Say plainly that forests are randomized — bootstrap rows plus a random feature subset per split — and that passing an integer random_state makes a fit reproducible. That answer is complete at this level.
Name both randomization sources and what each is for, note that max_features defaults differ between the classifier and the regressor, and explain why n_jobs does not affect the fitted model.
Talk about reproducibility end to end: every seed in the pipeline, not just the estimator's, and using seed-to-seed variance to judge whether a metric improvement is real rather than freezing it away.
Own reproducibility as a policy — seeds recorded alongside data and library versions in the experiment record, so a result can be re-derived months later, and a stated convention for reporting variance rather than single-run numbers.
## Where the randomness comes from A random forest is not a deterministic algorithm that happens to have a seed bolted on; randomization is the mechanism that makes it work. Averaging many trees only reduces variance if the trees are decorrelated, and scikit-learn decorrelates them two ways: 1. **Bootstrap resampling of rows.** With `bootstrap=True` (the default), each tree is fitted on a sample of the training rows drawn with replacement. Different trees see different data, and roughly a third of the rows are out-of-bag for any given tree — which is what makes the optional `oob_score=True` estimate possible. 2. **Feature subsampling at each split.** At every node, the tree considers only a random subset of the features when searching for the best split. `max_features` controls the size of that subset: `'sqrt'` for `RandomForestClassifier`, `1.0` (all features) for `RandomForestRegressor`. This is what prevents one dominant feature from appearing at the top of every tree. `ExtraTreesClassifier` adds a third source — split thresholds themselves are drawn at random rather than optimized — which is exactly why it is a different estimator. ## What random_state controls The estimator's `random_state` seeds both draws. The documented behaviour: it controls the randomness of the bootstrapping of the samples used when building trees (if `bootstrap=True`) and the sampling of the features to consider when looking for the best split at each node. Three spellings are accepted. An **integer** gives a fixed, reproducible fit. A **`numpy.random.RandomState` instance** shares state, so repeated fits with the same object advance the stream and produce different models — occasionally what you want, usually a surprise. **`None`**, the default, draws from NumPy's global random state, which is why `np.random.seed(0)` at the top of a script appears to make things reproducible; it works, but it is action-at-a-distance, and passing the seed explicitly to the estimator is the maintainable form. ## What does not cause the difference - **`n_jobs`.** It only parallelizes tree construction across processes or threads. Each tree's random draws are seeded deterministically from `random_state`, so `n_jobs=1` and `n_jobs=-1` with the same seed give bit-identical predictions. Someone who blames parallelism for nondeterminism here has the wrong model of how the seeding works. - **Row order.** Reordering the input rows does not change the fitted forest for a fixed seed in the way shuffling does — but do not lean on this; if you want determinism, seed it. ## The seeds you also have to pin A reproducible experiment usually needs more than one. The data split has its own randomness, cross-validation splitters have theirs, and any resampling or augmentation step has its own too. A pinned estimator seed with an unpinned upstream split gives you a model that is reproducible on data that is not — which produces the confusing situation where the code is deterministic but the reported metric moves anyway. The complementary mistake is over-pinning. If a metric swings noticeably when only the seed changes, that variance is real information about the model's stability, and freezing the seed hides it rather than fixing it. The professional habit is to pin seeds so results are auditable, and separately to measure how much the metric moves across several seeds so you know how much of a reported improvement is signal. ## Related determinism notes `RandomForestClassifier` defaults to `n_estimators=100` and `max_depth=None`, meaning trees are grown until leaves are pure or contain fewer than `min_samples_split` samples. Fully grown trees are large, so a forest pickle can run to hundreds of megabytes — worth knowing when you wonder why the model artifact is so big. And if two teammates get different numbers with the same seed and the same data, the remaining suspects are the scikit-learn version (defaults do change between releases) and floating-point differences across platforms, not the forest's randomization.
- Does setting n_jobs=-1 make the results nondeterministic?No. `n_jobs` only distributes tree-building across workers; each tree's random draws are seeded deterministically from the estimator's `random_state`, so a seeded forest gives identical predictions at `n_jobs=1` and `n_jobs=-1`. Parallelism changes wall-clock time and memory use, not the fitted model.
- What is the difference between passing random_state=0 and passing a RandomState instance?An integer reseeds identically on every fit, so repeated fits reproduce each other exactly. A `numpy.random.RandomState` instance carries mutable state that advances as it is consumed, so the second fit continues the stream and produces a different forest. Use an integer when you want reproducibility, and reserve the instance for deliberately varying runs from one source.
- You pinned random_state on the forest but your reported accuracy still moves. What else is unseeded?Almost certainly the data split or the cross-validation splitter, each of which carries its own `random_state`. Any shuffling, resampling or augmentation step upstream has one too. Pinning the estimator makes the model reproducible given fixed data; it says nothing about how that data was partitioned.
- Is a metric that swings across seeds a problem to hide or information to use?Information. Seed-to-seed variance measures how stable the model is on this dataset, and a change smaller than that spread is not evidence of improvement. Pin a seed so runs are auditable, but also run several seeds and report the spread, so you can tell a real gain from sampling noise.
saying these in an interview costs you the question
- Blames n_jobs or thread scheduling for the variation
- Thinks a random forest should be deterministic by default
- Confuses the estimator seed with the train/test split seed
- Believes random_state changes model quality, not just the draw
- Treats seed-to-seed metric swings as a bug to suppress