Why is output from random.random() predictable, and what does random.SystemRandom change?
answer
- Statistical quality is not secrecy
- A public, invertible transform on visible state
- A few hundred outputs reveal the generator
- Ask the operating system instead
- SystemRandom seed does nothing, getstate raises
basics
~20 sThe random module runs a Mersenne Twister whose whole state can be reconstructed from a few hundred observed outputs, making past and future values computable. random.SystemRandom draws each value from the operating system's cryptographic source instead, giving up seeding and reproducibility.
solid answer
~50 s`random`'s generator, MT19937, is designed for statistical quality, not secrecy: its state is 624 32-bit words, it is not one-way, and an observer who collects enough consecutive outputs can invert the tempering step, recover the full state, and then reproduce every value the generator will ever emit — plus the ones it already did. A seed derived from a timestamp or a small integer is worse still, since it can simply be brute-forced. `random.SystemRandom` is a `random.Random` subclass that overrides the low-level draws to read the operating system's CSPRNG instead of the twister, so there is no state to steal: `seed()` is a stub that does nothing, and `getstate()`/`setstate()` raise `NotImplementedError`. You keep the same convenience API — `choice()`, `sample()`, `shuffle()`, `randrange()` — and lose reproducibility by design. For tokens and keys specifically, the standard library's `secrets` module wraps the same source with the right helpers.
code
pycon · 7 lines>>> import random
>>> rng = random.SystemRandom()
>>> rng.seed(42) # accepted, but does nothing
>>> rng.getstate()
Traceback (most recent call last):
...
NotImplementedError: System entropy source does not have state.go deeper
Remember the split: random is for simulations, samples and test data, and never for a value someone must not guess. That single rule prevents most of the damage.
Explain why the generator is predictable — invertible tempering over a finite recoverable state — and name random.SystemRandom as the operating-system-backed alternative with the same helper methods.
Spot the pattern in review: any identifier, nonce or ordering that an outsider profits from predicting, sourced from a seedable generator. Be ready to argue why reseeding, hashing or wider values do not fix it.
Own the boundary as a standing rule — which values in a system are adversarial, where the OS-backed source is mandatory, and how the team keeps reproducibility for simulations without letting it leak into security-relevant paths.
## What "predictable" means here MT19937 keeps 624 32-bit words of state. Each draw pulls one word, applies a fixed, invertible bit-mixing step called tempering, and returns it; every 624 draws the state is refreshed by a deterministic twist. Two properties follow directly. First, the mapping from state to output is public and invertible — tempering is a sequence of shifts, XORs and masks with published constants, so an output word can be turned back into the raw state word. Second, the twist is deterministic, so knowing 624 consecutive raw words is knowing the generator forever, in both directions. So an observer who can see enough consecutive outputs can rebuild the state and then say exactly what comes next. `random.random()` builds a float from two 32-bit draws, so roughly 312 observed floats is the order of magnitude that matters — a few hundred values, not billions. This is not a hypothetical: it is a standard exercise, and no amount of reseeding or mixing values with a timestamp fixes it, because the weakness is in the generator, not the seed. Small seeds make it easier rather than harder. Seeding from the current time in seconds leaves only a few million candidates for a known window; a caller can simply try them all and match against one observed value. A concrete shape of the mistake: a chat-transcript archiver mints share identifiers for exported conversations with `''.join(random.choice(alphabet) for _ in range(12))`. Each identifier looks fine in isolation. But every identifier is a window onto the same shared stream, so a user who has legitimately exported a few hundred transcripts of their own has, in effect, been handed the generator — and can compute the identifiers other people's exports will get. ## What SystemRandom is `random.SystemRandom` is a subclass of `random.Random` that replaces the primitive operations — the raw float draw, `getrandbits()`, `randbytes()` — with reads from the operating system's cryptographic random source (`/dev/urandom` and its equivalents, the same source `os.urandom()` exposes). Everything built on those primitives is inherited unchanged, so the convenience API still works: ```python import random rng = random.SystemRandom() print(rng.choice("abcdefghjkmnpqrstuvwxyz")) print(rng.sample(range(1000), 3)) ``` Because there is no user-visible state, the state API is deliberately broken: - `seed()` exists but is a stub that does nothing, so calling it does not raise and does not make output reproducible. - `getstate()` and `setstate()` raise `NotImplementedError` with the message that the system entropy source has no state. That asymmetry is the design speaking: reproducibility and unpredictability are opposites, and `SystemRandom` picks unpredictability. ## Cost and choosing between them Each `SystemRandom` draw goes to the OS rather than doing arithmetic in memory, so it is slower — typically by a large factor per call, though still fast in absolute terms. For identifiers, nonces, shuffling something an adversary cares about, or picking a value that must not be guessable, that cost is irrelevant. For a Monte Carlo loop drawing tens of millions of values, it is the wrong tool and the twister is correct. The rule of thumb is about consequences, not volume: if someone benefits from guessing the next value, the value must come from the OS source. Otherwise use the seedable generator, which is faster and reproducible. For tokens, passwords and keys specifically, the standard library ships `secrets`, which is built on the same system source and adds the right-shaped helpers plus a constant-time comparison; reach for it rather than assembling tokens by hand from `SystemRandom`. ## Things that do not help Several popular non-fixes are worth naming, because they show up in review: - **Reseeding often from the clock.** The clock is low-entropy and observable; it narrows the search rather than widening it. - **Hashing the output.** The attacker recovers state from the pre-hash values they can see; where they cannot see values at all, hashing adds nothing. - **Mixing several `random` calls together.** All of them come from the same stream, so combining them reveals more state, not less. - **Using `random.getrandbits(256)`.** The width of a value says nothing about the entropy behind it; 256 bits from a 19937-bit deterministic state seeded weakly is still guessable. The only fix is the source. In review, the question to ask about any random value is simply: does anything bad happen if an outsider computes this before we do? If yes, it must not come from a seedable generator.
- Does calling random.seed() with a large, unpredictable value make the module-level generator safe?No. A strong seed hides the starting point but not the generator: an observer who collects a few hundred consecutive outputs reconstructs the state directly from those values and never needs the seed at all. The weakness is that MT19937's output-to-state mapping is public and invertible, which no choice of seed changes.
- What actually happens when you call seed() on a random.SystemRandom instance?Nothing. It is a stub that accepts the argument and returns, so the call is silently useless rather than an error — which is exactly the trap: code that looks like it pinned a sequence for a test has pinned nothing. `getstate()` and `setstate()` are blunter and raise `NotImplementedError`, because a system entropy source has no state to expose.
- Is SystemRandom slow enough to matter?Per draw it is substantially slower than the twister, because each value comes from the operating system rather than from arithmetic on in-process state. At the volumes where unpredictability is the requirement — identifiers, nonces, one-off picks — it is irrelevant. In a simulation drawing tens of millions of values it is the wrong tool, and there the seedable generator is the right one anyway.
saying these in an interview costs you the question
- Says a big or clock-based seed makes the twister unpredictable
- Thinks hashing the output hides the generator state
- Believes more bits from getrandbits means more entropy
- Assumes SystemRandom can be seeded for a reproducible test
- Uses the seedable generator for anything an outsider benefits from guessing
- Claims predicting the twister needs millions of observed values