skip to content

What does calling random.seed(42) do to the sequence returned by random.random()?

level: juniorimportance: must knowfreq 62%

answer

  1. Same input, same output every run
  2. One hidden generator behind the module functions
  3. Twister state, not fresh entropy
  4. No argument means operating-system entropy
  5. getstate and setstate checkpoint the stream

basics

~20 s

It resets the internal state of the one Mersenne Twister generator that backs the random module's functions. Any run that seeds with the same value then replays the identical sequence of numbers in the same order.

solid answer

~40 s

The module-level functions in `random` are bound methods of a single hidden `random.Random` instance holding a Mersenne Twister. `random.seed(42)` sets that generator's state deterministically, so the same seed followed by the same calls in the same order yields identical results — `random.random()`, `random.randint()`, `random.choice()`, `random.sample()` and `random.shuffle()` all draw from that one stream. `random.seed()` with no argument (or `None`) seeds from operating-system entropy instead, which is what happens at import, so every process differs. Because the stream is shared process-wide, any other code that draws from it between your calls consumes numbers and shifts what you see. Reproducibility holds for a given Python version and any platform; `random.getstate()` and `random.setstate()` checkpoint and rewind mid-stream. Seeding never makes the output unpredictable to someone who has seen enough of it.

code

pycon · 7 lines
pycon
>>> import random
>>> random.seed(42)
>>> [random.randint(1, 100) for _ in range(3)]
[82, 15, 4]
>>> random.seed(42)
>>> [random.randint(1, 100) for _ in range(3)]
[82, 15, 4]

go deeper

for a junior

Be ready to say in one sentence that a seed fixes the generator's starting state so the same sequence comes back, and to show it with two seeded loops printing the same list.

for a middle

Explain the mechanics: one hidden random.Random instance shared by the module functions, each call advancing the state, and seed(None) pulling operating-system entropy at import.

for a senior

Show where reproducibility breaks in a running system — other code drawing from the shared stream, threads interleaving draws — and the habit of logging a run-time seed so a failing run can be replayed exactly.

for a principal

Own the policy question: how far a team should promise byte-for-byte reproducibility given that only the core generator is version-stable, and where determinism must be recorded as data rather than assumed from a hardcoded constant.

## One generator, one state `random` is a pseudo-random generator, not a source of entropy. Its core is the Mersenne Twister MT19937, whose entire state is 624 32-bit words plus an index. Every number the module hands out is a pure function of that state, and every draw advances it. "Random" here means *statistically well-distributed*, not *unknowable*. The functions you call as `random.random()`, `random.randint()`, `random.choice()` and friends are not free functions with private state: they are bound methods of one hidden `random.Random` instance created when the module is first imported. That single instance is shared by every module in the process. Seeding it is therefore a global act. ## What seed() actually accepts `random.seed(a=None, version=2)`: - **`None` or no argument** — seed from the operating system's entropy source, falling back to the current time if none is available. This is what happens automatically at import, and it is why two runs of an unseeded script differ. - **an `int`** — used directly; arbitrarily large integers are fine. - **a `str`, `bytes` or `bytearray`** — converted to an integer. Under the default `version=2` the conversion uses every bit of the string; `version=1` exists only to reproduce streams from very old Python releases. - **a `float`** — accepted, and derived from its exact binary value. After seeding, the sequence is fixed: ```pycon >>> import random >>> random.seed(42) >>> [random.randint(1, 100) for _ in range(3)] [82, 15, 4] >>> random.seed(42) >>> [random.randint(1, 100) for _ in range(3)] [82, 15, 4] ``` The determinism covers *the calls you make, in the order you make them*. Insert one extra `random.random()` in the middle and everything after it shifts, because each call consumes state. ## Where reproducibility silently breaks The usual bug is not the seed but the sharing. Suppose a chat-transcript archiver seeds once at start-up so its nightly quality review always inspects the same 45 transcripts. Then some helper it imports calls `random.seed()` itself, or draws a few values while building a retry delay. The archiver's numbers change, nothing raises, and the "reproducible" sample quietly stops being reproducible. The fix is an instance of your own — `random.Random(seed)` — which is state nobody else can touch. Threads are the same problem in miniature. Concurrent draws from the shared generator interleave in a non-deterministic order, so seeding once and expecting a fixed per-thread sequence does not work; give each thread its own generator instead. A third point of confusion: `PYTHONHASHSEED` has nothing to do with this module. It randomizes `str`/`bytes` hashing to defend against collision attacks, and setting it does not make `random` deterministic. Finally, a seeded generator is *reproducible*, which is the exact opposite of *unguessable*. Anything an adversary must not predict has to come from an OS-backed source such as `random.SystemRandom`, never from a seeded stream. ## Checkpointing instead of reseeding When you want to replay part of a stream rather than restart it, capture the state: ```pycon >>> random.seed(42) >>> state = random.getstate() >>> random.random() 0.6394267984578837 >>> random.setstate(state) >>> random.random() 0.6394267984578837 ``` `getstate()` returns an opaque tuple; treat it as a token to hand back to `setstate()`, not something to inspect or persist across Python versions. ## The stability promise For a given Python version, a given seed produces the same numbers on any platform — the twister is integer arithmetic and the float conversion is exact, so there is no CPU-dependent drift. Across versions the guarantee is narrower: the documentation promises that the *core generator* stays compatible, not that every derived helper keeps its exact algorithm. Historically `random.shuffle()` and `random.sample()` have changed implementation, so a stream that must survive a Python upgrade byte-for-byte should be pinned to a version, or reduced to raw `random.random()` / `random.getrandbits()` calls whose mapping you control. In practice, the useful habit for test suites is to choose the seed at run time, log it, and accept it back as a parameter. You keep the coverage of varied data and, when something fails, you can replay exactly the run that failed. ## Where to put the seed Seed once, at the point that owns the run — a `main()`, a test set-up, a job entry point — and never inside the function that draws. Reseeding before each draw is the classic anti-pattern: `random.seed(int(time.time()))` inside a loop makes every call within the same second return the same value, which is the opposite of what the author intended. One seed at the top, then draw freely, is the whole discipline.

  • What does random.seed() with no argument do?
    It seeds from the operating system's entropy source, falling back to the current time when no such source exists. That is exactly what happens once, automatically, when the module is first imported, which is why an unseeded script produces different numbers on every run. Calling it explicitly is how you deliberately throw away a fixed seed and go back to per-process variation.
  • Does the same seed give the same numbers on another machine or another Python version?
    On the same Python version, yes on any platform — the twister is integer arithmetic and the conversion to a float is exact, so there is no hardware drift. Across versions the promise is only that the core generator stays compatible; derived helpers such as `random.sample()` and `random.shuffle()` may change algorithm. Pin the interpreter version if byte-for-byte reproducibility matters.
  • How do you replay part of a stream without reseeding it?
    Capture `random.getstate()` before the section you care about and pass that tuple back to `random.setstate()` to rewind. The state is an opaque snapshot of the twister's 624 words plus its index; treat it as a token you hand back, not something to parse or store across interpreter versions.

A seed is the starting position of a very long pre-printed table of numbers: pick the same starting row and you read the same numbers, and anyone else reading over your shoulder can follow along.

saying these in an interview costs you the question

  • Says seeding makes the numbers cryptographically strong
  • Thinks PYTHONHASHSEED controls the random module
  • Believes randint and choice each have their own generator
  • Expects a fixed seed to survive other code drawing from the same stream
  • Assumes seeded output is identical across Python versions
  • Calls random.seed() before every single draw

context