Why does a forked child repeat a `random.Random(1234)` sequence while `random.random()` differs?
answer
- The child copies the parent's memory
- Generator state is just data in that memory
- The standard library protects only its own instance
- Something reseeds one of them in the child
- os.register_at_fork with after_in_child
basics
~10 sForking copies the whole address space, and a generator's state is just memory, so your own random.Random instance continues the parent's stream. The random module reseeds only its own module-level generator in the child.
solid answer
~40 s`os.fork()` duplicates the parent's memory, and `random.Random(1234)` keeps its Mersenne Twister state in that memory, so parent and child both continue from the same point and produce identical sequences forever. The module-level functions such as `random.random()` are backed by a hidden instance for which the `random` module itself registers an at-fork callback that reseeds it from operating-system entropy in the child — hence the difference. The fix for a generator you own is the same mechanism: `os.register_at_fork(after_in_child=rng.seed)`, since `seed()` with no argument draws from OS entropy. Alternatively construct the generator inside the worker after the fork, or seed it deterministically per worker when you need reproducible runs. Duplicated streams are silent: synchronized retry backoff, identical sampling, colliding random identifiers.
code
python · 11 linesimport os
import random
rng = random.Random(1234)
if os.fork() == 0:
print("child instance:", rng.random(), "module:", random.random(), flush=True)
os._exit(0)
os.wait()
print("parent instance:", rng.random(), "module:", random.random())go deeper
Recall that a forked child is a copy of the parent's memory, and a seeded generator's state lives in that memory. Know the two easy fixes: build the generator inside the worker, or reseed it there.
Explain why the module-level functions behave differently — the random module registers its own at-fork callback for its hidden instance — and write the equivalent os.register_at_fork(after_in_child=rng.seed) for a generator you own.
Recognize the symptom set: synchronized retry backoff, identical sampling across workers, colliding random identifiers. Be able to keep per-worker streams both decorrelated and reproducible, and to say which start methods make the problem impossible.
Own randomness as a platform concern: where seeds come from, whether runs must be replayable across a fleet, which generators third-party libraries hold, and the standing rule that anything security-bearing uses a cryptographic source rather than a seeded generator.
## What is really being observed `os.fork()` copies the parent's entire address space into the child. A generator object is just state in that memory — for `random.Random`, the 624-word Mersenne Twister state array plus an index — so the child continues from precisely the same point the parent was at. If the parent had drawn two numbers from `random.Random(1234)`, both parent and child will hand out the *third* number of that stream next, and every number after it, forever. Two processes, one identical sequence. The module-level functions behave differently, and the reason is a hook rather than magic. CPython's `random` module registers its own at-fork callback for the hidden instance that backs `random.random()`, `random.choice()` and friends, reseeding it from operating-system entropy in the child. That is why the two lines below disagree: ```python import os import random rng = random.Random(1234) if os.fork() == 0: print("child instance:", rng.random(), "module:", random.random()) os._exit(0) os.wait() print("parent instance:", rng.random(), "module:", random.random()) ``` The `instance:` values match exactly; the `module:` values differ. The stdlib protected its own generator and could not protect yours, because it does not know your object exists. ## Why duplicate streams matter The failure is silent and statistical rather than loud. Retry backoff computed from a duplicated generator makes every worker retry at the same instant, converting a small hiccup into a synchronized thundering herd. Sampling one log line in a thousand from duplicated state means every worker samples the *same* positions, so coverage is a fraction of what the dashboard claims. Randomized identifiers or temporary filenames collide across workers. Shuffled work assignments come out identical, so several workers grind the same shard. None of these raise an exception; they show up as skew, collisions and correlated timings that look like an infrastructure problem. ## The fixes **Register a reseed hook for your own generator.** The same mechanism the stdlib uses is public: ```python import os import random rng = random.Random(1234) os.register_at_fork(after_in_child=rng.seed) ``` `random.Random.seed()` with no argument reseeds from operating-system entropy, so each child diverges. This is the right fix for a generator owned by a library that cannot control how it is used. **Or create the generator after the fork.** An object built inside the worker was never duplicated, so there is nothing to repair. Where the worker's entry point is under your control, this is simpler and works under every `multiprocessing` start method rather than only under `fork`. **Or seed deliberately per worker when you need reproducibility.** Reseeding from entropy destroys the ability to replay a run. Constructing `random.Random(base_seed + worker_index)` inside each worker keeps runs reproducible while giving each worker a different stream. Reseeding every child with the same constant is not a fix — it recreates the identical-stream problem with extra steps. ## The boundaries worth knowing * **Only `fork` is affected.** `multiprocessing` workers started with `spawn` or `forkserver` are fresh interpreter startups: module-level code runs again and generators are constructed anew, so nothing is duplicated. Since Python 3.14 the default start method on Linux is `forkserver` (macOS and Windows already used `spawn`), so this bug is now something you meet mainly when `fork` is requested explicitly or on an older interpreter — which makes it easy to reintroduce by "optimizing" a worker pool back to `fork`. * **Third-party generators are on their own.** An array library or a simulation framework that keeps its own generator state gets no help from the `random` module's hook; each such library must register its own, and several do. * **This is not a security control.** `random` is a Mersenne Twister: fast, uniform, and fully predictable from enough observed output whether or not it was reseeded. Tokens, password-reset links and session identifiers belong to `secrets` or `os.urandom` regardless of forking. ## The one-line summary to give an interviewer Fork duplicates memory, and a seeded generator is memory; the `random` module quietly reseeds *its own* instance in the child through an at-fork hook, and any generator you constructed yourself needs the same treatment — `os.register_at_fork(after_in_child=rng.seed)`, or build it after the fork.
- How do you give each forked worker a different stream while keeping the run reproducible?Do not reseed from entropy, which destroys replayability. Give each worker an index and build `random.Random(base_seed + worker_index)` inside the worker after the fork. Streams are decorrelated and the whole run can be replayed by recording the base seed. Reseeding every child with the same constant is not a fix — it recreates the identical-stream problem.
- Is the `random` module acceptable for security-sensitive values once it is properly reseeded per worker?No. `random` is a Mersenne Twister: fast and uniform, but its internal state can be reconstructed from enough observed output, so future values become predictable regardless of how it was seeded. Tokens, password-reset links and session identifiers belong to the `secrets` module or `os.urandom`, which draw from the operating system's cryptographic source.
- What tells you in production that workers are drawing from duplicated generator state?Correlated behaviour rather than an exception: retries from every worker landing in the same instant instead of spreading out, sampled logs covering the same positions in each worker, or randomized identifiers and temporary names colliding across processes. Print a few draws per worker with `os.getpid()` and identical values across pids confirm it immediately.
Forking a seeded generator is like handing two people identical copies of the same pre-shuffled deck: each deals honestly, and both deal exactly the same cards in exactly the same order.
saying these in an interview costs you the question
- Assumes every generator is reseeded automatically after a fork
- Reseeds each child with the same constant seed
- Blames the operating system's entropy pool
- Thinks seeding in the parent later affects running children
- Uses random for tokens because the output looks unpredictable
- Believes spawn-created workers inherit generator state too