skip to content

A video-metadata extractor calls `random.seed(1234)` at import; what does that do to the rest of the process?

level: seniorimportance: should knowfreq 33%

answer

  1. Where does the module keep its generator?
  2. The module functions are bound methods
  3. One hidden instance for the whole process
  4. Seeding is a process-wide side effect
  5. Construct random.Random for private state

basics

~20 s

It fixes the stream for the whole process. The module-level random functions are bound methods of one hidden random.Random instance, so seeding it makes every other component's jitter, sampling and shuffling replay the same sequence. Give the extractor its own random.Random(1234).

solid answer

~40 s

`random.seed` does not seed "the extractor's randomness" — there is no such thing. The module-level callables are bound methods of a single hidden `random.Random` instance created at import, so one `random.seed(1234)` anywhere reseeds the generator that retry jitter, sampling, shuffling and load spreading all draw from. Every worker process that imports that module then produces an identical sequence: the extractor picks the same frame indices in every worker, which looks like a win when the cache-hit rate jumps to 83%, and is actually lost diversity. Worse, the culprit is a one-line import side effect, so if the extractor's own failures are swallowed nothing points at it. The fix is isolation: `rng = random.Random(1234)` inside the module and call `rng.sample(...)`, leaving the global generator alone.

code

python · 13 lines
python
import random

def seed_globally():
    random.seed(1234)          # mutates the process-wide generator

private = random.Random(1234)  # independent state

seed_globally()
a = [random.randrange(1000) for _ in range(3)]
seed_globally()
b = [random.randrange(1000) for _ in range(3)]
print(a == b)                  # True: every consumer now replays one stream
print(len(private.sample(range(1000), 3)))

go deeper

for a junior

Know that random.seed affects the whole program, not just your file, and that you can create your own generator with random.Random(seed) when you want a repeatable sequence of your own.

for a middle

Explain that the module-level functions are bound methods of one hidden random.Random, and demonstrate isolating state with an instance plus saving and restoring the global with random.getstate and random.setstate.

for a senior

Diagnose it in production: connect synchronised backoff, duplicated sampling and a suspiciously high cache-hit rate back to one import-time seed, and describe the rollout of per-component generators without breaking existing reproducibility promises.

for a principal

Decide where determinism belongs in the architecture — injected generators as a convention, seeds recorded with results rather than reused, and a rule that libraries never touch process-wide state at import time.

## There is only one global generator The `random` documentation states it plainly: the functions the module exposes are bound methods of a hidden instance of `random.Random`, created when the module is first imported. `random.random`, `random.choice`, `random.shuffle`, `random.sample`, `random.randrange` and `random.seed` all operate on that one object. It is shared by every module in the process — your code, your dependencies, the framework, the client libraries. So `random.seed(1234)` executed at import time in a video-metadata extractor is not a local setting. It is a process-wide mutation performed as a side effect of importing a module, and it takes effect for whoever imports it, in whatever order imports happen to run. ## What the symptoms look like The extractor wanted reproducible frame sampling, and it gets it. So does everything else: * Retry and backoff jitter stops being jittered. Every worker that started from the same seed backs off on the same schedule, so a downstream dependency that a spread was meant to protect gets synchronised bursts instead. * Any sampling or A/B assignment built on `random` returns the same picks in every process. * The extractor's frame cache reports an 83% hit rate, up from a low baseline. That reads as a performance improvement in a dashboard, and it is really every worker requesting the same frame indices from the same videos — the sampling diversity the pipeline was designed around is gone. The diagnosis is made harder when the extractor swallows its own exceptions: the module that changed global state is precisely the module producing no error output, so nothing in the logs points at the import that caused it. When randomness "stops being random" across a fleet, grep for `random.seed` before anything else. ## The fix: own your generator ```python import random class FrameSampler: def __init__(self, seed=None): self._rng = random.Random(seed) # private state, nobody else touches it def pick(self, frame_count, k): return self._rng.sample(range(frame_count), k) ``` `random.Random(seed)` constructs an independent generator with its own state. Reproducibility is now a property of this object, not of the process. Libraries should never call the module-level `random.seed`; applications may, at the very top of `__main__`, and even then an explicit instance is clearer. If you must touch the global generator temporarily — in a test, say — bracket it with `random.getstate()` and `random.setstate()` so the previous state is restored, and note that `getstate` returns an opaque tuple describing the internal state, not a seed. ## Fork, workers and inherited state A forked child inherits the parent's generator state byte for byte, so two workers forked after the same seeding produce identical sequences. This is a classic source of duplicate "random" values across a worker pool. It matters less by default on 3.14 than it used to: the `multiprocessing` default start method is now `forkserver` on Unix other than macOS, while macOS and Windows use `spawn`, and both of those re-import `random` in a fresh interpreter which seeds from OS entropy. Choose `fork` explicitly and the inheritance is back — as it is for any hand-rolled `os.fork` usage. ## Two more things a senior is expected to add Threads share the instance too. Concurrent draws interleave against one state, so seeding buys you nothing reproducible once more than one thread is drawing — another reason per-component instances are the right unit. Reproducibility across Python versions has a narrow guarantee, and it is worth stating precisely: the documentation promises that `random.random()` will keep producing the same sequence for the same seed with a compatible seeder. The higher-level helpers built on top of it — `sample`, `choice`, `shuffle` — carry no such promise, so pinning a seed does not pin those results across upgrades. Persist the chosen indices if the selection itself must be reproducible years later.

  • How would you make one test deterministic without seeding the global generator?
    Inject the generator. Let the component take a `random.Random` instance (defaulting to a fresh one) and pass `random.Random(1234)` from the test, so determinism is scoped to the object under test. If the code cannot be changed yet, bracket the test with `random.getstate()` and `random.setstate()` so whatever the rest of the suite depended on is restored afterwards.
  • Two worker processes emit identical 'random' sequences. What do you check first, and what changed in 3.14?
    Check whether the workers were forked after the generator was seeded, since a forked child inherits the parent's state exactly. On 3.14 the `multiprocessing` default start method became `forkserver` on Unix other than macOS, with macOS and Windows on `spawn`; both re-import `random` in a fresh interpreter and reseed from OS entropy. If the code asks for `fork` explicitly, or forks by hand, the inherited state is back.
  • Does pinning a seed guarantee the same `random.sample` results after a Python upgrade?
    No. The documented guarantee covers `random.random()` producing the same sequence for the same seed with a compatible seeder; the higher-level helpers layered on it, including `sample`, `choice` and `shuffle`, may change their consumption pattern between versions. If a selection has to be reproducible across upgrades, persist the chosen items rather than the seed.

saying these in an interview costs you the question

  • Assumes each module gets its own generator instance
  • Seeds the global generator to make one test deterministic
  • Thinks seeding improves the quality of the randomness
  • Believes forked workers always get fresh generator state
  • Calls `random.seed()` at import to initialise the module
  • Treats identical output across workers as a caching win

context