When should you create your own random.Random instance instead of calling the module-level random functions?
answer
- The module functions are one shared object
- Seeding is global; ownership is not
- Who else draws from this stream?
- Per-thread streams need per-thread generators
- Construct random.Random and inject it
basics
~20 sWhenever the stream must belong to you: library code, anything that seeds for reproducibility, and per-thread draws. The module-level functions share one hidden generator that any other code in the process can reseed or consume.
solid answer
~40 s`random.random()`, `random.choice()` and the rest are bound methods of a single `random.Random` instance created at import and shared by every module in the process. That is fine for a throwaway script, and wrong the moment reproducibility or isolation matters: a library that calls `random.seed()` silently redirects the application's stream, and an application that seeds cannot stop an imported helper from consuming draws and shifting its results. Constructing `rng = random.Random(seed)` gives an object with its own state — the same API as the module (`rng.random()`, `rng.choice()`, `rng.sample()`, `rng.shuffle()`, `rng.gauss()`), immune to `random.seed()` elsewhere. It is also the answer for threads: concurrent draws from one shared generator interleave non-deterministically, so a thread that needs a repeatable sequence needs a generator of its own. Instances are cheap; make one per component and pass it in.
code
python · 10 linesimport random
sampler = random.Random(20260903)
first = sampler.sample(range(1000), 3)
random.seed(0) # some other module reseeds the shared generator
sampler2 = random.Random(20260903)
assert sampler2.sample(range(1000), 3) == first
print(first)go deeper
Know that random.choice() and friends all share one generator, and that random.Random(seed) gives you a separate one with the same method names.
Explain why seeding is global, what breaks when a library seeds or consumes draws, and how passing a random.Random into a function removes the shared state entirely.
Diagnose a run that stopped being reproducible without any code change: find who else touches the shared stream, and move the component onto an injected generator with a logged seed.
Set the convention for a codebase — libraries never seed, components take an injected generator, seeds are recorded with results — so reproducibility is a property of the design rather than of what happened to be imported.
## The hidden instance The `random` module documents its own trick: the module-level functions are bound methods of a hidden instance of `random.Random`, created when the module is first imported. `random.random` really is `some_instance.random`. Everything that imports `random` anywhere in the process therefore shares one generator state. That design is deliberate — it makes `random.choice(colours)` a one-liner in a script — and it becomes a liability as soon as two pieces of code care about the stream at once. ## The two failure modes **A library that seeds.** Seeding is global. If a helper package calls `random.seed(0)` at import time to make its own examples tidy, it has just made every other consumer of `random` in that process deterministic and, worse, deterministic in a way that resets whenever that import happens. Library code should never call `random.seed()`; if it needs randomness it should own an instance, and ideally accept one from its caller. **An application that seeds and gets drifted.** Consider a chat-transcript archiver that seeds once at start-up so its nightly quality review always pulls the same 45 transcripts out of the day's batch. During the run, a retry helper it imported jitters its backoff with `random.uniform()`. Now the number of draws consumed before the sample depends on how many retries happened — so the "fixed" sample changes whenever the network does. Nothing raises. The report simply stops being comparable to yesterday's. An instance closes both holes: ```python import random sampler = random.Random(20260903) picks = sampler.sample(range(1000), 45) ``` `random.seed()` called anywhere else cannot touch `sampler`, and `sampler` consumes nothing anyone else is watching. ## What an instance gives you `random.Random` is a normal class. `random.Random(seed)` seeds at construction; `rng.seed(x)` reseeds later; `rng.getstate()` / `rng.setstate()` checkpoint it. Every module-level function has a method twin — `rng.random()`, `rng.randint()`, `rng.randrange()`, `rng.choice()`, `rng.choices()`, `rng.sample()`, `rng.shuffle()`, `rng.gauss()`, `rng.getrandbits()`, `rng.randbytes()` — so switching from module functions to an instance is a mechanical edit, not a redesign. Because it is a class, it is also an injection point. A function that takes `rng: random.Random` can be handed a seeded generator in a test and an unseeded one in production, with no global state and no monkeypatching of the module. ## Threads and other concurrency Under threads the shared generator has a second problem beyond who seeds it: order. Two threads drawing from one stream interleave in a way the scheduler decides, so "seed once, get the same sequence" is false even when nothing else in the process touches `random`. Give each thread its own `random.Random`, seeded from a per-thread value you control, and each thread's sequence is stable regardless of how the threads interleave. One concrete wrinkle: `gauss()` generates normal deviates in pairs and caches the spare between calls, so it carries a little extra per-instance state on top of the twister. That is another reason not to share one instance across threads and expect tidy behaviour. Processes are different again: a child created with the `fork` start method inherits the parent's generator state, so several children can produce identical "random" sequences unless each reseeds. Under `spawn` and `forkserver` — and on Unix other than macOS `forkserver` is the default start method as of Python 3.14 — the child imports `random` fresh and seeds itself from OS entropy, so the clash does not arise. ## When the module functions are fine Do not over-correct. A short script, a REPL session, a one-off data generator: the module functions are exactly what they are for. The rule of thumb is ownership — if the code is a library, if it seeds, if it runs concurrently, or if someone will later ask "why did this run differ?", it should own a `random.Random`. If the answer to "who else draws from this stream?" is "nobody, and nobody ever will", the shared instance costs nothing. And for values that must be unguessable, neither choice helps: both are the same predictable twister. That case needs `random.SystemRandom`, which draws from the operating system instead. ## Making the seed a first-class input Once a component owns its generator, the seed becomes an ordinary parameter rather than a hidden constant. Accept it from configuration or generate it at start-up, log it beside the results, and a run that produced a surprising sample can be re-run exactly by passing the same value back. That is strictly better than hardcoding `random.Random(0)` everywhere: you keep varied data across runs, and you keep the ability to replay any single one of them.
- Should a library ever call random.seed()?No. Seeding is process-global, so a library that seeds hijacks the stream of every other consumer in the process and makes their output depend on when the library happened to be imported. A library should hold its own `random.Random`, and ideally accept one from the caller so the application decides whether the run is reproducible.
- Two threads share one seeded random.Random. Why is the output still not reproducible?Because the draws interleave in whatever order the scheduler chooses, so each thread sees a different slice of one stream from run to run. The sequence as a whole is fixed, but its distribution across threads is not. Give each thread its own generator, seeded from a value you control, if per-thread sequences must repeat.
- Is it expensive to create many random.Random instances?Not meaningfully at normal scale — construction allocates and seeds a 624-word state, so it is heavier than a single draw but trivial next to any real work. The cost that matters is seeding one per call inside a hot loop, which is both slow and statistically poor. Create one per component and reuse it.
The module-level functions are a shared notepad on a communal desk: anyone can flip it to a new page. An instance is your own notepad in your own drawer.
saying these in an interview costs you the question
- Calls random.seed() from inside library code
- Thinks each module draws from its own separate generator
- Seeds once and assumes threads get repeatable sequences
- Creates a new generator inside a hot loop for every draw
- Believes a private instance is any harder to predict