skip to content

When should you create your own random.Random instance instead of calling the module-level random functions?

level: middleimportance: should knowfreq 45%

answer

  1. The module functions are one shared object
  2. Seeding is global; ownership is not
  3. Who else draws from this stream?
  4. Per-thread streams need per-thread generators
  5. Construct random.Random and inject it

basics

~20 s

Whenever the stream must belong to you: library code, anything that seeds for reproducibility, and per-thread draws. The module-level functions share one hidden generator that any other code in the process can reseed or consume.

solid answer

~40 s

`random.random()`, `random.choice()` and the rest are bound methods of a single `random.Random` instance created at import and shared by every module in the process. That is fine for a throwaway script, and wrong the moment reproducibility or isolation matters: a library that calls `random.seed()` silently redirects the application's stream, and an application that seeds cannot stop an imported helper from consuming draws and shifting its results. Constructing `rng = random.Random(seed)` gives an object with its own state — the same API as the module (`rng.random()`, `rng.choice()`, `rng.sample()`, `rng.shuffle()`, `rng.gauss()`), immune to `random.seed()` elsewhere. It is also the answer for threads: concurrent draws from one shared generator interleave non-deterministically, so a thread that needs a repeatable sequence needs a generator of its own. Instances are cheap; make one per component and pass it in.

code

python · 10 lines
python
import random

sampler = random.Random(20260903)
first = sampler.sample(range(1000), 3)

random.seed(0)  # some other module reseeds the shared generator

sampler2 = random.Random(20260903)
assert sampler2.sample(range(1000), 3) == first
print(first)

go deeper

for a junior

Know that random.choice() and friends all share one generator, and that random.Random(seed) gives you a separate one with the same method names.

for a middle

Explain why seeding is global, what breaks when a library seeds or consumes draws, and how passing a random.Random into a function removes the shared state entirely.

for a senior

Diagnose a run that stopped being reproducible without any code change: find who else touches the shared stream, and move the component onto an injected generator with a logged seed.

for a principal

Set the convention for a codebase — libraries never seed, components take an injected generator, seeds are recorded with results — so reproducibility is a property of the design rather than of what happened to be imported.

## The hidden instance The `random` module documents its own trick: the module-level functions are bound methods of a hidden instance of `random.Random`, created when the module is first imported. `random.random` really is `some_instance.random`. Everything that imports `random` anywhere in the process therefore shares one generator state. That design is deliberate — it makes `random.choice(colours)` a one-liner in a script — and it becomes a liability as soon as two pieces of code care about the stream at once. ## The two failure modes **A library that seeds.** Seeding is global. If a helper package calls `random.seed(0)` at import time to make its own examples tidy, it has just made every other consumer of `random` in that process deterministic and, worse, deterministic in a way that resets whenever that import happens. Library code should never call `random.seed()`; if it needs randomness it should own an instance, and ideally accept one from its caller. **An application that seeds and gets drifted.** Consider a chat-transcript archiver that seeds once at start-up so its nightly quality review always pulls the same 45 transcripts out of the day's batch. During the run, a retry helper it imported jitters its backoff with `random.uniform()`. Now the number of draws consumed before the sample depends on how many retries happened — so the "fixed" sample changes whenever the network does. Nothing raises. The report simply stops being comparable to yesterday's. An instance closes both holes: ```python import random sampler = random.Random(20260903) picks = sampler.sample(range(1000), 45) ``` `random.seed()` called anywhere else cannot touch `sampler`, and `sampler` consumes nothing anyone else is watching. ## What an instance gives you `random.Random` is a normal class. `random.Random(seed)` seeds at construction; `rng.seed(x)` reseeds later; `rng.getstate()` / `rng.setstate()` checkpoint it. Every module-level function has a method twin — `rng.random()`, `rng.randint()`, `rng.randrange()`, `rng.choice()`, `rng.choices()`, `rng.sample()`, `rng.shuffle()`, `rng.gauss()`, `rng.getrandbits()`, `rng.randbytes()` — so switching from module functions to an instance is a mechanical edit, not a redesign. Because it is a class, it is also an injection point. A function that takes `rng: random.Random` can be handed a seeded generator in a test and an unseeded one in production, with no global state and no monkeypatching of the module. ## Threads and other concurrency Under threads the shared generator has a second problem beyond who seeds it: order. Two threads drawing from one stream interleave in a way the scheduler decides, so "seed once, get the same sequence" is false even when nothing else in the process touches `random`. Give each thread its own `random.Random`, seeded from a per-thread value you control, and each thread's sequence is stable regardless of how the threads interleave. One concrete wrinkle: `gauss()` generates normal deviates in pairs and caches the spare between calls, so it carries a little extra per-instance state on top of the twister. That is another reason not to share one instance across threads and expect tidy behaviour. Processes are different again: a child created with the `fork` start method inherits the parent's generator state, so several children can produce identical "random" sequences unless each reseeds. Under `spawn` and `forkserver` — and on Unix other than macOS `forkserver` is the default start method as of Python 3.14 — the child imports `random` fresh and seeds itself from OS entropy, so the clash does not arise. ## When the module functions are fine Do not over-correct. A short script, a REPL session, a one-off data generator: the module functions are exactly what they are for. The rule of thumb is ownership — if the code is a library, if it seeds, if it runs concurrently, or if someone will later ask "why did this run differ?", it should own a `random.Random`. If the answer to "who else draws from this stream?" is "nobody, and nobody ever will", the shared instance costs nothing. And for values that must be unguessable, neither choice helps: both are the same predictable twister. That case needs `random.SystemRandom`, which draws from the operating system instead. ## Making the seed a first-class input Once a component owns its generator, the seed becomes an ordinary parameter rather than a hidden constant. Accept it from configuration or generate it at start-up, log it beside the results, and a run that produced a surprising sample can be re-run exactly by passing the same value back. That is strictly better than hardcoding `random.Random(0)` everywhere: you keep varied data across runs, and you keep the ability to replay any single one of them.

  • Should a library ever call random.seed()?
    No. Seeding is process-global, so a library that seeds hijacks the stream of every other consumer in the process and makes their output depend on when the library happened to be imported. A library should hold its own `random.Random`, and ideally accept one from the caller so the application decides whether the run is reproducible.
  • Two threads share one seeded random.Random. Why is the output still not reproducible?
    Because the draws interleave in whatever order the scheduler chooses, so each thread sees a different slice of one stream from run to run. The sequence as a whole is fixed, but its distribution across threads is not. Give each thread its own generator, seeded from a value you control, if per-thread sequences must repeat.
  • Is it expensive to create many random.Random instances?
    Not meaningfully at normal scale — construction allocates and seeds a 624-word state, so it is heavier than a single draw but trivial next to any real work. The cost that matters is seeding one per call inside a hot loop, which is both slow and statistically poor. Create one per component and reuse it.

The module-level functions are a shared notepad on a communal desk: anyone can flip it to a new page. An instance is your own notepad in your own drawer.

saying these in an interview costs you the question

  • Calls random.seed() from inside library code
  • Thinks each module draws from its own separate generator
  • Seeds once and assumes threads get repeatable sequences
  • Creates a new generator inside a hot loop for every draw
  • Believes a private instance is any harder to predict

context