skip to content

secrets versus random

The default generator is a Mersenne Twister that becomes predictable once it has emitted enough output, so tokens and session ids come from elsewhere. The wrong import here is a real, shipped breach.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

Why must a password-reset token come from `secrets`, not `random`?

level: juniorimportance: must knowfreq 70%

answer

  1. Two random modules, two different purposes
  2. One of them is built for simulation
  3. Its outputs give away the generator state
  4. Mersenne Twister, 624 words of state
  5. The other wraps OS cryptographic entropy

basics

~20 s

The random module's default generator is a Mersenne Twister whose entire internal state can be reconstructed from a few hundred observed outputs, so later tokens become predictable. secrets draws each token from the operating system's cryptographic source instead.

solid answer

~40 s

`random` is built for simulation, not secrecy. Its module-level functions are bound methods of one hidden `random.Random` instance running MT19937, and that generator is completely determined by 19937 bits of state; roughly 624 consecutive 32-bit outputs are enough to recover the state and predict everything that follows. Password-reset links, session identifiers, API keys and CSRF tokens are exactly the values an attacker gets to collect many of, so this is a real compromise rather than a theoretical one. `secrets`, added in Python 3.6, wraps the operating system's cryptographic source and exposes `secrets.token_bytes`, `secrets.token_hex`, `secrets.token_urlsafe`, `secrets.choice` and `secrets.randbelow`. Use `secrets` for anything an attacker must not guess, and keep `random` for jitter, sampling, shuffling fixtures and simulation.

code

python · 7 lines
python
import secrets

token = secrets.token_urlsafe(32)
print(len(token), token.isascii())

reset_code = "".join(secrets.choice("ABCDEFGHJKLMNPQRSTUVWXYZ23456789") for _ in range(8))
print(len(reset_code))

go deeper

for a junior

Memorise the split and the reason: secrets for tokens, keys, passwords and session ids; random for jitter, sampling and test data. Be ready to name secrets.token_urlsafe and secrets.token_hex without looking them up.

for a middle

Explain the mechanism, not the rule. Say that the module-level random functions share one hidden random.Random running MT19937, that its outputs expose 19937 bits of state, and that recovering the state predicts every later value.

for a senior

Show you can find the bug in a codebase: grep for random in code paths that mint identifiers, spot the broad-except fallback to random, and describe the remediation — rotate every token already issued, not just fix the import.

for a principal

Own the guardrail rather than the review comment. Decide whether a lint rule bans random in security modules, where token generation is centralised so it is written once, and what the response is when a predictable-token report arrives against issued credentials.

## Two generators, two jobs Python ships two sources of randomness with almost identical-looking APIs, and the entire security question is which one a value came from. `random` exists for simulation and statistics. The module-level callables — `random.random`, `random.randint`, `random.choice`, `random.shuffle`, `random.getrandbits` — are bound methods of a single hidden `random.Random` instance created when the module is first imported. That instance runs the Mersenne Twister MT19937 algorithm. MT19937 is fast, has an enormous period and excellent statistical properties, and is *entirely deterministic*: its whole future is a pure function of 19937 bits of state, held as 624 32-bit words plus a position index. `secrets` exists for keys and tokens. Every draw goes through the operating system's cryptographic random source — the same source `os.urandom` exposes — which is designed so that observing output tells you nothing usable about what comes next. ## Why "it looks random" is not the test The failure with `random` is not that its output looks patterned; it passes statistical tests comfortably. The failure is *state recovery*. Because the tempering step MT19937 applies to each output word is invertible, an attacker who collects around 624 consecutive 32-bit outputs can undo the tempering, reassemble the internal state, and then generate the same stream the victim's process will generate — backwards as well as forwards. Values produced through `random.random()` discard some bits per draw, which costs the attacker more samples, not the attack itself. Now look at what a token *is*. A password-reset token is mailed to whoever asks for one. An attacker can request thousands for accounts they control, harvesting a long run of the generator's output, then predict the token that will be issued for an account they do not control. Session identifiers, invitation codes, "unguessable" upload URLs, CSRF values and API keys all share that shape: the attacker sees many outputs and needs to guess one more. This class of bug has shipped repeatedly in real products. Seeding does not save it. `random.seed(some_secret)` keeps the seed private, but the outputs still leak the state, and state recovery never needed the seed. Nor does length: a 64-character token built from `random.choice` over an alphabet is 64 characters of a predictable stream. ## The `secrets` API you actually use ```python import secrets candidates = ["alpha", "beta", "gamma"] secrets.token_urlsafe(32) # str, URL- and cookie-safe, 43 characters secrets.token_hex(32) # str, 64 hexadecimal characters secrets.token_bytes(32) # bytes, for keys and binary protocols secrets.choice(candidates) # unpredictable selection from a sequence secrets.randbelow(1_000_000) # unbiased integer in range(1_000_000) ``` Note the `str` versus `bytes` split: `token_bytes` returns `bytes` and is what you want for a key you will feed to a keyed hash; `token_hex` and `token_urlsafe` return `str` and are what you want in a URL, a header or a database column. All three take a byte count and default to 32 bytes when called with no argument. ## Where `random` is still the right import Keep `random` for retry jitter, load shaping, Monte Carlo work, picking a sample of records for a report, shuffling fixture data in tests, and anything else where the only requirement is variety. `random` is much faster than reading OS entropy on every draw, and reproducibility from a seed is a feature there rather than a flaw. The rule that survives an interview is a rule about consequence, not about the module names: if guessing the value gains an attacker anything — access, identity, a bypass, a leaked resource — it comes from `secrets`. Otherwise `random` is fine and cheaper. ## The anti-patterns to name out loud Two are worth being able to spot. The first is a fallback: code that tries `secrets`, catches a broad `Exception`, swallows it, and quietly falls back to `random` — a silent downgrade from secure to guessable that no test will ever notice. The second is post-processing: hashing or base64-ing a `random` value to "strengthen" it. Encoding adds no entropy; a predictable input hashed with a public algorithm produces a predictable output.

  • If I seed `random` from `os.urandom` so nobody knows the seed, is its output safe for tokens?
    No. The attack does not need the seed. Each emitted word exposes part of MT19937's internal state, and once about 624 consecutive outputs are observed the state is recoverable directly, so all subsequent values are predictable regardless of how the generator was seeded. Secrecy of the seed is not the property tokens need; a generator whose output does not reveal its state is.
  • Some code calls `secrets`, catches a broad exception and falls back to `random`. What is wrong with that?
    It converts a loud failure into a silent security downgrade. Reading OS entropy does not fail in normal operation, so the fallback path exists only to hide a genuine problem, and when it fires the service keeps issuing guessable tokens with no signal in the logs. Let the exception propagate: a service that cannot generate secure tokens should refuse to issue them.
  • Is `uuid.uuid4()` acceptable for a session identifier?
    It is defensible: `uuid.uuid4` draws from the same OS source `secrets` uses and carries 122 random bits, which is enough. But six of its 128 bits are fixed version and variant markers, and its canonical form is long and often assumed to be a harmless identifier that is safe to log. `secrets.token_urlsafe(32)` states the intent plainly and gives you full control of the length.

MT19937 is a card-shuffling machine with a glass case: watch enough cards come out and you can read the gears, then call every remaining card. The OS source is a fresh deck cut behind a curtain for each card.

saying these in an interview costs you the question

  • Says `random` is fine as long as you seed it with the current time
  • Argues a long token from `random` is safe because it is long
  • Thinks hashing a `random` value makes it unpredictable
  • Believes the risk is statistical bias rather than state recovery
  • Calls `random.seed()` before each token to refresh the randomness
  • Falls back to `random` when `secrets` raises, and logs nothing

context

open as a page

What does the argument to `secrets.token_hex(16)` count, and how long is the result?

level: middleimportance: should knowfreq 40%

basics

~10 s

It counts raw random bytes, not output characters. secrets.token_hex(16) requests 16 bytes — 128 bits — of entropy and returns a 32-character hexadecimal string, because hex encoding spends two characters per byte.

open as a page

A video-metadata extractor calls `random.seed(1234)` at import; what does that do to the rest of the process?

level: seniorimportance: should knowfreq 33%

basics

~20 s

It fixes the stream for the whole process. The module-level random functions are bound methods of one hidden random.Random instance, so seeding it makes every other component's jitter, sampling and shuffling replay the same sequence. Give the extractor its own random.Random(1234).

open as a page

When would you reach for `random.SystemRandom` instead of the `secrets` functions?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

When you need the wider random API from an unpredictable source. random.SystemRandom is a random.Random subclass fed by os.urandom, so it offers shuffle, sample, randrange and uniform; secrets exposes only choice, randbelow, randbits and the token helpers.

open as a page