How many random bits does the session identifier in a cookie need, and how do you derive that number?
answer
- derive it, do not quote it
- many live identifiers, not one
- guess rate times the window
- N x R x T over 2^B
- 64 is the floor, 128 the norm
basics
~20 sDerive it from four measured quantities: live sessions, the guess rate your front door permits, the window identifiers stay valid, and the hit probability you accept. Expected hits are about N x R x T / 2^B, which puts the floor near 64 bits and the norm at 128.
solid answer
~50 sAn attacker guessing a session identifier is not guessing one presenter's value — **any** live identifier is a hit, so the whole live population works for them. That makes the answer a budget, not a constant. Take `N` live identifiers, `R` guesses per second your rate limiting and connection budget actually allow, `T` the window those identifiers stay valid, and `ε` the probability of one hit you are willing to accept. Expected successes are roughly `N × R × T / 2^B`, so `B ≥ log2(N × R × T / ε)`. For a station with a couple of hundred sessions live, 64 bits clears any honest ε by more than ten orders of magnitude, which is why it is quoted as a floor; 128 is the norm because the headroom costs a few bytes and means nobody re-runs the arithmetic when the numbers change.
code
pseudocode · 12 lines# size the identifier, do not quote a number
N = 200 # live identifiers at the busiest hour
R = 50 # guesses/second the front door will pass
T = 12 * 3600 # seconds an identifier stays valid
epsilon = 1e-9 # acceptable probability of one hit
attempts = R * T # 2.16e6
chances = attempts * N # 4.32e8
bits_needed = log2(chances / epsilon) # ~58.6
# 58.6 -> round up past the 64-bit floor, then take the 128-bit norm
identifier = base64url(csprng_bytes(16)) # 128 bitsgo deeper
Know that the value in the cookie must be long and unpredictable, drawn from a cryptographically secure generator, and that sixteen random bytes is the usual answer. The derivation can wait; the reflex not to shorten it cannot.
Produce the arithmetic on demand: live identifiers times guess rate times window, divided by two to the power of the bits, compared against an acceptable probability. Explain why the live population is a multiplier rather than a detail.
Bring real numbers from your own service — measured peak sessions, the rate your front door actually passes — and say what you would monitor to know someone is guessing, since sizing prevents but never detects.
Decide the acceptable probability explicitly and record it, because everything downstream is derived from a number nobody wrote down. Then size for growth so the argument is not reopened annually as the population changes.
"Use 128 bits" is the right answer given for the wrong reason, and this question exists to find out whether a candidate can derive it. The number falls out of four quantities you can measure on your own service. ## What is actually being guessed Guessing a session identifier is not guessing *a particular* presenter's identifier. The attacker sends a request carrying a made-up value, and the scheduler either finds a record filed under it or does not. **Any** live identifier is a win, so the entire population of live sessions works in the attacker's favour: a station with two hundred people signed in is two hundred times easier to hit than one with a single session open, at identical identifier length. That multiplier is what separates this from sizing a one-off value such as a link mailed to a single person, where one value is live at a time. It is also why the honest answer is a budget rather than a constant. ## The four quantities 1. **N — live identifiers.** Everything currently in the store, including automation accounts and records that have expired logically but not yet been swept. Size for the busiest hour, not the average one. 2. **R — guesses per second the front door will actually accept.** Not the attacker's imagination: what survives your rate limiting and your connection budget, counted per source and then multiplied by the number of sources you believe they can bring. 3. **T — the window.** How long the identifiers being guessed stay valid, because guesses accumulate over the whole period during which a hit is still worth having. 4. **ε — the acceptable probability** of one successful guess across that window. This is a business number, and writing it down is most of the value of the exercise. The expected number of successful guesses is about **N × R × T / 2^B**, and you want it comfortably under ε, which rearranges to **B ≥ log2(N × R × T / ε)**. ## Working it on the scheduler Take two hundred live sessions, a front door that will pass fifty requests a second to a determined attacker, and a twelve-hour window: about 2.2 million guesses against 200 live values, so roughly 4.3 × 10^8 chances to land on something. | bits | distinct values | expected hits in that window | |---|---|---| | 64 | 1.8 × 10^19 | ~2 × 10^-11 | | 96 | 7.9 × 10^28 | ~5 × 10^-21 | | 128 | 3.4 × 10^38 | ~1 × 10^-30 | Sixty-four bits already clears any ε a station this size would write down, by about ten orders of magnitude. That is the honest basis for calling 64 a **floor**: it is not a magic threshold, it is the point below which the arithmetic starts to depend on numbers you cannot guarantee. ## Why the norm is 128 anyway - Every input to the arithmetic drifts. The population grows, rate limits get relaxed during an incident, someone lengthens the window — and nobody re-derives the number when they do. - The headroom is nearly free: sixteen random bytes instead of eight, once per sign-in, on a value that rides in a header. - It ends the argument. A reviewer asked "is this enough?" can answer without reconstructing the population estimate that was true two years ago. - It absorbs quiet losses. If part of the value is later used as a lookup prefix, or trimmed by something in the path, 128 bits leaves enough behind that the identifier is still unguessable. ## Where the bits have to come from Length is a ceiling on unpredictability, not a substitute for it. Two applied traps matter more here than any theory of generators: - **A digest of a predictable input has the entropy of the input, not of the digest.** A 256-bit hash of an incrementing counter is a 256-bit string that anyone can compute for the next twenty sessions. - **Composite values leak structure.** An identifier built as an account id plus a timestamp plus a short random tail is only as strong as the tail, and the tail is usually the part someone shortened for readability. ## What the number does not buy - Nothing about **theft**. An identifier copied off a shared studio machine is a valid identifier; entropy answers guessing and only guessing. - Nothing about **detection**. Sizing tells you a hit is improbable; counting failed lookups per source tells you someone is trying, and the two are separate controls. - Nothing about how long the identifier stays useful once issued, which is a different decision on a different clock.
- Why does the number of live sessions belong in the sum at all?Because the attacker does not need a specific session. Every identifier currently in the store is a winning guess, so the population multiplies their chance directly: two hundred live sessions make a blind guess two hundred times more likely to land than one live session does, at the same identifier length.
- A team generates identifiers as a 256-bit digest of an incrementing counter. How strong is that?Effectively not at all. A digest is deterministic, so anyone who works out the counter can compute the next several identifiers exactly. The output width is irrelevant; unpredictability comes from the input, which is why the bytes must be drawn from a cryptographically secure generator rather than derived from something knowable.
- Does raising entropy reduce the risk of an identifier being stolen?No. Entropy answers one attack — blind guessing — and a stolen identifier was never guessed. A copied value is indistinguishable from the issued one at any length, so theft is answered by how the value is carried, how the session is bound and how quickly the record can be destroyed.
Sizing a combination lock for a corridor of lockers, not for one locker. A thief trying one combination a second wins if it opens any occupied locker, so the number of occupied lockers belongs in the sum alongside the thief's speed and the hours they get.
saying these in an interview costs you the question
- 128 bits because that is what everyone uses
- A 32-character identifier means 32 bytes of entropy
- Hashing a counter produces a full-width unpredictable value
- Only one identifier is at risk, so the population is irrelevant
- More bits make a stolen identifier harder to replay