skip to content

What does the argument to `secrets.token_hex(16)` count, and how long is the result?

level: middleimportance: should knowfreq 40%

answer

  1. The argument is not a character count
  2. Ask what unit the parameter uses
  3. Each encoding expands bytes differently
  4. Hex doubles; base64url is four per three
  5. No argument means thirty-two bytes

basics

~10 s

It counts raw random bytes, not output characters. secrets.token_hex(16) requests 16 bytes — 128 bits — of entropy and returns a 32-character hexadecimal string, because hex encoding spends two characters per byte.

solid answer

~40 s

All three token helpers take the same argument: a number of random **bytes**, which is the entropy you are asking for, and each then encodes those bytes differently. `secrets.token_bytes(16)` returns 16 raw `bytes`. `secrets.token_hex(16)` returns a 32-character `str`, two hex characters per byte. `secrets.token_urlsafe(16)` returns a 22-character `str` using base64url characters with the padding stripped, so roughly four characters per three bytes. Called with no argument all three use 32 bytes, which is the documented reasonable default. Pick by destination: `token_bytes` for keys and binary protocols, `token_hex` for logs and fixed-width columns, `token_urlsafe` for URLs, cookies and headers. Never truncate the encoded string to fit a column — that throws away entropy; ask for fewer bytes instead.

code

python · 7 lines
python
import secrets

raw = secrets.token_bytes(16)
print(type(raw).__name__, len(raw))          # bytes 16
print(len(secrets.token_hex(16)))            # 32 hex characters
print(len(secrets.token_urlsafe(16)))        # 22 url-safe characters
print(len(secrets.token_urlsafe()))          # 43: default is 32 bytes

go deeper

for a junior

Remember that the number you pass is bytes of randomness, and that secrets.token_hex gives you twice that many characters. Knowing the three helper names and their return types is enough at this level.

for a middle

Be able to do the arithmetic aloud for all three helpers, including the four-characters-per-three-bytes expansion of secrets.token_urlsafe, and explain why encoding never changes how guessable a token is.

for a senior

Show judgement about sizing and storage: pick a byte count rather than slicing output, keep the random part separate from any readable prefix, and catch case-insensitive columns or comparisons in review.

for a principal

Set the house standard once — default byte count, encoding per destination, a scannable prefix convention for secret detection — so individual services stop inventing token formats and rotation tooling has one shape to handle.

## One argument, three encodings `secrets` gives you three token functions and they differ only in how the random bytes are dressed for output. The argument — `nbytes` — is always a count of raw random bytes, never a count of characters in the result. That single distinction is what the question is really testing, because teams routinely write `secrets.token_hex(64)` believing they asked for a 64-character token and quietly ship a 128-character one that overflows a column. The arithmetic per function: * `secrets.token_bytes(n)` returns a `bytes` object of exactly `n` bytes. * `secrets.token_hex(n)` returns a `str` of exactly `2 * n` characters, drawn from `0-9a-f`. * `secrets.token_urlsafe(n)` returns a `str` of `ceil(4 * n / 3)` characters, drawn from `A-Z`, `a-z`, `0-9`, `-` and `_`, with base64 `=` padding removed. So 16 bytes becomes 16 bytes, 32 hex characters, or 22 URL-safe characters. Thirty-two bytes becomes 32, 64 and 43 respectively. Calling any of them with no argument uses 32 bytes, which the standard library documents as a reasonable default for tokens. ## Entropy is bytes; length is presentation Sixteen bytes is 128 bits of entropy, and that is the number that matters for guessability. The encoding is pure presentation — hex expands the same 128 bits over 32 characters and base64url over 22, and neither adds a single bit. A useful sanity check when reviewing a token scheme: convert the string length back to bytes before judging it. A 32-character hex token carries 128 bits; a 32-character base64url token carries about 192; a 32-character token assembled by hand from a 10-digit alphabet carries about 106. The practical floor for a bearer value an attacker can guess at online is 16 bytes; 32 bytes is the comfortable default and costs nothing measurable. ## `str` versus `bytes` — the choice that bites This is Python's most reliable source of avoidable bugs, and the token API puts it front and centre. `token_bytes` returns `bytes`: that is what you want for a symmetric key, a salt, a nonce, or anything you will feed to a byte-oriented API. `token_hex` and `token_urlsafe` return `str`: that is what you want in a URL, an HTTP header, a JSON body or a text column. Passing the wrong one produces either a `TypeError` at the boundary, or — worse — a value silently stringified as `b'...'` and stored with the `b` and the quotes baked in. A detail specific to `token_urlsafe`: its alphabet is case-sensitive and includes `-` and `_`. Both characters survive URLs, path segments and cookie values untouched, which is the point of the function, but the case sensitivity means the value must not live in a case-insensitive database column or be compared after `str.lower()` — either silently collapses the space of distinct tokens. ## Truncation is the mistake to name When a token does not fit the column, the wrong fix is `token_hex(32)[:16]` and the right fix is `token_hex(8)`. They produce the same length, but reviewers reading the first cannot tell how much entropy survived, and the pattern generalises badly: someone later slices a `token_urlsafe` result to a byte boundary that does not exist and quietly halves the keyspace. Ask for the bytes you intend to keep. The same discipline applies to prefixes. Adding a readable prefix such as an environment marker is fine and useful for scanning secrets out of source control, as long as the prefix is understood to contribute zero entropy and the random part is sized on its own. ```python import secrets api_key = "kj_live_" + secrets.token_urlsafe(32) # 8 fixed chars + 43 random ones print(len(api_key), api_key.startswith("kj_live_")) ```

  • Which of the three helpers would you use for a symmetric key, and why?
    `secrets.token_bytes`, because a key is binary and the byte-oriented APIs that consume it expect `bytes`. Encoding it to hex first only to decode it again adds a round trip and a chance to store the encoded form by mistake. Reserve `token_hex` and `token_urlsafe` for values that have to travel through text: URLs, headers, JSON and database columns.
  • A reviewer asks why the token is 43 characters rather than a round number. What do you tell them?
    Because it was sized in bytes: `secrets.token_urlsafe(32)` asks for 32 bytes of entropy and base64url encodes them as `ceil(4 * 32 / 3)` = 43 characters with padding stripped. The awkward length is a consequence of the encoding, not of a chosen string length, and rounding it down by slicing would silently discard entropy.
  • Is `secrets.token_urlsafe` safe to put in a case-insensitive column or a URL path segment?
    The path segment is fine — that is exactly what the alphabet of letters, digits, `-` and `_` is chosen for, and the base64 padding is stripped. A case-insensitive column is not: the alphabet is case-sensitive, so folding case collapses distinct tokens onto each other and shrinks the effective keyspace considerably.

saying these in an interview costs you the question

  • Reads the argument as the length of the returned string
  • Slices the encoded token to fit a database column
  • Thinks hex encoding adds entropy over the raw bytes
  • Cannot say which helper returns `bytes` and which return `str`
  • Stores a `token_urlsafe` value in a case-insensitive column

context