Why must a PKCE `code_verifier` be fresh random data rather than a value the client can derive?
answer
- 43 to 128, but of what?
- characters, not octets or bits
- length is not entropy
- 32 random octets, base64url encoded
- one verifier per authorization request
basics
~20 sBecause the only thing PKCE guarantees is that the redeemer can produce a value nobody else can. A verifier derived from data an attacker can obtain, or reused across flows, can be recomputed or replayed, and an intercepted code becomes redeemable again.
solid answer
~40 sRFC 7636 constrains the `code_verifier` twice over. The grammar says 43 to 128 **characters** drawn from the unreserved set — A-Z, a-z, 0-9, hyphen, period, underscore and tilde. The security requirement is separate and stronger: the verifier should come from a cryptographic random source, and the specification RECOMMENDS 32 octets of random data, base64url-encoded, which yields 43 characters and 256 bits of entropy. Length alone is not entropy — a 128-character string derived from an install identifier or a timestamp satisfies the grammar and protects nothing, because whoever can reproduce the rule can reproduce the verifier. It must also be fresh per authorization request, so that one leaked verifier compromises one flow rather than all of them.
code
pseudocode · 9 lines# once per authorization request, never reused
octets = cryptographic_random(32) # not a general-purpose generator
verifier = base64url_no_padding(octets) # 43 characters, 256 bits
challenge = base64url_no_padding(sha256(ascii(verifier)))
store verifier against THIS flow only
send authorization request with challenge and method "S256"
# ... later, on the token request
send verifier, then discard it whether or not the exchange succeededgo deeper
Remember the rule of thumb: a new random verifier for every login attempt, from a source meant for secrets, and never a value the application can work out again.
Separate the two requirements cleanly — the 43 to 128 character grammar against the recommendation of 32 cryptographically random octets — and say why the second is what protects the flow.
Review a client for the source, the octet count and the lifetime of the pending verifier, and recognise that a weak generator produces a flow that succeeds every time while protecting nothing.
Decide how this is proven rather than assumed across many clients: where the random source is specified, who reviews it, and what evidence you would accept that no client ships a derived verifier.
## What the specification actually constrains Two separate requirements hide behind one parameter, and candidates routinely collapse them into each other. - **The grammar.** A `code_verifier` is 43 to 128 **characters** long, drawn from the unreserved set: `A-Z`, `a-z`, `0-9`, and the four characters hyphen, period, underscore and tilde. Those bounds are on the string that is sent, not on the random value behind it, and not on bits of entropy. - **The security property.** The verifier must be **cryptographically random**. RFC 7636 RECOMMENDS taking 32 octets from a suitable random number generator and base64url-encoding them, which produces exactly 43 characters carrying 256 bits of entropy — the shortest string the grammar allows, and a strong one. So the minimum length and the recommended construction happen to coincide at 43 characters, which is why so much code gets the shape right and the source wrong. ## Length is not entropy The grammar is a syntax check, and syntax checks cannot see where a value came from. Every one of these passes it and destroys the mechanism: 1. A constant compiled into the application, so every copy on every member's machine sends the same verifier. 2. A value derived from an install identifier or a device identifier the client already stores — stable, and often readable by other software on the same machine. 3. A timestamp, a counter, or a sequence seeded at startup, all of which a nearby attacker can narrow to a handful of candidates. 4. A digest of something public, such as the `client_id` and the current date, which anybody can recompute. 5. Output from a general-purpose random routine intended for sampling or shuffling rather than for secrets. | Generation approach | Satisfies 43-128 characters | What an attacker needs to redeem an intercepted code | |---|---|---| | constant shipped in the build | yes | one copy of the application | | derived from an install identifier | yes | the identifier, which is often local and stable | | timestamp or counter | yes | an approximate time and a few guesses | | 32 octets from a cryptographic source | yes | 256 bits of luck | ## What a predictable verifier hands over Remember what PKCE checks: the authorization server compares a value derived from the presented verifier against the challenge stored with the code. It has no idea how the verifier was produced. If an attacker can compute the verifier independently, they intercept the code, produce the matching verifier, and redeem it — and the comparison succeeds. `S256` does not help here at all, because the attacker never needed to reverse the digest; they reconstructed the input. That is the useful way to hold the two properties apart: - **One-wayness** protects the verifier from an observer of the authorization request. - **Entropy** protects the verifier from someone who never saw the request. A flow needs both, and only one of them is a parameter you can read off the wire. ## Freshness, one per authorization request The verifier is per authorization request, not per client, per user or per session. Reuse turns a single exposure into a standing key: a verifier that leaks once — from a crash dump, a debug log, a device that was examined — covers every code that was, or will be, bound to the same value. Generating it at the moment the authorization request is built, remembering it only until that flow's token request is answered, and discarding it afterwards keeps the blast radius at one login attempt. There is a practical corollary for clients that can have more than one flow in the air. The pending verifier has to be stored per flow, not in a single slot, or a second authorization request will overwrite the first and the earlier flow will fail its comparison at redemption with `invalid_grant`. ## Reviewing a client quickly Three questions settle it, and none of them requires reading the whole client: - Where does the random value come from, and is that source documented as cryptographic? - How many octets are drawn before encoding — 32, or whatever happened to make the length check pass? - When is the verifier discarded, and can two concurrent flows in the same client collide on it? If the first answer names a general-purpose generator, nothing after that matters: the string will be the right length, the challenge will be the right shape, every login will succeed, and the protection will be theatre.
- A client generates a 128-character verifier from a device identifier. Does the extra length help?No. The grammar is satisfied and the security property is not: an attacker who learns the device identifier reproduces the verifier at any length. Entropy comes from the source, not the string size, and 43 characters of cryptographic randomness is stronger than 128 derived ones.
- Is there any reason to use more than the recommended 32 octets?Not for security. 32 octets is 256 bits, which is already far past anything guessable, and encoding more pushes the string toward the 128-character ceiling for no benefit. The bound worth remembering is the floor: fewer octets means a shorter string that may fail the 43-character minimum outright.
saying these in an interview costs you the question
- Reuses one code_verifier across every authorization request
- Derives the verifier from the client_id or an install identifier
- Reads the 43 to 128 bound as octets or bits of entropy
- Uses a general-purpose random routine rather than a cryptographic one
- Thinks a 128-character predictable string is strong enough