Why does a server mask or re-randomise a session-bound CSRF token in every response when the value is already unguessable?
answer
- the response channel, not guessability
- a secret repeated verbatim across responses
- reflected guess changes the compressed size
- random pad per response, same stored value
- masking is not rotation
basics
~20 sBecause a secret that appears byte-for-byte in many responses is recoverable through a compression side channel: an attacker whose text is reflected into the same response and who can watch its compressed size learns the value piece by piece. Masking makes the transmitted bytes differ each time.
solid answer
~40 sThe concern is not guessability, it is the response channel. If the same value appears verbatim in response after response, and an attacker can get text of their choosing reflected into one of those responses and observe its compressed size, then a guess that matches part of the value compresses better and the response comes out shorter. Repeating that recovers the secret without ever guessing it outright. Masking breaks the repetition: the server generates a random pad per response, sends the pad alongside the pad combined with the session value, and reverses the operation on the way in. The stored value never changes - only its representation on the wire does. This is why masking is not rotation and adds no entropy: it closes a side channel in the response, nothing more.
code
pseudocode · 12 lineson rendering a page:
pad = random_bytes(length_of(session.csrf_token))
masked = xor(pad, session.csrf_token)
write encode(pad) + encode(masked) into the page
# session.csrf_token is NOT changed
on unsafe request:
pad, masked = split(decode(submitted value))
candidate = xor(pad, masked)
if not constant_time_equal(candidate, session.csrf_token):
reject
proceedgo deeper
Recall the shape only: the same secret repeated in many responses can be teased out through how well those responses compress, so the bytes on the wire are randomised each time while the stored value stays put.
Explain the three conditions the channel needs - a compressed response, attacker-chosen text reflected into it, and an observer of the response size - and why a matching guess makes the response shorter.
Be exact about direction and scope: it is a leak through the response, it is scoped to this one value, the stored copy is unchanged, and any other secret sharing that compressed response is still exposed.
Frame the real decision as what may be compressed alongside reflected input at all. Per-value masking is a point fix; the estate-wide question is which responses mix untrusted input with secrets and whether compression belongs on them.
## The problem is repetition, not guessability A session-bound token is unguessable by construction. The weakness masking addresses is somewhere else entirely: the value is a **secret that appears verbatim in many responses**. Every page of the booking flow that renders a form carries the same bytes, and that repetition is what a compression side channel feeds on. ## How the channel works The attack needs three conditions to hold at once, and the third is the one people forget: 1. **The response is compressed.** Compression works by replacing repeated sequences with references to earlier occurrences, so a response containing two copies of the same string is smaller than one containing two different strings of the same length. 2. **The attacker gets text of their own choosing reflected into that response** - a search term echoed back, a name field redisplayed, a path fragment in an error message. 3. **The attacker can observe the size of the response** - as an observer on the network path, for instance. Encryption of the response does not hide its length. Given those, the attack is a guessing game with feedback. The attacker causes a response to include their guess at a prefix of the secret. If the guess matches, that text now duplicates part of the secret already in the response, compression exploits the duplication, and the response is a few bytes shorter than it would otherwise have been. Wrong guesses give no such saving. Character by character, the size signal walks the secret out - without any attempt ever being submitted to the check itself. ## What masking does Masking removes condition one's leverage by making the bytes on the wire different every time, while leaving the stored value alone: - generate a **random pad** of the same length as the value, fresh on each response; - combine pad and value with a reversible operation - an exclusive-or is the usual choice; - send **both** the pad and the combined result, so the page carries a pair that differs on every response; - on the way back in, split the pair, reverse the combination to recover the candidate value, and compare that against the stored one in constant time. Two responses rendered a second apart now share no long repeated sequence, so a reflected guess has nothing to duplicate and the size signal disappears. ## What masking is not | claim | reality | |---|---| | "it makes the value harder to guess" | the value is unchanged; the pad adds no entropy to the secret | | "it is a form of rotation" | the stored value is untouched, so an older page's token still validates | | "it fixes the compression channel" | it protects this one value; anything else secret in a compressed response is still exposed | | "it replaces the constant-time comparison" | the comparison is still needed, on the unmasked candidate | That last row is the one to internalise: masking is scoped to the token. If a response mixes attacker-influenced text with any other repeated secret, the same channel is open on that secret, and the structural answer is to stop compressing responses that reflect attacker input alongside secrets. ## Where this came from The channel was demonstrated first against compression at the transport layer, in the attack named **CRIME** (CVE-2012-4929), and then against compression of the response body itself, in the attack named **BREACH**. RFC 7457, the summary of known attacks on TLS and DTLS, collects these in its section 2.6 along with TIME, and states that HTTP implementations using CSRF tokens will need to randomise them - which is exactly the masking described above, written into the record as a requirement on the token, not on the compression. ## In an interview The distinguishing answer is the one that names the **direction**: this is a leak through the response, not a weakness in the value. A candidate who says "we mask it so it cannot be guessed" has the right ritual and the wrong model, and will not be able to explain why the stored value is left untouched or why masking changes nothing about an older page's validity.
- Is masking a substitute for not compressing secrets alongside reflected input?No. Masking protects one value. Any other secret that appears verbatim in a compressed response containing attacker-influenced text is exposed through the same channel - a session identifier printed into a page, an account number, an authorisation value. The structural answer is to keep reflected input and secrets out of the same compressed response.
- Does masking change what the server stores or when an older page stops working?Neither. The stored value is untouched, so a page rendered ten minutes ago carries a differently masked pair that still unmasks to the same value and still validates. That is precisely what separates masking from rotation, and why a system can do one, both or neither.
- What does an attacker need besides a compressed response to run this?Text of their choosing reflected into the same response as the secret, and a way to observe that response's size - being on the network path is enough, since encryption hides content but not length. Remove either condition and the channel closes, which is why some deployments simply refuse to compress responses that reflect request input.
saying these in an interview costs you the question
- Says masking makes the value harder to guess
- Treats masking as a form of rotation that expires older pages
- Thinks encryption of the response hides its length
- Believes masking closes the compression channel for every secret in the page
- Claims the unmasked candidate no longer needs a constant-time comparison