What does `hmac.compare_digest` still leak, and how do you compare variable-length secrets?
answer
- The guarantee is narrower than people assume
- Contents are hidden; sizes are not
- Docs say types and lengths, not values
- Make both sides the same width first
- A fixed-size digest before comparing
basics
~20 sIt hides the contents but not the sizes: a timing attack could still reveal the lengths and types of the two arguments, never their values. Hash both sides to a fixed-size digest first, then compare the digests.
solid answer
~50 sThe documented guarantee is narrower than people assume. `hmac.compare_digest` never returns early on a content mismatch, but its loop runs a number of times that follows the size of what it was given, so a length mismatch or an error could in principle reveal the lengths and types of the arguments — not their values. For message authentication tags and digests this is a non-issue: both sides are the same fixed size by construction. It bites when the secret itself is variable length — historical API keys of assorted formats, for instance. The fix is one line: run both sides through a fixed-size digest, say `hashlib.blake2b(value, digest_size=32).digest()`, and compare the two 32-byte results. Lengths then always match, and the encoding question disappears at the same time. That is safe for high-entropy random tokens; anything guessable belongs in a slow key-derivation function instead.
code
python · 12 linesimport hashlib
import hmac
def tokens_match(presented: str, stored: str) -> bool:
a = hashlib.blake2b(presented.encode("utf-8"), digest_size=32).digest()
b = hashlib.blake2b(stored.encode("utf-8"), digest_size=32).digest()
return hmac.compare_digest(a, b)
print(tokens_match("short", "a-much-longer-stored-secret"))
print(tokens_match("same-token", "same-token"))go deeper
Take away the headline: the fixed-time comparison protects the contents of the two values, not their sizes. If the two sides can differ in length, hash both to the same width before comparing them.
State the caveat in the documented terms — types and lengths may leak, values never do — and show the fix as code: a fixed-size digest of each side, then one call to the constant-time comparison.
Show the whole verification path: lookup by a non-secret identifier, a dummy stored value so the miss path costs the same, fixed-width normalisation, one comparison, one generic rejection, rate limiting in front. Say when the hashing step is unnecessary.
Own the tradeoff between a fast digest and a slow key-derivation function, and the policy behind it: which credentials are machine-generated with real entropy, which are human-chosen, and what the migration off the second class looks like.
### The guarantee, stated precisely `hmac.compare_digest` promises that the comparison does not short-circuit on content: it will not return sooner because the third byte differs than because the thirtieth does. What it does not promise is size-independence. The loop covers a number of positions derived from the arguments it was handed, so the total time still scales with how much data there is, and the documentation says so plainly: if the two arguments have different lengths, or an error occurs, a timing attack could theoretically reveal information about the *types and lengths* of the arguments — but not about their values. That is the whole caveat, and it is worth stating in an interview in exactly those terms, because the two wrong summaries are common in both directions. 'It leaks nothing' is over-claiming. 'It leaks the position of the first difference when the lengths differ' is under-claiming in a way that misunderstands the implementation — there is no early exit for an attacker to locate. ### When the caveat is irrelevant Almost always. A message authentication tag from a given algorithm is a fixed number of bytes. A digest is a fixed number of bytes. A token minted by your own code with a fixed-width generator is a fixed number of bytes. When both sides are structurally the same size, there is no length to leak, and passing the two values straight to the comparison is complete and correct. Adding a hashing step there buys nothing and costs a reviewer's attention. ### When it is not It matters when the value on the stored side is not fixed-width. A fraud-scoring service that has accumulated four generations of integration credentials — an early 16-character key, a later 43-character URL-safe token, a couple of hand-issued values of whatever length someone typed — has a stored side whose length varies per caller. Compare directly, and the response time carries a hint about the size of the credential the caller is being checked against, which narrows an attacker's search space before they have guessed a single byte. It is a small leak. It is also entirely avoidable. ### Hash first, then compare The fix is to give the comparison two values that are always the same size: ```python import hashlib, hmac def match(presented: str, stored: str) -> bool: a = hashlib.blake2b(presented.encode('utf-8'), digest_size=32).digest() b = hashlib.blake2b(stored.encode('utf-8'), digest_size=32).digest() return hmac.compare_digest(a, b) ``` Both arguments are now 32 bytes whatever went in, so no length is observable. Two other problems fall out at the same time: text of any encoding is accepted, since the encode step happens before the comparison, and the comparison itself only ever sees bytes, so the TypeError paths are gone. The cost is one fast hash per request, which is invisible next to a network round trip. ### The limits of hashing first Two caveats a senior candidate should raise unprompted. First, a plain fast digest is the right tool only when the underlying secret has real entropy — a randomly generated token of adequate width. If the secret is anything a human chose, a fast digest of it is offline-crackable should the stored side ever leak, and the value belongs in a purpose-built slow key-derivation function instead. Second, hashing does not hide *whether a record exists*: if the code returns early when no credential is on file for a caller, the existence of the account leaks through timing and through the shape of the response, regardless of how the comparison is done. Look the record up, and when there is none, compare against a stored dummy value of the same shape so both paths do the same work and return the same rejection. ### The order to do things in The complete verification path looks like: parse the request, take the non-secret identifier, look up the stored value (falling back to a dummy), normalise both sides to fixed-size digests, run exactly one constant-time comparison, and return one generic result for every kind of failure — with rate limiting in front of the whole thing. Every step there addresses a different leak, and the constant-time comparison is only one of them. ### You cannot unit-test this One practical consequence worth raising: there is no reliable way to assert constant-timeness from a Python test. Wall-clock measurements inside a test process are dominated by scheduling, garbage collection and the interpreter's own adaptive behaviour, so a timing assertion is a flaky test that will eventually be deleted. Test the behaviour — the right value returns True, every wrong one returns False, an absent record rejects identically — and enforce the property structurally instead: one audited helper that every secret comparison goes through, and a review rule against `==` on anything the codebase treats as a credential.
- If both sides are fixed-width tags already, is the hashing step still worth adding?No. When the stored value and the presented value are structurally the same size — two tags from the same algorithm, two digests of the same width — there is no length to leak, and the extra hash adds a step a reviewer must reason about for no gain. Add it when the stored side genuinely varies in width, or when it doubles as your normalisation point for untrusted text.
- What leaks when there is no stored credential for the caller at all?The existence of the account. Returning early on a missing record makes the miss path measurably cheaper than the compare path and often differently shaped in the response. Look the record up, substitute a stored dummy value of the same width when it is absent, run the same comparison anyway, and return the identical rejection either way.
- Why not just hash the secret with a slow key-derivation function every time?Because the cost lands on you, not the attacker, on every single request. For a high-entropy random token a fast fixed-size digest is enough — there is nothing to crack offline. A slow derivation function is the right answer when the secret is human-chosen and therefore guessable, and there the per-request cost is the price of that weakness.
saying these in an interview costs you the question
- Claims the comparison hides everything including length
- Thinks a length mismatch causes an early return
- Hashes fixed-width tags for no reason and calls it hardening
- Returns early when no credential exists for the caller
- Uses a fast digest to store human-chosen secrets
- Treats constant-time comparison as the whole defence