Which argument types does `hmac.compare_digest` accept, and when does it raise TypeError?
answer
- Two shapes are legal, nothing else
- Text is allowed, but only narrow text
- A wire value is untrusted text
- Non-ASCII or mixed types raise TypeError
- Encode at the boundary with str.encode
basics
~20 sBoth arguments must be the same flavour: two bytes-like objects such as bytes, bytearray or memoryview, or two str values containing only ASCII. Mixing a str with bytes, or passing a str holding a non-ASCII character, raises TypeError.
solid answer
~40 s`hmac.compare_digest` works over a flat run of bytes, so it accepts either two bytes-like objects (`bytes`, `bytearray`, `memoryview`) or two `str` values restricted to ASCII — the `str` case exists for hex-digest strings. A `str` containing a non-ASCII character raises TypeError, because CPython stores such a string with a wider internal representation and refuses to pretend a comparison over it is constant time. Mixing a `str` with a bytes-like object also raises TypeError. This matters at a request boundary: a signature or token arriving in a header or JSON body is a `str` an attacker controls, so if it goes straight into the comparison, a crafted non-ASCII value turns a clean authentication failure into an unhandled exception. Encode to bytes at the boundary and compare bytes.
code
python · 9 linesimport hmac
print(hmac.compare_digest("9f86d081884c7d65", "9f86d081884c7d65"))
for a, b in [("abc", b"abc"), ("caf\u00e9", "caf\u00e9")]:
try:
hmac.compare_digest(a, b)
except TypeError as exc:
print("TypeError:", exc)go deeper
Remember the two legal shapes — two bytes-like objects, or two ASCII-only strings — and that anything else raises TypeError rather than returning False. Encode text with str.encode('utf-8') before comparing.
Explain the ASCII restriction from CPython's variable-width string storage, and name both TypeError cases: mixed str/bytes, and a str carrying a non-ASCII character.
Frame it as an availability bug on the authentication path: attacker-chosen header text reaching the comparison turns a 401 into an unhandled exception, skips the failed-attempt counter, and floods the logs. Normalise at the edge.
Own the boundary contract — where untrusted text becomes bytes, which helper every secret comparison goes through, and the rule that all authentication failures are indistinguishable in status, body and, as far as practical, latency.
### The accepted shapes `hmac.compare_digest(a, b)` compares two sequences of bytes in fixed time. Two argument shapes are legal: * **Two bytes-like objects.** `bytes` is the usual one, but anything exposing a buffer works — `bytearray`, and `memoryview`, which lets you compare a slice of a larger buffer without copying it. * **Two `str` values containing only ASCII.** This case exists so that hex digests and URL-safe tokens, which are already text, can be compared without a manual encode. Anything else is a TypeError: a `str` on one side and bytes on the other, a `str` holding a character outside ASCII, or an object with no buffer at all such as `None` or an `int`. ### Why ASCII, specifically CPython stores a `str` compactly, choosing one, two or four bytes per character according to the widest code point present. Two strings that look similar can therefore have completely different in-memory widths, and comparing them as raw memory is either wrong or not uniform in time. Rather than silently do something whose timing properties it cannot promise, the implementation restricts the `str` path to the case where the representation is guaranteed to be one byte per character, and raises TypeError otherwise. The restriction is a feature: it refuses to give you a comparison whose constant-time claim it cannot honour. ### Why this is an availability bug, not a style nit Consider a fraud-scoring service whose scoring callbacks arrive over HTTP with an authentication tag in a header. The header value reaches the handler as a `str`, and every byte of it is attacker-chosen. If that `str` goes directly into `hmac.compare_digest` against a stored hex digest, then a request carrying a single non-ASCII character in that header does not produce a 401 — it produces an unhandled TypeError. Concretely that means an error response rather than a clean rejection, a stack trace in the logs for every such request, a skipped failed-authentication counter (so the rate limiter and the alerting never see the attempt), and a trivially scriptable way to fill the log pipeline. A 4-person team debugging this at 3 a.m. sees a burst of 500s from an endpoint whose authentication looked correct in review. ### The boundary rule Normalise once, at the edge, and compare bytes from there inward: ```python presented = request_header.encode('utf-8') # str -> bytes, never raises for text ``` `str.encode` with UTF-8 accepts any string, so the TypeError path disappears entirely and the comparison sees two bytes objects. If the two sides can also differ in length — which the comparison does not hide — run both through a fixed-size digest before comparing, which solves the encoding and the length question in one step. ### Getting the failure mode right Whichever normalisation you choose, the handler must still return the same generic rejection for every failure: a malformed tag, a wrong tag, an unknown key id and a missing header should be indistinguishable to the caller in status code, body and — as far as is practical — latency. Wrapping the comparison in a `try`/`except TypeError` that returns the same rejection is acceptable as a belt-and-braces measure, but it is a backstop, not the fix: the real fix is that untyped text from the network never reaches the comparison in the first place. ### What to say in an interview The two accepted shapes, the ASCII restriction and the reason behind it, the two TypeError cases, and the observation that the risky argument is always the one that came off the wire. Candidates who have only ever seen the function in a tutorial know that it compares digests; candidates who have shipped it know that the input needs encoding first and that the failure is an exception, not a `False`. ### The third TypeError, and the one people actually hit Mixed types and non-ASCII text are the documented cases; the one that reaches production most often is `None`. A missing header read with a mapping's `get` returns `None`, and `None` has no buffer, so it raises TypeError just as surely as a mismatched pair does. That means the *absent* credential and the *wrong* credential take different code paths — one an exception, the other a clean `False` — which is both an availability bug and, because the two are distinguishable from outside, a small information leak. Check for a missing value explicitly and route it into the same generic rejection as a wrong one.
- Why is a `memoryview` accepted, and when is that useful?The comparison reads a buffer, and `memoryview` exposes one without copying. That lets you compare a tag embedded inside a larger frame — a slice of a received message — without slicing out a fresh `bytes` object, which avoids leaving another copy of the secret material on the heap for the garbage collector to move around.
- Is catching TypeError around the comparison a sufficient fix?It stops the 500, so it is a reasonable backstop, but it is not the fix. It leaves untrusted text flowing into a function that has type preconditions, and it is easy for the next caller to forget. Normalise at the boundary with `str.encode`, or hash both sides to fixed-size digests, so the comparison only ever sees two bytes objects of equal length.
saying these in an interview costs you the question
- Thinks compare_digest silently encodes a str for you
- Assumes a bad argument returns False instead of raising
- Passes a raw header string straight into the comparison
- Believes any str is accepted, not just ASCII
- Cannot say why the ASCII restriction exists