A service authenticates messages by transmitting the message together with SHA-256(secret || message). Explain the structural weakness in that construction, and state the general rule for when a bare hash actually gives you integrity.
answer
- digest IS the internal state → resume and append
- immune: sponge, BLAKE2/3, SHA-512/256, HMAC
- unkeyed = anyone can recompute = no authenticity
- integrity only vs an authenticated reference digest
- MAC gives authenticity, not freshness — add a nonce
basics
~20 sSHA-256 is a Merkle–Damgård chain whose output is its full internal state, so an attacker who knows the digest and the secret's length can append data and compute a valid digest without the secret. More fundamentally, a bare hash gives integrity only against a digest obtained over a channel the attacker cannot rewrite.
solid answer
~50 sTwo separate failures. **Structural:** MD5, SHA-1, SHA-256 and SHA-512 are Merkle–Damgård constructions that emit their entire internal chaining state as the digest. Given `H(secret || m)` and the length of `secret`, an attacker resumes the computation from that state and produces `H(secret || m || padding || suffix)` — a valid tag for a message they extended, with no knowledge of the key. Sponge constructions (SHA-3), BLAKE2/BLAKE3, and truncated variants such as SHA-512/256 withhold state and are immune; HMAC's nested construction is immune regardless of the inner hash. **Conceptual:** an unkeyed hash supplies no authenticity at all — anyone can recompute it. It gives integrity only *relative to a reference digest the attacker cannot alter*: a digest published on the same mirror as the artifact detects accidental corruption and nothing else. Authenticity needs a key (a MAC) or a signature; freshness needs a nonce or timestamp on top.
code
text · 6 linest0 server: tag = H(K || "amount=10") sends msg + tag
t1 attacker: state <- tag (no key needed)
t2 attacker: absorb(glue_padding + "&amount=1000000")
t3 attacker: tag' = finalize(state)
t4 server: H(K || "amount=10"+glue+"&amount=1000000") == tag' ✓
parser takes last-wins ⇒ amount = 1000000go deeper
Say a bare hash is unkeyed so anyone can recompute it, and that authenticating a message requires a keyed construction rather than a hash of secret plus message.
Explain length extension mechanically — the digest is the internal state — and name HMAC as the correct construction and which hash families are immune.
Generalise to the property table: which of confidentiality, integrity, authenticity and freshness each primitive supplies, and where the reference digest's own trust comes from in a real distribution pipeline.
Frame it as trust-anchor design: identify what ultimately supplies authenticity in each artifact flow, ensure digests are only ever an efficiency layer beneath a signature or key, and rule out home-grown keyed constructions by policy.
## The construction and why it leaks MD5, SHA-1, SHA-256 and SHA-512 all use the Merkle–Damgård design: the message is padded to a whole number of blocks, an internal state is initialised to a fixed value, and each block is mixed into that state by a compression function. When the blocks run out, **the internal state is the digest**. That last sentence is the vulnerability. The digest is not a summary of the state; it *is* the state. An attacker who sees `t = H(secret || m)` can load `t` as the starting state of the same function, feed in any suffix, and obtain the correct digest for a longer message — specifically for `secret || m || glue || suffix`, where `glue` is the padding the original computation appended (length-encoded, so the attacker must guess or brute-force the secret's length; the search space is tiny). At no point do they learn or need the secret. Whether that is exploitable depends on how the receiver parses the extended message. In practice, formats that ignore trailing junk or take last-wins semantics — trailing query parameters, appended key/value pairs, permissive serializers — make it directly exploitable: append `&role=admin` and the tag still verifies. ``` text known: tag = H(secret || "user=alice") attacker computes, without the secret: tag' = resume(tag) over ("padding-of-original" + "&role=admin") sends: "user=alice" + padding + "&role=admin" , tag' receiver recomputes H(secret || received) == tag' ✓ verifies ``` ## Which functions are immune, and why - **Sponge constructions (SHA-3/Keccak).** The internal state is larger than the output; the unexposed "capacity" portion cannot be reconstructed from the digest, so the chain cannot be resumed. - **Truncated Merkle–Damgård (SHA-512/256, SHA-384).** The withheld bits play the same role: the attacker lacks part of the state. - **BLAKE2 / BLAKE3.** Designed with the finalisation flag that makes extension impossible. - **HMAC over any of them.** HMAC hashes twice with two derived keys, so the outer hash's input is the inner digest and the exposed state is not the tag over the message. A candidate who says "switch to SHA-3 and the construction is fine" has spotted the mechanism but missed the design point: you would be relying on an accidental property of a chosen algorithm to make a home-made scheme safe. Use a construction whose security is stated for the purpose — HMAC, or a purpose-built keyed hash. ## The general rule about hashes and integrity Separate the four properties an interviewer expects you to name: - **Confidentiality** — nobody else can read it. Hashes give none. - **Integrity** — the data has not changed since the reference was taken. - **Authenticity** — the data came from the party holding the key. - **Freshness** — this is not a replay of an older, legitimately produced message. A bare hash is unkeyed, so anyone — including the attacker — can compute it. It therefore provides authenticity never, and integrity **only relative to a digest the attacker could not have rewritten**. That is the rule worth memorising: > A hash moves the integrity problem from the artifact to the digest. It is a control only if the digest travels a path the attacker cannot influence. Apply it across stacks and the pattern is identical: - A checksum on the same download page as the tarball: mirrors the artifact's own trust, so it detects a truncated download and nothing adversarial. - A package repository: digests live in an index that is *signed*, so the trust ultimately rests on a signature, and the digests are an efficiency mechanism that lets you verify big artifacts against a small signed statement. - TLS certificates: the digest of the certificate body is signed by a CA key; the hash is the compression step inside an authenticity mechanism, not the mechanism. - Content-addressed storage and version-control object graphs: hash chaining gives you tamper-*evidence* relative to a root you obtained elsewhere; if the attacker controls the root, the chain proves nothing. In every case the hash is doing the same job — shrinking a large artifact to a small commitment — and the authenticity is supplied by a key somewhere else. ## Freshness is separate again Even a correct MAC does not stop replay: a captured message with its valid tag stays valid forever. Freshness requires a nonce, sequence number or timestamp *inside* the authenticated data. Candidates often stop at "use HMAC" and are surprised by the replay follow-up. ## How to answer Name the Merkle–Damgård state exposure, name the immune constructions, then step up a level: the real error was building an authenticity mechanism out of a primitive that offers none, and the general rule is that a hash gives integrity only against an authenticated reference value.
- Would switching from SHA-256 to SHA-3 make `H(secret || message)` a sound MAC?It removes the length-extension attack, because a sponge withholds part of its state, but it does not make the scheme sound. You would be depending on an incidental structural property rather than on a construction with a stated security claim for keyed authentication, and the scheme still offers no analysis against related-key or key-recovery issues. Use HMAC or a dedicated keyed hash whose security is proved for this purpose.
- Publishing a SHA-256 checksum next to a download link — what does it actually protect against?Accidental corruption: a truncated transfer, a bad mirror disk, a proxy that mangled the bytes. It does not protect against an attacker who can modify the download, because the same attacker can modify the checksum published beside it. It becomes a real control only when the digest arrives over a path with independent authenticity — a signed release index, a signature over the digest, or a value pinned in a build configuration reviewed in source control.
saying these in an interview costs you the question
- "It has a secret in it, so it's a MAC" — the secret's position matters; prefixing leaks to length extension.
- Believing length extension recovers the secret; it does not, it forges an extended message.
- "Anyone can verify the hash, so integrity is covered" — verifiable by anyone means forgeable by anyone.
- Treating a checksum published beside the artifact as a security control rather than a corruption check.
- Stopping at "use HMAC" and missing that replay still needs a nonce or timestamp.