MD5 and SHA-1 are routinely called "broken". State precisely which property fails, explain why a system that compares an attacker-supplied artifact against an attacker-influenced digest fails immediately while a keyed construction over the same function does not, and say what you would tell a team that answers "we only use MD5 as a cache key".
answer
- collision gone, preimage intact
- chosen-prefix = two meaningful files collide
- adversary picks both inputs → exposed
- keyed construction ≠ collision-dependent
- store sha256: prefix — agility, not bare hex
basics
~20 sTheir collision resistance is destroyed; their preimage resistance is essentially intact. So uses where the adversary chooses both inputs — signing, deduplication, approval lists — fail now, while uses that require matching a fixed unknown value do not. Brokenness is a statement about a property, never about an algorithm in the abstract.
solid answer
~50 sMD5 collisions are computable in seconds and chosen-prefix collisions — where the attacker fixes two meaningful prefixes and appends colliding blocks — are practical; SHA-1 fell the same way, identical-prefix in 2017 and chosen-prefix in 2019. Preimage resistance is a different story: the best MD5 preimage attack is around 2^123, no practical improvement. So the correct question is which property a use depends on. Anything where the adversary supplies both inputs — certificate or document signing, dedup, content addressing, "this binary is approved because its digest is listed" — is broken today. A keyed construction such as HMAC-MD5 does not rest on collision resistance of the raw function in the same way and has no practical break, which is an argument about *urgency*, not about *keeping it*. "Only a cache key" is fine until the key carries a security decision: a collision then means one tenant's entry serves another's request. Store an algorithm identifier next to every digest so the choice is revisable.
go deeper
Say that MD5 and SHA-1 have practical collisions and must not be used for signatures or integrity decisions, and name SHA-256 as the default.
Separate collision from preimage, cite the birthday bound, and identify which uses in a codebase depend on which property.
Drive the exposure assessment: enumerate the places where an attacker supplies both inputs, rank them, and give a migration path that keeps verification working during the transition.
Talk about crypto agility as a design property — algorithm identifiers stored with digests, dual-verify/single-write rollout, a dated removal, and a written rule for when a non-cryptographic hash is acceptable.
## What actually broke Collision resistance. Nothing else, and that distinction is the whole answer. For MD5, colliding pairs can be generated in well under a second on ordinary hardware. More importantly, **chosen-prefix** collisions are practical: the attacker picks two arbitrary, meaningful prefixes — say, a benign document and a malicious one — and computes appendages that make the two full messages hash to the same value. That is what turns a mathematical curiosity into an exploit, because the colliding pair no longer has to be two blobs of noise; it can be two files that each parse and each mean something. SHA-1 followed the same path: an identical-prefix collision was demonstrated in 2017 with two PDFs, and a chosen-prefix collision in 2019 at roughly 2^63 work — expensive but purchasable. Meanwhile preimage resistance survived. The best published MD5 preimage attack sits around 2^123, barely below the generic 2^128. Nobody can take an MD5 digest and produce an input for it. This is why "MD5 is broken" and "MD5 reveals its input" are different claims and only the first is true. ## Mapping the break onto uses Use the invariant: how much of the input does the adversary control? **Broken now — adversary controls both inputs.** - Certificate and code signing. The signature covers a digest; a colliding pair means a signature over the benign artifact is a valid signature over the malicious one. This is not hypothetical; it is how MD5-based certificate forgery worked. - Deduplication and content-addressed storage. Upload artifact A, then upload malicious B that collides; the store believes it already has that content and serves the wrong bytes, potentially across tenants. - Allow/deny lists keyed by digest. Get the benign twin approved, ship the evil twin. - Any "same digest means same object" trust decision, including cache keys once they influence authorization or content selection. **Not immediately broken — the adversary must hit a value they did not choose.** - Verifying that a file you already hold matches a digest you already hold, against non-adversarial corruption. That needs second-preimage resistance, still intact. - HMAC-MD5. HMAC's security proof leans on the compression function behaving as a pseudorandom function under a secret key rather than on raw collision resistance, and there is no practical forgery. It should still be retired: it is a liability you are choosing to defend rather than remove, and every reviewer will flag it forever. ## The cache-key answer "It is only a cache key" is a legitimate category — non-cryptographic hashing for distribution and lookup is a real, correct use, and MD5 is merely a slow choice for it. Two things convert it into a vulnerability: 1. **The key becomes a decision.** If the cache stores rendered responses keyed by a digest of the request, a collision serves one requester's content to another. If it deduplicates uploads, a collision is a data-integrity break. If it memoizes an authorization result, a collision is an authorization bypass. 2. **Someone reads the code later.** A digest in a database column is untyped: an engineer three years on cannot tell whether it is load-bearing. This is the real argument for uniform migration — not that today's use is exploitable, but that the property boundary is invisible in the code and will be crossed by accident. So the answer is: it is defensible today, it is not defensible as a standard, and if you keep it you must be able to point at the reason the collision property is not in play, in writing, next to the code. ## Crypto agility The durable lesson is that hash choices outlive the reasoning behind them. Practices that make the next migration cheap: - **Store the algorithm with the digest** — a prefixed identifier such as `sha256:...` rather than a bare hex string. Bare digests are the reason migrations stall. - **Support verifying against multiple algorithms while writing only the new one**, then backfill and drop the old verifier on a date. - **Size the new digest by the property you need** — collision security is half the digest length, so a 256-bit digest gives 128-bit collision security. - **Do not treat SHA-3 as an upgrade over SHA-2.** SHA-2 has no practical collision attack; SHA-3 exists as a structurally different backup (a sponge, not a Merkle–Damgård chain), not because SHA-2 is weak. ## How to answer Lead with the property, not the algorithm. "Collision resistance is gone; preimage resistance is not. So the exposure is exactly the set of places where the attacker picks both inputs — here is that set in our system." That answer survives the follow-up; "MD5 is insecure, use SHA-256" does not.
- What is the difference between an identical-prefix and a chosen-prefix collision, and why does the distinction matter for exploitability?An identical-prefix collision requires both messages to start with the same bytes, with the colliding difference confined to attacker-generated blocks; it proves the function is broken but constrains what the two files can be. A chosen-prefix collision lets the attacker fix two entirely different meaningful prefixes — a real certificate subject and a forged one, a benign installer and a malicious one — and compute appendages that reconcile them. Only the second reliably turns into a forged certificate or a swapped artifact, which is why the chosen-prefix result is the one that ends an algorithm's usable life.
- Does SHA-256 need replacing, and is SHA-3 the replacement?No practical collision or preimage attack exists against SHA-256, so it does not need replacing on cryptanalytic grounds. SHA-3 was standardised as structural insurance: it is a sponge construction rather than a Merkle–Damgård chain, so an advance against the MD design would not affect it, and it is naturally immune to length extension. Choosing between them is a portfolio-diversification and interoperability decision, not a strength comparison.
saying these in an interview costs you the question
- "MD5 is broken so it leaks the original value" — preimage resistance is intact; the break is collisions.
- Declaring an algorithm safe or unsafe without naming the property the use depends on.
- Assuming a keyed construction inherits the raw function's collision break, or conversely that it is fine to keep indefinitely.
- "It's just a cache key" without checking whether the key drives a trust, dedup or authorization decision.
- Migrating to SHA-3 on the belief that SHA-2 is weak, or truncating the new digest and re-creating a low collision bound.