skip to content

questions

4

A team stores user passwords as SHA-256 digests and argues that SHA-256 is cryptographically strong and irreversible. Explain why that reasoning is wrong for passwords, and what property a password-storage function must have instead.

level: juniorimportance: must knowfreq 82%

answer

  1. attacker guesses candidates, never inverts
  2. fast hash = billions of guesses/sec on a GPU
  3. salt kills precomputation + cross-user amortisation
  4. salt is unique and random, stored in the clear
  5. cost parameter stored with the record so it can rise

basics

~20 s

Passwords are low-entropy, so nobody inverts the hash — they guess candidates and hash them. A general-purpose hash is designed to be fast, so an attacker tests billions of guesses per second. A password function must be deliberately slow and hardware-hostile, with a per-user salt and a tunable cost.

solid answer

~50 s

SHA-256 is strong against the attacks it was designed for: preimage and collision resistance against *arbitrary* inputs. Passwords are not arbitrary — the human-chosen space is small and skewed, so an attacker never inverts anything. They run the candidate list through the same function and compare. The property that matters is therefore **cost per guess**, and general-purpose hashes are optimised for the opposite: throughput. On commodity GPUs a fast hash yields billions of candidates per second; a purpose-built password function yields thousands. So the requirements are: a **unique random salt per password**, so precomputed tables are useless and one cracking run cannot be amortised across all users; a **tunable work factor**, so the defender can raise cost as hardware improves without changing algorithm; and ideally **memory-hardness**, so parallel hardware loses its advantage. That is what bcrypt, scrypt, Argon2 and PBKDF2 provide and what SHA-256 by construction does not.

go deeper

for a junior

Core recall: passwords are guessed, not reversed; use a slow purpose-built function with a unique random salt per user, never a bare fast hash.

for a middle

Explain the two distinct jobs — salt removes precomputation and cross-user amortisation, work factor raises per-guess cost — and why the parameters are stored with the record.

for a senior

Quantify the asymmetry (defender pays once per login, attacker pays once per guess), note the memory-hardness argument against parallel hardware, and separate offline from online threat models.

for a principal

Frame it as a cost-of-attack budget that must be reviewed as hardware moves, with an upgrade path built into the record format and layered controls — breached-password screening, second factors — because storage strength alone cannot rescue weak human choices.

## The attack is guessing, not inversion A cryptographic hash maps input to a fixed-size digest and is designed so that, given a digest, finding *any* input that produces it is infeasible. That guarantee is about the space of all possible inputs. Passwords occupy a tiny, heavily skewed corner of that space: real users choose from a distribution dominated by common words, names, dates, keyboard patterns and predictable mutations. An attacker holding a stolen digest therefore never attempts inversion. They generate candidates — from leaked password corpora, from dictionaries with mutation rules, from targeted personal data — hash each one and compare. The mathematics of preimage resistance is untouched and irrelevant. This reframes the entire problem. The only defensive lever is the **cost of testing one candidate**, multiplied by the number of candidates the attacker must test. You cannot raise the second term much — users choose the passwords they choose — so you raise the first. ## Speed is a feature everywhere else and a defect here General-purpose hashes are engineered for throughput: they are used to checksum files, build Merkle trees and sign large documents, so implementers optimise them hard, and hardware vendors add dedicated instructions. The result is that a single commodity graphics card evaluates a fast hash on the order of billions of candidates per second, and rented fleets multiply that. Against that rate, an entire realistic dictionary of human passwords is exhausted in a time measured in minutes to hours. A password function inverts the design goal. It is built so that one evaluation costs a configurable amount of work — typically tuned so the defender spends on the order of a tenth of a second per login. The defender pays that cost once per authentication; the attacker pays it once per guess, and the attacker makes billions of guesses. The asymmetry is the whole design. ## Why the salt is not optional and not a secret A **salt** is a unique, random value generated per password and stored alongside the digest. It is not secret; assume the attacker has it the moment they have the digest. It defeats two things. *Precomputation.* Without a salt, an attacker can compute a table of digests for a candidate list once — in the general case a time/memory trade-off structure such as a rainbow table — and then reverse any number of stolen digests by lookup. Salting means the table would have to be recomputed per salt, which destroys the economics of precomputation entirely. *Amortisation across users.* Without a salt, one hash evaluation of the candidate "summer2024" tests it against every user in the dump simultaneously, so cracking ten million accounts costs the same as cracking one. With unique salts, each user must be attacked separately, restoring a linear relationship between accounts and attacker cost. It also stops the trivially observable fact that two users share a password, which is itself a leak. The salt must be **random and unique per password**, from a cryptographically secure generator, and long enough that collisions do not occur in practice. Deriving it from the username or the user id is a classic mistake: usernames repeat across sites, so cross-site precomputation returns, and a rename changes the salt. ## The work factor and why it must be adjustable Hardware gets faster; a fixed cost silently erodes. A password function therefore exposes a **cost parameter** — iteration count, or a cost exponent, or explicit time, memory and parallelism settings — and stores it in the record next to the salt. That single design choice enables two things: the parameters can be raised over time, and each stored value carries the parameters it was created with, so old and new records verify correctly side by side and can be upgraded opportunistically when the user next authenticates. ## What the stored record looks like conceptually A password record is not just a digest. It carries an algorithm identifier, the parameters used, the salt, and the derived output — commonly encoded as a single self-describing string. That structure is what makes migration and parameter increases possible at all; a bare digest column has thrown away the information needed to evolve. ## The two boundary claims First, a password function protects only the *offline* attack that follows a database compromise. It does nothing about guessing against your live login endpoint, which is a rate-limiting and detection problem, nor about credential reuse from other breaches, which needs breached-password screening and a second factor. Second, adding iterations to a fast hash by hand — hashing repeatedly in a loop — is not equivalent to using a purpose-built function. It raises cost linearly while leaving the operation entirely parallel-friendly and cheap in memory, which is exactly the property that modern password functions target. Home-rolled schemes also routinely lose the salt handling and the parameter encoding along the way.

  • If the salt is stored in the database right next to the hash, how does it help once the database is stolen?
    It was never meant to be secret. Its job is to make each stored password a separate cracking problem: precomputed tables become worthless because they would have to be rebuilt per salt, and a single hash evaluation can no longer be tested against every account at once. Attacker cost therefore scales with the number of accounts instead of being amortised across them.
  • Is hashing the password ten thousand times with SHA-256 in a loop an acceptable substitute?
    It is better than a single pass and it is roughly what PBKDF2 formalises, but rolling it by hand is a poor idea: the construction stays cheap in memory and perfectly parallel, so specialised hardware retains its advantage, and hand-rolled versions routinely mishandle the salt, the parameter encoding and the output length. Use a standard construction with an explicit cost parameter instead.
  • Does slow hashing protect against someone guessing passwords at your login form?
    No. It only raises the cost of the offline attack that follows a database compromise. Online guessing is bounded by rate limiting, lockout or progressive delays keyed to the account and the source, plus anomaly detection and a second factor. In fact an expensive hash makes the login endpoint a resource-exhaustion target, so the rate limit protects the server as well as the account.

A vault door rated against drilling is useless if the lock accepts a thousand key guesses a second. Password storage does not make the door thicker; it makes each attempt take a measurable amount of time.

saying these in an interview costs you the question

  • "SHA-256 is irreversible, so it is fine" — confusing preimage resistance with guessing cost
  • Treating the salt as a secret, or deriving it from the username or user id
  • Using one global salt for all users, which restores cross-user amortisation
  • Believing encryption of passwords is an upgrade over hashing — it introduces a recoverable key
  • Assuming a slow hash also defends the live login endpoint against guessing

context

open as a page

In password storage, a per-user salt, a secret "pepper", and the work factor of the key-derivation function are three separate mechanisms. Explain what attack each one defeats, which of them still helps after a full database dump, and how the pepper must be constructed if you ever want to rotate the secret.

level: middleimportance: must knowfreq 62%

basics

~20 s

Salt is unique and public: it kills precomputed tables and stops one guess testing every account. Work factor makes each guess expensive — the only thing that still helps once the attacker holds everything. Pepper is a secret key stored outside the database; it defeats database-only theft, and it is rotatable only if applied as a reversible keyed layer.

open as a page

Compare bcrypt, scrypt, Argon2 and PBKDF2 as password-storage functions. On what axes do they actually differ, and how would you choose and parameterise one for a service handling heavy login traffic?

level: seniorimportance: should knowfreq 52%

basics

~20 s

They differ in what resource they force the attacker to spend. PBKDF2 costs only iterations of a fast hash, so specialised hardware wins; bcrypt costs a modest fixed memory footprint; scrypt and Argon2 let you demand large tunable memory, which is what neutralises massively parallel hardware. Parameters come from a measured latency budget at peak concurrency.

open as a page

Verifying a submitted password against a stored value looks trivial, yet several failures cluster there — timing side channels, account enumeration, denial of service, and inconsistent handling of the submitted string. Walk through what a correct verification path must do.

level: seniorimportance: should knowfreq 46%

basics

~20 s

Compare derived values in constant time; do the same amount of work for an unknown user as for a known one, using a dummy record, and return an identical response; normalise and length-bound the submitted string consistently; rate-limit because each verification is expensive; and never log or echo the credential.

open as a page