After deriving a PBKDF2 hash, what do you store, and how do you verify a password on login?
answer
- Store salt + iterations + algo + hash, never the password
- Verify = recompute with stored params, then compare
- MessageDigest.isEqual for constant-time compare
- Self-describing record: algo$iters$salt$hash
- Rehash on successful login to upgrade work factor
basics
~20 sStore the salt, iteration count, algorithm, and the hash bytes (not the password). On login, re-derive the hash from the entered password using the stored salt and iterations, then compare it to the stored hash with a constant-time check.
solid answer
~40 sPBKDF2 is deterministic: same password + same salt + same iterations + same algorithm yields the same bytes. So verification means recomputing. You persist a self-describing record - typically `algo$iterations$base64(salt)$base64(hash)` - which holds every parameter except the password. On login you parse that record, build a PBEKeySpec with the entered password and the stored salt/iterations, derive the candidate hash, and compare it to the stored one using MessageDigest.isEqual (constant-time) to avoid timing leaks. Because the iteration count travels with the record, you can raise it later: when a user with an old, weak record logs in successfully, re-hash with the new parameters and overwrite (transparent upgrade). Never store or log the plaintext, and wipe the char[] after use.
go deeper
Understands the hash is one-way and that login means recomputing and comparing, not decrypting.
Stores salt + iterations + algo + hash, re-derives correctly on login, and knows not to store plaintext.
Uses MessageDigest.isEqual for constant-time compare, designs a self-describing record, and implements rehash-on-login to raise the work factor.
Defines org-wide hash-format versioning and migration policy, plans the work-factor uplift cadence, and accounts for compliance/audit and zero-downtime rotation across services.
## The core idea: hashing is one-way, verification is recompute-and-compare A password hash is **one-way** - you can't reverse it to get the password. So how do you check a login? You don't decrypt anything; you **recompute the hash of the entered password using the exact same inputs** and check whether the result matches what you stored. For that to work, every input PBKDF2 used must be recoverable at login time - except the password, which the user supplies. ## What to store (and why each piece) PBKDF2's output is fully determined by four things: the **password**, the **salt**, the **iteration count**, and the **algorithm/PRF** (plus key length). You store all of them except the password: - **salt** - needed because each user has a different random salt; without it you can't reproduce the hash. It's not secret. - **iteration count** - needed to reproduce the hash, and crucially it lets you *change* the work factor over time without invalidating existing hashes. - **algorithm + key length** - so a future code change (e.g. SHA-256 -> SHA-512) doesn't silently break old records; the record says how it was made. - **the derived hash bytes** - the value you'll compare against. A conventional, self-describing encoding packs these into one column, e.g. `pbkdf2-sha256$600000$<base64 salt>$<base64 hash>`. This pattern (sometimes called a Modular Crypt Format-style string) means the database row carries its own recipe. ## The verification sequence 1. Look up the stored record for the username. 2. Parse out algorithm, iterations, salt, and stored hash. 3. Build `new PBEKeySpec(enteredPassword, salt, iterations, keyLengthBits)` and derive the candidate hash with the matching `SecretKeyFactory`. 4. Compare candidate vs stored with a **constant-time** comparison. ## Constant-time comparison - what and why A naive byte-by-byte comparison (`Arrays.equals`) returns as soon as it finds the first differing byte, so it takes *less time* for a wrong guess that differs early than for one that matches a long prefix. An attacker could, in principle, measure these tiny timing differences to learn the hash byte-by-byte (a **timing side-channel**). `MessageDigest.isEqual(a, b)` is designed to take time independent of where the first difference is, closing that channel. (Deeper rationale lives in the password-storage foundations topic; here the point is: use `MessageDigest.isEqual`, not `==`/`equals`.) ## Transparent upgrade (rehash on login) Because the parameters travel with each record, you can strengthen security gradually. Set a current target (algorithm + iterations). When a user logs in successfully against an *old* record, immediately re-derive a hash with the current parameters and overwrite the stored record. Over time, active accounts migrate to the strong settings with zero user friction. ## Don't-do list - Don't store or log the plaintext password (or the derived key in logs). - Don't reuse one global salt - it re-enables rainbow tables and reveals equal passwords. - Don't compare with `==` or `Arrays.equals`. - Don't forget `clearPassword()` / zeroing the char[].
- Why use MessageDigest.isEqual instead of Arrays.equals when comparing hashes?Arrays.equals short-circuits on the first differing byte, so its runtime depends on how many leading bytes match - a timing side-channel. MessageDigest.isEqual compares in (near) constant time regardless of where bytes differ, so an attacker can't infer the hash from timing.
- How can you increase the iteration count later without forcing every user to reset their password?Store the iteration count with each hash. On a successful login against an old record, re-derive the hash using the new, higher iteration count and overwrite the stored value. Active users transparently migrate to stronger parameters over time.
saying these in an interview costs you the question
- Trying to 'decrypt' the hash to compare (it's one-way)
- Comparing with Arrays.equals or == instead of MessageDigest.isEqual
- Not storing the iteration count, so you can never raise it
- Storing the password alongside the hash 'just in case'
- Using a hardcoded salt shared by all users