skip to content

How do you store a pbkdf2_hmac password record so the iteration count can be raised later?

level: seniorimportance: should knowfreq 44%

answer

  1. The row must describe itself
  2. Parameters cannot live only in config
  3. One-way means no offline migration
  4. Only one moment holds the plaintext
  5. Re-derive inside the successful-login branch

basics

~20 s

Store the algorithm name, iteration count, salt and derived key together in one field. Verify with the record's own parameters, and when they fall below current policy re-derive at the new count during a successful login and rewrite the row.

solid answer

~40 s

The stored value must be self-describing: something like `pbkdf2_sha256$600000$<salt hex>$<key hex>`, where the salt is 16 bytes from `os.urandom` generated when the password was set. Verification parses the record and calls `hashlib.pbkdf2_hmac` with the parameters *that record* carries — never today's constant — then compares with a constant-time comparison. Raising the work factor is only possible at one moment: inside the successful-login branch, where the plaintext password is briefly in hand. There you re-derive with a fresh salt and the new iteration count and overwrite the row. You cannot migrate offline, because the derivation is one-way and the plaintext is not stored. Accounts that never log in stay on the old parameters, so pair the lazy upgrade with a deadline that forces a reset.

code

python · 25 lines
python
import hashlib
import hmac
import os

ITERATIONS = 600_000


def store(password: str) -> str:
    salt = os.urandom(16)
    key = hashlib.pbkdf2_hmac("sha256", password.encode("utf-8"), salt, ITERATIONS)
    return f"pbkdf2_sha256${ITERATIONS}${salt.hex()}${key.hex()}"


def verify(password: str, record: str) -> tuple[bool, str | None]:
    _, iterations, salt_hex, want = record.split("$")
    key = hashlib.pbkdf2_hmac(
        "sha256", password.encode("utf-8"), bytes.fromhex(salt_hex), int(iterations)
    )
    if not hmac.compare_digest(key, bytes.fromhex(want)):
        return False, None
    upgraded = store(password) if int(iterations) < ITERATIONS else None
    return True, upgraded


print(verify("hunter2", store("hunter2")))

go deeper

for a junior

Remember that the salt and the iteration count are stored with each password record, not hard-coded in the code, and that a password hash can never be turned back into the password.

for a middle

Explain the self-describing record format and why verification must use the parameters carried by the record rather than the current constant, or every existing user is locked out on the next deployment.

for a senior

Demonstrate the lazy upgrade inside the successful-login branch, the fresh salt on rewrite, the fact that a bulk offline migration is impossible, and the failure handling that never turns a correct password into an error.

for a principal

Own the migration as a programme: measurable coverage from the per-row parameters, a deadline for dormant accounts, a story for non-interactive accounts, and a record format that lets algorithm changes ride the same path.

Consider the internal portal in front of a genome-annotation pipeline: 6,800 accounts, PBKDF2-HMAC-SHA256 records written years ago at an iteration count that was defensible then and is not now. The question is how the records were stored, because that determines whether raising the cost is a routine change or a mass password reset. ## The record is self-describing or it is stuck A password field should carry four things: 1. which **algorithm** produced it, 2. the **parameters** used, 3. the per-user **salt**, 4. and the **derived key**. A single delimited string is the conventional shape — `pbkdf2_sha256$600000$<salt hex>$<key hex>` — and any equivalent encoding works as long as nothing lives only in application configuration. The reason is simple: verification must reproduce the exact derivation that created the value. If the iteration count comes from a module-level constant, then the day someone raises that constant, every existing record silently fails to verify and every user is locked out. **The parameters belong to the row, not to the deployment.** ## The salt Generate 16 fresh bytes per user with `os.urandom(16)` at the moment the password is set, and store them in clear inside the record. It is not a secret; it exists so that one precomputed table cannot cover all 6,800 rows and so that two accounts sharing a password do not look identical in a dump. Reusing one salt for the whole table, or deriving it from the username, throws away both properties. ## Verification Parse the record, call `hashlib.pbkdf2_hmac(hash_name, password.encode("utf-8"), salt, iterations)` with the record's own values, and compare the result to the stored key with a **constant-time comparison** over the raw bytes. Encode the submitted password explicitly so the same input verifies identically everywhere. ## The upgrade window A KDF is one-way, and you do not store plaintext, so there is exactly one moment per user when a stronger record can be produced: immediately after a successful verification, while the submitted password is still in memory. The pattern is: 1. verify against the record's parameters; 2. if it matches and the record's parameters are below current policy: - derive a new key with a fresh salt and the new iteration count, - write the new record, - and continue the login. The user notices nothing but a slightly longer login. Do not skip the fresh salt — you are rewriting the record anyway, and a new salt costs nothing. ## What you cannot do - You cannot run a background job over 6,800 rows to upgrade them, because the job has no plaintext to derive from. - You also should not "upgrade" by feeding the stored key back through more iterations: it works arithmetically, but it invents a construction that is not PBKDF2, you must then record the whole chain in the record format and replay it at verification time forever, and you have built something no reviewer can check against a specification. The honest choices are **lazy upgrade on login**, or a **forced reset**. ## Coverage and the long tail Lazy upgrade is only as good as your login rate. Some fraction of a 6,800-account portal will not sign in for months, and service accounts may never sign in interactively at all. So make the migration measurable: - record the parameters per row, - count how many rows remain below policy, - and set a date after which stale records are invalidated and those users must reset. Without the deadline, the migration never finishes and nobody notices, which is how a system ends up with records at three different cost levels and no plan. ## Handling several algorithms at once Because the record names its algorithm, the same verification path can carry `pbkdf2_sha256` records and, say, `scrypt` records side by side while a migration runs. That is the real payoff of the self-describing format: changing algorithm and changing parameters become the same operation, dispatched on what the row says rather than on what the code assumes. ## A note on failure paths The re-hash belongs strictly inside the success branch, and it must not turn a successful login into an error: if writing the upgraded record fails, log it and let the user in on the old record rather than rejecting a correct password. The upgrade is **opportunistic maintenance**, not part of the authentication decision.

  • Why can a background job not raise the iteration count on all stored records overnight?
    Because hashlib.pbkdf2_hmac is one-way and the plaintext is not stored, so the job has nothing to derive from. It could only re-hash the existing derived key, which produces a bespoke chained construction rather than PBKDF2 at a higher count, and commits you to replaying that chain at every future verification. The only inputs that allow a genuine re-derivation arrive at login time.
  • What happens to accounts that never log in during a lazy upgrade?
    They keep their old parameters indefinitely, which is why lazy upgrade alone is not a migration plan. Track how many records remain below policy — the parameters are in each row, so it is a query — and set a cut-off date after which stale records are invalidated and the account must go through a password reset. Service accounts that never log in interactively need rotating by their own process.
  • Should the upgraded record reuse the original salt or generate a new one?
    Generate a new one with os.urandom. You are rewriting the row anyway, so a fresh 16-byte salt costs one call and removes any link between the old stored value and the new one — useful if the old table was ever exposed. Reusing the salt is not catastrophic, but there is no reason to preserve it once the record is being replaced.

It is like re-cutting a key you can only copy while the owner is standing at the door: you cannot upgrade the lock from the filing cabinet of old key blanks, only in the moment someone presents the original.

saying these in an interview costs you the question

  • Reads the iteration count from application config at verify time
  • Claims a background job can re-derive stored hashes
  • Re-hashes the stored digest to raise the work factor
  • Uses one shared salt for the whole users table
  • Compares derived keys with == on hex strings
  • Treats lazy upgrade alone as a finished migration

context