Which hashlib functions are built for password storage, and why not hashlib.sha256?
answer
- Not every hashlib function fits here
- Speed is the attacker's friend
- Salt defeats tables, not speed
- Two functions take a work factor
- pbkdf2_hmac and scrypt
basics
~20 shashlib offers two password functions: pbkdf2_hmac and scrypt. Both take a per-user salt and a work factor you choose, so one guess costs real time. hashlib.sha256 is built to be fast, which is exactly wrong for passwords.
solid answer
~40 s`hashlib` mixes two different tools. `hashlib.sha256` and the other digest constructors are general-purpose hashes designed to be as fast as the hardware allows — a leaked table of `sha256(salt + password)` values is an offline oracle an attacker runs at hundreds of millions of guesses per second. The password-storage functions are `hashlib.pbkdf2_hmac(hash_name, password, salt, iterations)` and `hashlib.scrypt(password, salt=..., n=..., r=..., p=...)`: key-derivation functions whose cost the caller sets deliberately. The salt, 16 bytes from `os.urandom` per user, defeats precomputed tables and hides identical passwords; the work factor defeats raw speed. You need both, and salting a bare `hashlib.sha256` only buys the first. Both KDFs take `bytes`, so encode the password explicitly.
code
python · 7 linesimport hashlib
import os
salt = os.urandom(16)
fast = hashlib.sha256(salt + b"correct horse").hexdigest()
slow = hashlib.scrypt(b"correct horse", salt=salt, n=2**14, r=8, p=1, dklen=32).hex()
print(len(fast), len(slow))go deeper
Be ready to name hashlib.pbkdf2_hmac and hashlib.scrypt as the password functions and to say plainly that hashlib.sha256 is too fast for the job. Know that the salt comes from os.urandom and is stored, not hidden.
Explain the mechanics: the salt defeats precomputed tables, the iteration or cost parameter defeats raw guessing speed, and the two solve different problems. Mention that both functions take bytes and that dklen sets the output length.
Show the production judgment: where the salt is generated and stored, what the stored record must carry so parameters can change later, and why a hand-rolled loop over a fast digest is worse than the stdlib KDF on every axis.
Own the framing that this is a cost-asymmetry decision under a threat model, and that the standard library deliberately stops at two KDFs and no record format — so someone on your team owns the encoding and upgrade policy that hashlib does not provide.
`hashlib` exposes two families of function, and only one of them belongs anywhere near a password. ## The digest family `hashlib.sha256`, `hashlib.sha512` and `hashlib.blake2b` are **general-purpose cryptographic digests**. They take bytes and return a fixed-length digest as fast as the machine can manage — speed is their design goal, because they exist to fingerprint files, messages and cache keys. On a current CPU with hardware SHA instructions a single core computes SHA-256 at hundreds of megabytes per second, and purpose-built hardware does far better. If a users table storing `sha256(salt + password)` leaks, the attacker owns an **offline oracle** they can drive at that rate against a wordlist, one row at a time. The salt is still doing its job: - it stops one precomputed table from covering every row; - and it stops two accounts with the same password showing the same stored value. But a salt does not make a single guess any slower. Salting is not the missing ingredient. **Cost is.** ## The password family `hashlib.pbkdf2_hmac(hash_name, password, salt, iterations, dklen=None)` and `hashlib.scrypt(password, *, salt, n, r, p, maxmem=0, dklen=64)` are **password-based key-derivation functions**. Their defining feature is that the caller chooses how expensive one evaluation is. - **PBKDF2** repeats an HMAC over the password `iterations` times. - **scrypt** additionally forces the evaluation through a large working area of memory, which is what makes it awkward to parallelise on hardware built for cheap hashing. The **asymmetry** is the whole point: the defender evaluates the function once per login, the attacker once per guess, so a cost that is a barely noticeable delay for you is a wall for them. ## Arguments and types - Both functions take `bytes`, not `str`. Passing a `str` password raises `TypeError: a bytes-like object is required, not 'str'`, which is the right failure — encode explicitly with `password.encode("utf-8")` so the encoding is a decision rather than an accident, and so the same password verifies identically on every machine. - The **salt** is bytes too: generate 16 fresh bytes per user with `os.urandom(16)` (or `secrets.token_bytes(16)`) when the password is set, and store them in clear beside the derived key. A salt is not a secret; it is a uniqueness device, and it must be stored or you can never verify again. - `dklen` picks the output length — PBKDF2 defaults to the digest size of `hash_name`, scrypt defaults to 64 bytes. ## Anti-patterns that look like the fix - Writing a Python loop that calls `hashlib.sha256` a hundred thousand times is slow for you and still cheap for an attacker running optimised native code, and it is a home-made construction nobody has analysed. - Using the username as the salt makes the salt predictable and shared across systems. - Storing the derived key without recording which algorithm and parameters produced it makes every later change impossible. - And re-hashing the stored digest with a stronger function does not upgrade the record; it just builds a bespoke chain you must then carry forever. ## What the standard library does not give you On 3.14 there is no Argon2 and no bcrypt in `hashlib`, no encoded record format, and no built-in "does this stored value need re-hashing" check. You write the string that carries algorithm, parameters, salt and derived key, and you write the upgrade path. Both KDFs are thin wrappers over the interpreter's linked **OpenSSL** — CPython dropped the slow pure-Python PBKDF2 fallback in 3.12 — so an interpreter built without suitable OpenSSL support does not expose them at all. ## Verification To check a password you re-derive with the stored salt and the stored parameters and compare the two byte strings with a **constant-time comparison**, never with `==` on the hex text. And the comparison direction matters: you never decrypt a stored password, because a KDF output is not reversible; you only recompute and compare. ## Choosing the cost Neither function has a safe default, because the right value depends on your hardware and your tolerance for login latency. 1. Measure: time one derivation on a machine like production. 2. Raise `iterations` or `n` until a single call costs the largest delay you are willing to add to every login. 3. Then write the chosen value into the stored record rather than into a constant, so it can move later. A number copied from a tutorial written years ago is the most common way a correct call ends up providing a fraction of the protection it appears to. ## What to say in an interview - Name the two functions. - Say that the salt defeats precomputation while the **work factor** defeats speed. - And point out the trap in the question: `hashlib.sha256` sits in the same module behind the same import, so the wrong primitive is always one attribute away from the right one.
- If the per-user salt is stored in clear next to the hash, what is it actually protecting?Two things, neither of them secrecy. A unique salt per user means an attacker cannot build one precomputed table and test it against every row — the work must be redone per account. It also means two users with the same password produce different stored values, so a dump does not reveal which accounts share a password. It never slows down a targeted guess against one account; only the work factor does that.
- Why is looping hashlib.sha256 a hundred thousand times in Python not a home-made KDF?Because the cost asymmetry runs the wrong way. Your loop pays Python interpreter overhead per iteration while the attacker runs a tight native or GPU implementation, so you buy far less attacker time than you spend. It is also an unanalysed construction with no memory-hardness and no published parameter guidance. Call hashlib.pbkdf2_hmac or hashlib.scrypt, which do the iteration inside OpenSSL.
- Does hashlib give you a way to reverse a stored password hash for support purposes?No, and that is the design. Both hashlib.pbkdf2_hmac and hashlib.scrypt are one-way derivations with no inverse; there is no key that recovers the input. A support flow can only reset the password and let the user set a new one. If a system can email a user their existing password, it is not hashing at all — it is storing the plaintext or encrypting it under a recoverable key.
A digest is a photocopier: one press, instant copy. A password KDF is a copier deliberately geared to take a quarter-second per page — irrelevant when you copy one page a day, ruinous when you need a billion.
saying these in an interview costs you the question
- Says salting hashlib.sha256 makes it safe for passwords
- Calls a KDF output decryptable with the right key
- Uses the username or a shared constant as the salt
- Loops hashlib.sha256 by hand instead of a KDF
- Thinks the salt must be kept secret or encrypted
- Passes a str password and expects silent UTF-8 encoding