When is hashlib enough for password storage in a Python service, and when do you take on a native third-party hasher?
answer
- Two functions and nothing around them
- The surrounding code is yours either way
- A compiled dependency has deployment costs
- Threat model, login rate, review capacity
- Self-describing records keep the exit open
basics
~10 shashlib gives you PBKDF2 and scrypt, no Argon2 and no record format. Staying stdlib means owning that small amount of security-critical code; a native hasher means owning a compiled wheel on every deployment target.
solid answer
~40 sThe standard library's answer is two functions. `hashlib.pbkdf2_hmac` and `hashlib.scrypt` are solid, memory-hard in scrypt's case, and available anywhere the interpreter's OpenSSL provides them — but `hashlib` ships no Argon2, no encoded record format and no "needs re-hashing" helper, so roughly fifty lines of encoding, verification and lazy-upgrade code become yours to write, review and keep correct. A native third-party hasher supplies all of that plus Argon2id, at the cost of a compiled extension: a wheel for every platform you build and run on, an ABI to track across interpreter upgrades including free-threaded builds, and one more link in your supply chain. Decide on the value of the credential store, your deployment constraints, whether anyone will review a bespoke record format, and whether federation removes the store entirely.
code
python · 7 linesimport hashlib
if hasattr(hashlib, "scrypt"):
scheme = "scrypt"
else:
scheme = "pbkdf2_sha256"
print(scheme, "sha256" in hashlib.algorithms_available)go deeper
Know that Python's hashlib contains PBKDF2 and scrypt but no Argon2 or bcrypt, and that anything beyond those two functions comes from a third-party package you must install.
Explain what a team must write themselves on top of hashlib — the record format, the comparison, the upgrade path — and that a native package removes that code but adds a compiled dependency.
Weigh login rate, memory per concurrent derivation, deployment images and review capacity, and show that a self-describing record keeps a later switch between schemes cheap.
Own the decision as a dependency and ownership tradeoff under a stated threat model, name what would change your mind, and keep the migration path open rather than debating which algorithm wins.
This is a judgement call about dependencies and ownership, not about which key-derivation function is theoretically strongest. Frame it that way in an interview. ## What the standard library actually gives you Two functions: `hashlib.pbkdf2_hmac` and `hashlib.scrypt`. Both are thin bindings over the interpreter's linked OpenSSL, both let you choose a work factor, and scrypt is memory-hard. What `hashlib` does not give you is everything around them — no Argon2, no bcrypt, no encoded record format, no helper that answers "was this record produced with parameters below current policy", no place to put a secret pepper, no parameter presets that someone else keeps current. That surrounding code is small but **load-bearing**: - the delimited record string, - the parse, - the constant-time comparison, - the lazy upgrade inside the login success branch, - and the tests that prove an old record still verifies after a parameter change. ## The case for staying on the standard library It is already installed, it has no wheel to build, no ABI to track and no new supply-chain link, and it works identically in a minimal container, on a locked-down base image with no compiler, and on an unusual interpreter build. Some regulated environments prefer or require the algorithms their approved cryptographic module provides, which is precisely what `hashlib` exposes. If your login volume is modest, your threat model is ordinary, and someone competent will review those fifty lines once, PBKDF2 or scrypt from `hashlib` with well-chosen parameters is a defensible, boring answer — and **boring is the correct aesthetic** for credential storage. ## The case for the dependency A maintained password-hashing package gives you Argon2id, an encoded record format that already carries the parameters, a rehash check, and parameter defaults that track current guidance without anyone on your team tracking it. If a credential compromise is a company-level event, buying a well-reviewed implementation and a maintained parameter policy is a better use of scarce security attention than writing a bespoke record format nobody will re-read. The price is operational: a native extension needs a wheel for every operating system, architecture and interpreter version you deploy — including containers, developer laptops on a different architecture, and any free-threaded build you adopt — or a compiler in the build image. It becomes something that can block an interpreter upgrade, and something a supply-chain review must cover. ## The axes that actually decide it - What does a full dump of the credential store cost the business? - What is the login rate at peak, and what CPU time and memory per login can you afford — remembering scrypt's roughly `128 * n * r` bytes multiplied by concurrent logins? - Do you control the deployment images, or do you ship into environments where a compiled dependency is a negotiation? - Who reviews the record format and the upgrade path if you write them? - How long do you expect this store to live, and how many parameter or algorithm migrations will it see? - And the question that dissolves the problem: do you need to store passwords at all, or does delegating authentication to an identity provider remove the asset — at the price of a different dependency and a different failure mode. ## The migration argument Whichever you pick, the exit must exist. A **self-describing record** — algorithm name, parameters, salt, key — makes moving between `hashlib.pbkdf2_hmac`, `hashlib.scrypt` and a future third-party hasher the same operation: 1. dispatch on what the row says, 2. verify with the old scheme, 3. re-derive with the new one at the next successful login, 4. and measure how many rows remain. Teams that hard-code a single scheme are the ones for whom this decision becomes irreversible, which is what makes it worth deciding deliberately rather than by whichever import someone reached for first. ## How to answer State the axes, pick a default for a stated context — for most services, stdlib scrypt with measured parameters and a self-describing record is enough — and name the conditions that would change your mind: - a high-value credential store, - a team without review capacity for hand-written security code, - or an existing platform standard. A principal-level answer names the operational costs of the dependency as clearly as its benefits, and refuses to turn the question into an algorithm beauty contest.
- What concretely breaks first when a native password-hashing extension enters a Python deployment?The build and upgrade path, long before anything cryptographic. You need a wheel for every operating system, architecture and interpreter version in play, or a compiler and headers in the image; a developer on a different architecture hits it first, and an interpreter upgrade can stall until the wheel exists. Add the supply-chain review of a compiled artefact, and that is the real cost of the dependency.
- If you stay on hashlib, what code do you now own that a hashing package would have provided?The encoded record — algorithm name, parameters, salt and key in one field — the parser, the constant-time comparison, the lazy re-hash inside the login success branch, the parameter choices and the evidence behind them, and tests proving an old record still verifies after parameters change. It is perhaps fifty lines, but it is security-critical and it needs a named owner and a review, not a copy from a blog post.
- Does moving authentication to an identity provider remove this decision or relocate it?It removes the credential store, which is the asset you were protecting, and that is a genuine reduction in blast radius. It relocates the risk into an integration you do not control: token validation, session lifetime, the provider's own availability and the break-glass path when it is down. It also rarely covers every account — service and legacy logins tend to remain — so most teams end up with a smaller store rather than none.
saying these in an interview costs you the question
- Turns the question into which algorithm is strongest
- Ignores wheel and ABI costs of a native extension
- Assumes hashlib includes Argon2 or bcrypt
- Writes a bespoke record format with no reviewer
- Forgets memory per login multiplies by concurrency
- Treats the choice as irreversible from day one