skip to content

What do hashlib.scrypt's n, r and p control, and what is maxmem for?

level: middleimportance: should knowfreq 34%

answer

  1. Three cost knobs, one safety ceiling
  2. One knob must be a power of two
  3. Memory scales with two parameters together
  4. 128 * n * r bytes
  5. Default maxmem is OpenSSL's 32 MiB

basics

~20 s

In hashlib.scrypt, n is the cost parameter and must be a power of two, r is the block size, and p is parallelism. Working memory is roughly 128 * n * r bytes, and maxmem is the ceiling OpenSSL enforces on it.

solid answer

~40 s

The signature is `hashlib.scrypt(password, *, salt, n, r, p, maxmem=0, dklen=64)` — everything after the password is keyword-only. `n` is the CPU and memory cost parameter and must be a power of two greater than one, or you get `ValueError: n must be a power of 2`. `r` is the block size in 128-byte units, and `p` is parallelism, which multiplies CPU work without multiplying memory. Working memory is about `128 * n * r` bytes, so `n=2**14, r=8` needs roughly 16 MiB. `maxmem` defaults to 0, meaning OpenSSL's own 32 MiB limit — so `n=2**16, r=8` fails with a memory-limit `ValueError` until you pass a large enough `maxmem` explicitly. `dklen` is the output length, 64 by default.

code

python · 11 lines
python
import hashlib
import os

salt = os.urandom(16)
try:
    hashlib.scrypt(b"pw", salt=salt, n=2**16, r=8, p=1, dklen=32)
except ValueError as exc:
    print("rejected:", exc)

key = hashlib.scrypt(b"pw", salt=salt, n=2**16, r=8, p=1, dklen=32, maxmem=2 * 128 * 2**16 * 8)
print(len(key))

go deeper

for a junior

Know that hashlib.scrypt needs a salt plus three cost parameters passed by keyword, that n must be a power of two, and that dklen only sets output length rather than strength.

for a middle

Explain what each parameter does, derive the roughly 128 * n * r memory figure, and recognise that a memory-limit ValueError points at the default maxmem of 32 MiB rather than at the machine's RAM.

for a senior

Show that you size parameters by measuring on production-like hardware and by multiplying memory per call by peak concurrent logins, and that verification must replay the parameters stored with each record.

for a principal

Own the tradeoff between login latency, per-login memory footprint and attacker cost as a capacity decision, including what a burst of simultaneous logins does to a service sized for the average.

`hashlib.scrypt(password, *, salt, n, r, p, maxmem=0, dklen=64)` is the standard library's **memory-hard password-based key derivation function**. Only `password` is positional; `salt`, `n`, `r` and `p` are keyword-only and must be supplied. ## The parameters `n`, `r` and `p` - **`n` — the cost parameter.** scrypt fills a table of `n` blocks and then walks it in a data-dependent order, so `n` scales both the time and the memory of one evaluation. It must be a power of two greater than one; anything else raises `ValueError: n must be a power of 2`. Doubling `n` doubles both the work and the footprint, which is why it is the knob you actually turn. - **`r` — the block size.** Each of those `n` blocks is `128 * r` bytes. Raising `r` makes each memory access move more data, which is a way to lean on memory bandwidth rather than on the number of accesses. It is conventionally left at 8. - **`p` — parallelism.** `p` independent instances of the core mixing function are run and combined. It multiplies CPU work but not the peak memory of a single instance, so it is the parameter for buying more cost when you cannot afford more RAM. It is conventionally 1. ## The memory formula Peak working memory is approximately `128 * n * r` bytes. With `r=8`: | `n` | Peak memory | |---|---| | `n=2**14` | about 16 MiB | | `n=2**15` | about 32 MiB | | `n=2**16` | about 64 MiB | That is memory **per concurrent call**, which is the number that matters operationally — a hundred simultaneous logins at 64 MiB each is 6.4 GiB of transient allocation, and Python's threads do not make that cheaper. ## `maxmem` — the parameter that surprises people It is not a tuning knob for the algorithm; it is a **safety ceiling** handed to OpenSSL, which refuses to allocate beyond it. The default `maxmem=0` means "OpenSSL's own default", which is 32 MiB. So the first time you try realistic parameters — `n=2**16, r=8`, needing 64 MiB — the call fails with `ValueError: [digital envelope routines] memory limit exceeded`, and the fix is to pass `maxmem` explicitly at or above `128 * n * r` (with headroom), not to reduce your security parameters. This trips up almost everyone who moves from a tutorial's `n=2**14` to production values, and recognising the error message is genuinely useful interview knowledge. ## `dklen` The derived key length in bytes, default 64. For password storage 32 is plenty; the length has no effect on cost and is not a security parameter you tune for hardness. ## Choosing values Measure rather than copy. 1. Wrap a call in `time.perf_counter` on hardware resembling production. 2. Raise `n` until one derivation costs the largest delay you are willing to add to a login and the largest allocation you are willing to multiply by your peak concurrency. A common starting point is `n=2**14` or `2**15`, `r=8`, `p=1`, then adjust. The parameters are a **capacity decision** as much as a security one: every login pays them, and so does every attacker guess. ## Verification uses the stored parameters, not today's Because the derivation is a pure function of password, salt and parameters, verifying a stored value means calling `hashlib.scrypt` with exactly the `n`, `r`, `p` and `dklen` that produced it. That is why the stored record must carry them; a global constant read from configuration silently invalidates every record the moment someone raises it. ## Contrast with the other stdlib KDF `hashlib.pbkdf2_hmac` has a single cost knob, `iterations`, and uses negligible memory, which is exactly why memory-hard functions were invented: cheap-memory hardware parallelises PBKDF2 far better than it parallelises scrypt. - If a compliance regime or an old record format pins you to PBKDF2, the equivalent conversation is about iteration counts. - If you have a free choice inside `hashlib`, scrypt's memory cost is the reason to prefer it. ## Everything is keyword-only for a reason Because `salt`, `n`, `r` and `p` must be named at the call site, you cannot accidentally pass the salt where the cost belongs or reorder two integers into a silently weaker call. It also makes a code review of the call site trivial: every security-relevant input is visible with its name attached, and a missing one is a `TypeError` rather than a default that quietly does less work than you think. ## Availability `hashlib.scrypt` comes from the interpreter's linked OpenSSL, so a build without scrypt support does not expose the attribute at all — worth a guarded check rather than an assumption in code that must run on unfamiliar images.

  • Why raise n rather than p when you want a hashlib.scrypt call to cost more?
    Because n buys memory hardness and p does not. Raising n enlarges the table the algorithm must fill and walk, so an attacker's custom hardware needs proportionally more RAM per guess — that is the property scrypt exists for. Raising p only runs more instances of the same small computation, which parallel hardware absorbs cheaply. Use p when you want more CPU cost but cannot afford more memory per concurrent login.
  • A hashlib.scrypt call that works locally raises ValueError about a memory limit in a container. What changed?
    Almost certainly nothing about the parameters — it is the maxmem ceiling, not the container's memory. maxmem=0 means OpenSSL's 32 MiB default, so any call needing more than that fails identically everywhere; a local success usually means the local code passed a larger maxmem or smaller n. Pass maxmem explicitly, sized at or above 128 * n * r with headroom, and keep the value alongside the parameters you chose.
  • Does raising dklen in hashlib.scrypt make the stored password harder to attack?
    No. dklen only sets how many bytes of output the function emits; it does not change the number of memory accesses or the cost of one evaluation. An attacker still pays the same price per guess and compares whatever prefix you stored. Thirty-two bytes is a sensible choice for password storage, and cost lives entirely in n, r and p.

saying these in an interview costs you the question

  • Treats maxmem as the algorithm's memory cost knob
  • Thinks any integer is valid for n
  • Says p multiplies memory the way n does
  • Raises dklen to make the hash stronger
  • Verifies stored hashes with today's parameters, not the record's
  • Ignores that memory cost multiplies by concurrent logins

context