skip to content

Cryptographic Primitives

Picking the right primitive for randomness, secret comparison, password storage, digests and transport using only what ships in the box. Interviewers probe the wrong-primitive answer hard.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

22

How does hashlib's update, digest and hexdigest cycle hash data incrementally?

level: juniorimportance: must knowfreq 65%

answer

  1. The object accumulates; it is not a function
  2. Chunked feeding equals one big call
  3. Two renderings of the same value
  4. Reading a digest does not reset it
  5. Bytes in, TypeError on str

basics

~20 s

A hashlib hash object accumulates bytes across repeated update() calls, so feeding data in chunks matches hashing it in one call. digest() returns the raw bytes, hexdigest() the same value as a hex string; reading either does not reset the object.

solid answer

~40 s

`hashlib.sha256()` (or `hashlib.new("sha256")`) returns a mutable hash object, not a number. Every `update()` call appends to the running state, so `h.update(a); h.update(b)` is identical to `hashlib.sha256(a + b)` — that is what lets you hash a stream you never hold in memory. `digest()` returns `digest_size` raw bytes (32 for SHA-256); `hexdigest()` returns the same bytes rendered as a lowercase hex `str` of twice that length. Both are read-only snapshots: you can call them, keep updating, and call again. Input must be bytes-like — passing a `str` raises `TypeError`. `copy()` forks the current state, which is how you hash many messages that share a long common prefix without re-reading it.

code

python · 8 lines
python
import hashlib

h = hashlib.sha256()
h.update(b"lat,lon\n")
h.update(b"51.5,-0.12\n")
print(h.hexdigest() == hashlib.sha256(b"lat,lon\n51.5,-0.12\n").hexdigest())
print(len(h.digest()), h.digest_size, h.name)
print(h.hexdigest()[:16])

go deeper

for a junior

Be ready to build a SHA-256 digest in two lines and say what each call returns: update() feeds bytes, digest() gives raw bytes, hexdigest() gives the hex string. Remember that a str must be encoded first.

for a middle

Explain the mechanics: the object holds running state, chunk boundaries do not affect the result, reading a digest snapshots rather than finalizes, and digest_size fixes the output length regardless of input size.

for a senior

Show where the incremental form earns its keep in production — streaming input you never buffer, running checkpoints, copy() for a shared prefix — and be clear that an unkeyed digest proves integrity against accidents, not against an attacker.

for a principal

Own the standard: which digest the organisation uses for content addressing, whether digests are stored with their algorithm name so they can be migrated, and what it costs to rehash an existing corpus when that choice changes.

### A hash object is a state machine, not a function `hashlib` gives you one constructor per algorithm — `hashlib.sha256()`, `hashlib.sha512()`, `hashlib.blake2b()` — plus a lookup-by-name constructor, `hashlib.new("sha256")`, for when the algorithm arrives as configuration. All of them return the same kind of object: a mutable accumulator holding the digest's internal state. The convenience form passes the whole message to the constructor: ```python import hashlib hashlib.sha256(b"lat,lon\n51.5,-0.12\n").hexdigest() ``` The incremental form builds the same state in pieces: ```python h = hashlib.sha256() h.update(b"lat,lon\n") h.update(b"51.5,-0.12\n") h.hexdigest() ``` Both produce the identical digest, because the algorithm processes a byte stream and does not care where your chunk boundaries fell. That property is the whole point of `update()`: a digest over a 4 GB export, a socket, or a batch of records arriving one at a time never requires the full message in memory. What it does *not* mean is that the chunks are independent — the state carries forward, so the order of `update()` calls is part of the input. ### digest() versus hexdigest() `digest()` returns `bytes` of exactly `digest_size` bytes: 32 for SHA-256, 64 for SHA-512, 20 for SHA-1. `hexdigest()` returns a `str` of `2 * digest_size` lowercase hex characters — the same value, re-encoded for anywhere bytes are inconvenient: a log line, a filename, a JSON field, a URL. Neither is "more secure"; they are two renderings of one number. Choose `digest()` when the value goes back into binary machinery (a MAC comparison, a fixed-width column) and `hexdigest()` when a human or a text protocol will see it. The critical property for beginners: **calling either one does not finalize or reset the object.** Internally CPython snapshots the state and finishes a copy, so the accumulator survives untouched. This is legal and gives a running digest per line: ```python h = hashlib.sha256() for row in (b"a\n", b"b\n"): h.update(row) print(h.hexdigest()) ``` There is no `reset()` — to start over, build a new object. That asymmetry trips people who expect a stream API. ### Bytes only Hash functions are defined over octets, and Python refuses to guess an encoding for you: `hashlib.sha256("abc")` raises `TypeError: Strings must be encoded before hashing`. Any bytes-like object works — `bytes`, `bytearray`, a `memoryview` over a buffer — which is what makes zero-copy slicing of a read buffer possible. ### The useful extras Three attributes are worth knowing. `digest_size` is the output length in bytes; `block_size` is the algorithm's internal block length (64 for SHA-256), which matters when keying a digest; and `name` is the lowercase algorithm string, so a value that arrived through `hashlib.new()` can identify itself downstream when you store the digest alongside its algorithm. The last method is `copy()`. It clones the accumulator's state, so a long shared prefix is hashed once: ```python base = hashlib.sha256(b"a very long shared header...") for suffix in (b"-1", b"-2"): h = base.copy() h.update(suffix) ``` Without `copy()` you would re-hash the header per item. With it, each variant costs only its own tail. This is also the trick that makes a running checkpoint digest cheap: copy, finish the copy, keep feeding the original. ### What this cycle is and is not for The digest tells you two byte strings are the same (or, given a strong algorithm, that they are almost certainly not). It carries no key, so anyone can recompute it — a digest alone proves integrity against accidents, not against an attacker who can also rewrite the digest. It is also not a password store: a fast digest is fast for the attacker too, and both keyed digests and password hashing are separate APIs with their own rules. For everyday work — content addressing, dedupe keys, cache keys, "did this file change" — the update/digest cycle over a modern algorithm such as SHA-256 or BLAKE2 is exactly the right tool. ### Two mistakes that survive review The first is comparing across renderings. `h.hexdigest() == stored_bytes` is always `False`, because a `str` never equals a `bytes` object in Python and no exception warns you — the comparison simply reports "different" for every input. Decide once whether a stored digest is hex text or raw bytes, and convert at the boundary rather than at the comparison. The second is truncation done casually. Taking the first eight hex characters of a SHA-256 digest is fine as a log-friendly label and is not fine as an identity: 32 bits collide by birthday at roughly sixty-five thousand items. When you genuinely want a short fingerprint, ask the algorithm for one — `hashlib.blake2b(data, digest_size=16)` produces a proper 16-byte digest rather than a slice of a longer one — and never shorten a value that anything treats as unique.

  • Does calling hexdigest() finalize the object so further update() calls fail?
    No. The implementation finishes a copy of the internal state, so the object is untouched and you may keep calling update() afterwards. That is what makes a running per-chunk digest possible. There is no reset() counterpart, though: to hash a fresh message you construct a new object.
  • What is the difference between hashlib.sha256() and hashlib.new("sha256")?
    Nothing in the result — both return a SHA-256 hash object. The named constructors are direct module attributes and read better in fixed code; hashlib.new() takes the algorithm as a string, so it is the one to use when the algorithm comes from configuration or from a stored record, and it raises ValueError for a name this build does not support.
  • Why does hashlib.sha256("abc") raise TypeError instead of encoding the text?
    Hash functions are defined over bytes, and the digest of a str depends entirely on which encoding you pick — UTF-8 and Latin-1 give different digests for any non-ASCII character. Python refuses to choose silently, so you must call .encode() with an encoding your whole system agrees on.

saying these in an interview costs you the question

  • Thinks hexdigest() finalizes or resets the hash object
  • Believes chunked update() calls give a different digest than one call
  • Says digest() and hexdigest() are different algorithms or strengths
  • Passes a str and expects automatic UTF-8 encoding
  • Assumes there is a reset() method to reuse the object
  • Treats a bare digest as proof against a tampering attacker

context

open as a page

Which hashlib functions are built for password storage, and why not hashlib.sha256?

level: juniorimportance: must knowfreq 68%

basics

~20 s

hashlib offers two password functions: pbkdf2_hmac and scrypt. Both take a per-user salt and a work factor you choose, so one guess costs real time. hashlib.sha256 is built to be fast, which is exactly wrong for passwords.

open as a page

Why must a password-reset token come from `secrets`, not `random`?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The random module's default generator is a Mersenne Twister whose entire internal state can be reconstructed from a few hundred observed outputs, so later tokens become predictable. secrets draws each token from the operating system's cryptographic source instead.

open as a page

Why does `hmac.compare_digest` exist when `==` already compares two bytes objects?

level: juniorimportance: must knowfreq 50%

basics

~20 s

== on two bytes objects stops at the first differing byte, so the time it takes reveals how much of a secret an attacker guessed right. hmac.compare_digest always does the same work; secrets.compare_digest is the same function.

open as a page

What does setting ssl.SSLContext.verify_mode to ssl.CERT_NONE actually give up?

level: juniorimportance: must knowfreq 60%

basics

~10 s

It turns off certificate verification, so a Python client accepts any certificate at all, including one an attacker generated. The connection stays encrypted but is no longer authenticated, which defeats the point of TLS.

open as a page

How do you verify an inbound webhook's HMAC signature header in Python?

level: middleimportance: must knowfreq 55%

basics

~10 s

Recompute the tag with hmac.new over the exact raw request body bytes plus whatever else the sender signed, decode the header value to the same form, and compare with hmac.compare_digest. Never re-serialize parsed JSON.

open as a page

What does ssl.create_default_context() set up, and why prefer it to a hand-built ssl.SSLContext?

level: middleimportance: must knowfreq 50%

basics

~20 s

It returns a client context that already verifies: hostname checking on, certificates required, the system trust store loaded, and a TLS 1.2 floor. It carries the standard library current security policy, so it tightens as Python releases tighten.

open as a page

How does Python's hmac module produce an HMAC-SHA256 tag over a payload?

level: juniorimportance: should knowfreq 45%

basics

~10 s

Build an object with hmac.new(key, message, "sha256"), then read .digest() for raw bytes or .hexdigest() for a hex string. hmac.digest(key, message, "sha256") is the one-shot form. The key and message must be bytes.

open as a page

How do you hash a text address with hashlib so two services agree on the digest?

level: middleimportance: should knowfreq 38%

basics

~10 s

Encode explicitly and identically everywhere: pin one encoding (UTF-8), normalise the text first, and hash the resulting bytes. hashlib refuses a str with TypeError precisely because the encoding choice changes the digest.

open as a page

Why use hashlib.file_digest to hash a 4 GB file instead of reading it first?

level: middleimportance: should knowfreq 40%

basics

~20 s

hashlib.file_digest() streams the file through a reused 256 KiB buffer, so peak memory stays flat instead of holding the whole 4 GB. It returns the hash object, so you still call hexdigest() on the result.

open as a page

What do hashlib.scrypt's n, r and p control, and what is maxmem for?

level: middleimportance: should knowfreq 34%

basics

~20 s

In hashlib.scrypt, n is the cost parameter and must be a power of two, r is the block size, and p is parallelism. Working memory is roughly 128 * n * r bytes, and maxmem is the ceiling OpenSSL enforces on it.

open as a page

What does the argument to `secrets.token_hex(16)` count, and how long is the result?

level: middleimportance: should knowfreq 40%

basics

~10 s

It counts raw random bytes, not output characters. secrets.token_hex(16) requests 16 bytes — 128 bits — of entropy and returns a 32-character hexadecimal string, because hex encoding spends two characters per byte.

open as a page

Which argument types does `hmac.compare_digest` accept, and when does it raise TypeError?

level: middleimportance: should knowfreq 35%

basics

~20 s

Both arguments must be the same flavour: two bytes-like objects such as bytes, bytearray or memoryview, or two str values containing only ASCII. Mixing a str with bytes, or passing a str holding a non-ASCII character, raises TypeError.

open as a page

How do you set a minimum TLS version on a Python ssl.SSLContext, and what does that not guarantee?

level: middleimportance: should knowfreq 32%

basics

~20 s

Assign an ssl.TLSVersion member to SSLContext.minimum_version, for example ssl.TLSVersion.TLSv1_3. It only bounds what your side will negotiate; the peer, the linked OpenSSL build and the platform policy can each refuse more than you asked for.

open as a page

How do you build a canonical message so an HMAC verifies on both sides?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Define one exact byte encoding both sides implement: fixed field order, an unambiguous framing such as length prefixes, a pinned text encoding, and fixed formatting for numbers and timestamps. Sign those bytes, never a language object.

open as a page

How do you store a pbkdf2_hmac password record so the iteration count can be raised later?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Store the algorithm name, iteration count, salt and derived key together in one field. Verify with the record's own parameters, and when they fall below current policy re-derive at the new count during a successful login and rewrite the row.

open as a page

A video-metadata extractor calls `random.seed(1234)` at import; what does that do to the rest of the process?

level: seniorimportance: should knowfreq 33%

basics

~20 s

It fixes the stream for the whole process. The module-level random functions are bound methods of one hidden random.Random instance, so seeding it makes every other component's jitter, sampling and shuffling replay the same sequence. Give the extractor its own random.Random(1234).

open as a page

What does `hmac.compare_digest` still leak, and how do you compare variable-length secrets?

level: seniorimportance: should knowfreq 30%

basics

~20 s

It hides the contents but not the sizes: a timing attack could still reveal the lengths and types of the two arguments, never their values. Hash both sides to a fixed-size digest first, then compare the digests.

open as a page

A Python ticket-triage bot hits ssl.SSLCertVerificationError against an internal service behind a private CA. How do you fix it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Add the private CA to the client trust anchors with ssl.SSLContext.load_verify_locations, rather than disabling verification. Then decide deliberately whether that context should also keep the public roots, and make sure the CA file is actually deployed everywhere the bot runs.

open as a page

When is hashlib enough for password storage in a Python service, and when do you take on a native third-party hasher?

level: principalimportance: should knowfreq 26%

basics

~10 s

hashlib gives you PBKDF2 and scrypt, no Argon2 and no record format. Staying stdlib means owning that small amount of security-critical code; a native hasher means owning a compiled wheel on every deployment target.

open as a page

When would you reach for `random.SystemRandom` instead of the `secrets` functions?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

When you need the wider random API from an unpredictable source. random.SystemRandom is a random.Random subclass fed by os.urandom, so it offers shuffle, sample, randrange and uniform; secrets exposes only choice, randbelow, randbits and the token helpers.

open as a page

What does usedforsecurity=False do in hashlib.new("md5", usedforsecurity=False)?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

It declares that the digest is not being used for a security purpose, so a hardened or FIPS-restricted build permits a blocked algorithm such as MD5 instead of raising ValueError. On an ordinary build it changes nothing observable.

open as a page