skip to content

Why use hashlib.file_digest to hash a 4 GB file instead of reading it first?

level: middleimportance: should knowfreq 40%

answer

  1. One buffer, not the whole file
  2. The loop moved into the standard library
  3. It hands back the object, not the string
  4. Binary mode or it refuses
  5. readinto plus a memoryview slice

basics

~20 s

hashlib.file_digest() streams the file through a reused 256 KiB buffer, so peak memory stays flat instead of holding the whole 4 GB. It returns the hash object, so you still call hexdigest() on the result.

solid answer

~40 s

`f.read()` materialises the entire file as one `bytes` object before hashing, so a 4 GB file needs 4 GB of resident memory and a large allocation that may fail outright. `hashlib.file_digest(f, "sha256")`, added in Python 3.11, does the read-and-update loop for you: it allocates one 256 KiB `bytearray`, calls the file's `readinto()` repeatedly, and updates the hash from a `memoryview` slice — so no per-chunk copies and constant memory whatever the file size. It requires a file object opened in **binary** mode and readable, and raises `ValueError` otherwise; a text-mode file has no `readinto()`. It hashes from the current position to EOF and leaves the position there. It returns the hash object, not a string, so you finish with `.hexdigest()`.

code

python · 17 lines
python
import hashlib, os, tempfile

with tempfile.NamedTemporaryFile(delete=False) as f:
    f.write(b"lat,lon\n51.5,-0.12\n")
    path = f.name

with open(path, "rb") as f:
    digestobj = hashlib.file_digest(f, "sha256")
print(digestobj.name, digestobj.hexdigest()[:16])

with open(path, "r") as f:
    try:
        hashlib.file_digest(f, "sha256")
    except ValueError as exc:
        print("rejected text mode:", exc)

os.remove(path)

go deeper

for a junior

Remember that hashing a big file should never start with read(). Open it with "rb", pass the file object to hashlib.file_digest(), and call hexdigest() on what comes back.

for a middle

Explain the mechanics: one reused buffer filled by readinto(), a memoryview slice fed to update(), constant memory regardless of file size, binary mode required, and hashing from the current position to EOF.

for a senior

Show the operational judgement — bounded memory under a container limit, hashing streams you cannot seek, storing the digest with its algorithm name, and being explicit that an unkeyed checksum detects corruption, not tampering.

for a principal

Own the content-addressing policy: which digest the platform standardises on, whether digests are recorded with their algorithm so a future migration is possible, and what rehashing an existing corpus would cost.

### The problem it replaces Before Python 3.11 every codebase carried its own version of this loop: ```python h = hashlib.sha256() with open(path, "rb") as f: while chunk := f.read(65536): h.update(chunk) ``` and a depressing number of codebases carried the shortcut instead: ```python h = hashlib.sha256(open(path, "rb").read()) ``` The shortcut is correct for a 40 KB config and catastrophic for a 4 GB export: `read()` builds one `bytes` object holding the whole file, so peak resident memory tracks the largest input you will ever be handed. On a container with a memory limit it does not degrade — it dies, and it dies on the one day someone uploads an unusually large batch. `hashlib.file_digest(fileobj, digest)` is that loop, written once, in the standard library. ### What it actually does The second argument is either an algorithm name (`"sha256"`) or a callable returning a hash object (`hashlib.sha256`), so it composes with `hashlib.new()`-style configuration. Then: * If the object exposes `getbuffer()` — an `io.BytesIO`, for instance — it hashes that buffer directly, with no copy at all. * Otherwise it requires a readable binary file. It allocates a single `bytearray` of 256 KiB, wraps it in a `memoryview`, and loops on `readinto(buf)`, feeding `view[:size]` to `update()` until a read returns 0. Two details make that better than the naive loop. `readinto()` fills an existing buffer rather than allocating a fresh `bytes` per chunk, so the allocator sees one buffer for the whole file instead of tens of thousands of short-lived objects. And the `memoryview` slice hands the hash a view of the filled prefix without copying the tail. It returns the **hash object**, not a digest — a small API decision people trip over exactly once: ```python with open(path, "rb") as f: digestobj = hashlib.file_digest(f, "sha256") print(digestobj.hexdigest()) ``` Returning the object means you can also ask for `digest()`, `digest_size` or `name` without re-reading the file. ### The constraints worth knowing **Binary mode is mandatory.** A file opened in text mode is a decoding wrapper with no `readinto()`, and the call raises `ValueError` naming the object as not being a file-like object in binary reading mode. That refusal is a feature: hashing decoded text would silently make the digest depend on the platform's default encoding and on newline translation, so the same file hashed on two hosts could disagree. **It reads from the current position to EOF and does not seek back.** If you have already consumed part of the stream, you hash only the remainder; if you need the digest and then the contents, seek to 0 yourself. This also makes it usable on a pipe or socket wrapper, where seeking is not an option at all. **It is a hashing helper, not a file API.** It does not open the file, does not close it, and does not care about the path — which is what lets it hash anything readable and binary, including an in-memory buffer, a decompressing stream, or a network body reader. ### Where it fits in real work Content addressing and change detection are the everyday uses: store `hexdigest()` alongside an artefact and you can answer "has this changed?" without a byte-by-byte comparison, and you can dedupe identical uploads that arrived under different names. For an integrity check against a published checksum, the digest must be compared to a value delivered over a channel you trust — the digest itself carries no key and an attacker who can rewrite the file can usually rewrite a checksum sitting next to it. Throughput is respectable because the hashing work happens in C and the interpreter releases the GIL around a sufficiently large buffer, so a thread pool hashing several files at once genuinely overlaps. The remaining cost is I/O, and the 256 KiB default is already tuned for it; the buffer size is a private parameter, not a public tuning knob, so if you truly need different chunking you write the loop by hand. On Python 3.10 and earlier there is no `file_digest` — write the `read()`/`update()` loop, and keep the chunk size in the hundreds of kilobytes rather than a few bytes. ### When you need more than a digest `file_digest` is for the case where hashing is the *only* job: it consumes the stream to EOF and you get a digest back. When you must also process the bytes — parse rows, write them elsewhere, compress them — go back to the explicit loop and call `update()` on each chunk as it passes through, so the data is read exactly once and the digest falls out for free at the end. Hashing on write is the same pattern in the other direction: update the digest as you emit bytes, then record it beside the artefact so a later reader can confirm what it received matches what was produced.

  • What does hashlib.file_digest return, and what is the most common mistake with it?
    It returns a hash object, not a digest string, so the usual mistake is logging the object or comparing it to a stored hex value. Call .hexdigest() (or .digest()) on the result. Returning the object also lets you read .name or .digest_size without touching the file again.
  • Why does it reject a file opened in text mode rather than encoding for you?
    A text-mode file is a decoding wrapper with no readinto(), so the streaming path does not apply, and hashing decoded text would make the digest depend on the platform's default encoding and on newline translation. Refusing keeps the digest a property of the bytes on disk rather than of the host that read them.
  • You need digests for a directory of large files as fast as possible. Does threading help?
    Usually yes, because the hashing runs in C and the interpreter releases the GIL around a large buffer, so threads genuinely overlap hashing with each other and with I/O waits. A ThreadPoolExecutor over the file list is the simple answer; the ceiling is then storage throughput, not the interpreter.

saying these in an interview costs you the question

  • Calls f.read() on a multi-gigabyte file before hashing
  • Expects file_digest to return a hex string directly
  • Passes a path string instead of an open file object
  • Opens the file in text mode and is surprised by ValueError
  • Assumes it rewinds the file before hashing
  • Thinks a plain checksum protects against a tampering attacker

context