Why can Python's text-mode open() change a file's bytes, and when must you use binary mode?
answer
- A text stream is only a view
- Two transformations, both information-losing
- Many endings collapse to one character
- The digest disagrees across machines
- One argument fixes lines, not bytes
basics
~20 sText mode both decodes and rewrites line endings, so the characters you read are not the bytes on disk. Whenever the exact bytes matter — checksums, signatures, byte-for-byte copies, non-text formats — open in binary mode.
solid answer
~40 sA text-mode file object performs two transformations: a codec turns bytes into `str` and back, and newline translation rewrites line endings. Both are lossy in the sense that matters here — `'\r\n'` and `'\n'` in the file both arrive as `'\n'` in your string, so re-encoding the string cannot reproduce the original bytes. Any operation defined over bytes rather than characters must therefore use `'rb'`/`'wb'`: hashing, digital signatures, byte-for-byte comparison or copying, and every non-line-oriented format. `newline=''` removes the translation half but not the codec half, so it makes text mode faithful to the line endings, not byte-exact. When a file misbehaves, look at `pathlib.Path(p).read_bytes()` rather than the decoded string — the layer that caused the discrepancy is the same one hiding it.
code
python · 14 linesimport hashlib
import os
import tempfile
from pathlib import Path
path = os.path.join(tempfile.mkdtemp(), "meta.txt")
with open(path, "wb") as f:
f.write(b"duration=12\r\n")
with open(path, encoding="utf-8") as f: # '\r\n' collapses to '\n' in the str
text = f.read()
print(hashlib.sha256(text.encode("utf-8")).hexdigest()[:16]) # digest of 12 bytes
print(hashlib.sha256(Path(path).read_bytes()).hexdigest()[:16]) # digest of 13 bytesgo deeper
Recall that text mode decodes and rewrites line endings, so it is for human-facing text. Anything that is not line-oriented text — images, archives, media — is opened with 'rb' or 'wb'.
Explain both transformations and why translation is many-to-one, so a decoded string cannot be re-encoded back to the original bytes. Be able to say why newline='' fixes line endings but not byte-exactness.
Show the production judgment: recognise a checksum or comparison that silently disagrees across platforms, diagnose it from bytes rather than characters, and reject the tempting fix of normalising endings before hashing.
Own the boundary rule for the codebase: byte-defined operations use binary streams, character-defined ones state their encoding and newline explicitly. That one convention removes an entire class of defect that only part of a mixed-platform team can ever reproduce.
## Characters are not bytes `open(path)` gives you a text stream, and a text stream is a *view*: bytes are decoded to `str` on the way in, encoded on the way out, and line endings are rewritten in both directions. Each transformation is useful and each destroys information. Decoding is only round-trippable if you re-encode with the same codec and the text survives it. Newline translation is worse: with the default `newline=None`, `'\r\n'`, `'\r'` and `'\n'` all become the single character `'\n'`, so the mapping is many-to-one and the original is simply gone. A 13-byte file can produce a 12-character string, silently. That is why any operation whose meaning is defined over *bytes* must not go through text mode. ## A concrete failure Take a video-metadata extractor that writes a small sidecar file next to each asset and records a SHA-256 digest so a later stage can verify the sidecar was not modified. The digest is computed by reading the file back and hashing the encoded string. On an 11-person team, nine develop on Linux and two on Windows, and the build agent is Linux. Every sidecar written and verified on Linux checks out. The two Windows machines produce sidecars whose endings are CRLF on disk; reading them in text mode collapses each `'\r\n'` to `'\n'`, so the string that gets hashed is shorter than the file and the digest does not match the one another machine computed over the same asset. The verifier reports tampering on files nobody touched, and only for two of eleven people. Nothing raises, no log line points at line endings, and the reproduction rate looks random until someone notices it correlates with who produced the file. The root cause is worth naming precisely, because the fix people reach for first is wrong. It is not that CRLF is bad; it is that the digest was computed over a *decoded and translated view* rather than over the file. Normalising endings before hashing would paper over it and break the moment a sidecar legitimately contains a carriage return. The fix is to hash bytes: `hashlib.sha256(Path(p).read_bytes())`. A second, sharper variant of the same story: the extractor's writer helper is declared as `def write_sidecar(path, rows, opts={})` and stashes the chosen `newline` value in `opts`. Because that default is a single mutable object shared across every call, one early call that set a CRLF ending silently changed the behaviour of every later call in the process — so the same helper produced different bytes depending on call order. Byte-level divergence and a mutable default are a nasty pairing: the first makes the damage invisible, the second makes it non-deterministic. ## The rule for choosing a mode Use **binary mode** when the byte sequence is the thing: * hashing, HMACs, digital signatures, content-addressed storage; * byte-for-byte copies, uploads and downloads (or just use `shutil.copyfile`); * any format that is not line-oriented text — media containers, archives, images, wire protocol frames; * comparing two files for exact equality. Use **text mode with `newline=''`** when you want characters but the endings are data or belong to a layer above you: record-oriented writers and parsers, formats that specify CRLF, anything that must round-trip endings exactly while still decoding. Use **plain text mode** for ordinary human-facing text, where normalising endings is a feature. ## Diagnosing The discipline is: when a file surprises you, stop reading it as text. `pathlib.Path(p).read_bytes()` and `repr()` show the truth in one line; `len()` of the bytes versus `len()` of the decoded string localises the problem to translation immediately. In tests, assert on `read_bytes()` when the file's format is a contract with another system — a test that reads its own output back in text mode is blind to exactly the class of bug it exists to catch, because the same normalisation runs on both sides and cancels out. ## What newline='' does and does not buy you `newline=''` disables translation, so line endings survive intact. It does not make text mode byte-exact: the codec still runs, so the bytes you get back depend on the encoding you specify, and a mismatch between the encoding used to write and the one used to read still changes everything. Byte-exactness is a property of binary mode alone. Treat `newline=''` as "faithful to the lines" and binary as "faithful to the file".
- Does newline='' make a text-mode read byte-exact?No. It removes the newline translation, so the line endings survive, but the codec still runs: the characters you get depend on the encoding you named, and re-encoding with a different one produces different bytes. Only a binary mode gives you the file itself. Think of `newline=''` as faithful to the lines and binary as faithful to the file.
- How would you copy a file through Python without disturbing a single byte?Open source and destination in binary and stream between them, or simply call `shutil.copyfile`, which does exactly that. A copy written by reading with `read()` in text mode and writing the string back can change both the encoding and every line ending, and on a mixed-platform team it does so only for some contributors.
- Your test writes a file and reads it back in text mode, and it passes. Why is that weak?Because the same translation runs on both sides and cancels out: whatever ending was written is normalised back to `'\n'` on the way in, so the assertion holds no matter which bytes landed. Where the file is a contract with another system, assert on `pathlib.Path(p).read_bytes()` so the test sees what that system will see.
- How do you localise a suspected line-ending problem quickly?Compare lengths and look at the raw bytes: `len(Path(p).read_bytes())` against `len(Path(p).read_text())`, then `repr()` of a slice of the bytes. A difference equal to the line count points straight at CRLF collapsing under universal-newline reading, and a `b'\r\r\n'` in the output points at a doubled translation on the write side.
Reading in text mode is like reading a translated transcript of a speech: fine for understanding it, useless for verifying the speaker's signature on the original.
saying these in an interview costs you the question
- Assumes reading text mode returns the file's exact bytes
- Hashes a decoded string instead of the file's bytes
- Thinks newline='' makes text mode byte-exact
- Normalises line endings before hashing to make digests agree
- Copies a binary file by reading and writing it as text
- Asserts on read_text() when the bytes are the contract