What does Python's open() do to line endings when newline is left at its default?
answer
- Text mode does more than decode
- Three conventions, one in-memory form
- Read normalizes, write platform-izes
- The default is newline=None
- os.linesep appears on the write side
basics
~10 sWith the default newline=None, open() reads in universal-newline mode: '\r\n', '\r' and lone '\n' all arrive in your string as '\n'. On write, every '\n' you emit is translated to os.linesep.
solid answer
~40 sA file opened in text mode is wrapped in an `io.TextIOWrapper`, which decodes bytes to `str` **and** translates line endings. With the default `newline=None` this is universal-newline mode: on read, any of `'\r\n'`, `'\r'` or `'\n'` in the file becomes a single `'\n'` in the string you get back, so your code never has to know which convention wrote the file. On write it runs the other way: every `'\n'` you pass to `write()` is replaced with `os.linesep`, which is `'\r\n'` on Windows and `'\n'` elsewhere. Pass `newline=''` to keep the original characters, or `newline='\r\n'` to force CRLF output on every platform. Binary mode (`'rb'`/`'wb'`) does none of this — it moves bytes through untouched, and passing `newline=` with a binary mode raises `ValueError`.
code
python · 12 linesimport os
import tempfile
path = os.path.join(tempfile.mkdtemp(), "data.txt")
with open(path, "wb") as f:
f.write(b"a\r\nb\rc\n")
with open(path, encoding="utf-8") as f: # newline=None (the default)
print(repr(f.read())) # 'a\nb\nc\n'
with open(path, encoding="utf-8", newline="") as f: # translation switched off
print(repr(f.read())) # 'a\r\nb\rc\n'go deeper
Be ready to say what the default does in both directions: reading normalizes every convention to '\n', writing turns '\n' into the platform ending. Knowing that binary mode skips all of it is enough at this level.
Explain the mechanics: a TextIOWrapper decodes and translates, and name the accepted newline values and what each does on read versus write. Expect to be asked which one preserves the original characters.
Show that you know when the translation is a hazard — byte-exact round-trips, checksums, formats with their own terminators — and that you reach for read_bytes() to diagnose rather than guessing from the decoded string.
Own the convention: pick one policy for the codebase (explicit encoding, explicit newline at every text boundary, binary for anything byte-exact) so platform differences never become a class of bug the team rediscovers per file format.
## Two jobs hiding behind one call `open(path)` in text mode returns an `io.TextIOWrapper` sitting on top of a raw byte stream, and that wrapper does two separate transformations. The first is the codec: bytes are decoded to `str` on the way in and encoded on the way out. The second, much less visible, is **newline translation**. Both are on by default, and both can change what you observe relative to what is on disk. This section is about the second one. Three conventions for ending a line are still in circulation. Unix-descended systems use a single line feed, `'\n'` (0x0A). Windows, DOS and a long list of network protocols use a carriage return followed by a line feed, `'\r\n'` (0x0D 0x0A). Classic Mac OS used a bare carriage return, `'\r'`. Python's answer is not to pick a winner but to normalize: inside your program a line ends with `'\n'` and nothing else, and the boundary layer converts. ## Reading With `newline=None` — the default — the wrapper is in universal-newline mode. It recognises all three sequences as line terminators, and it **replaces every one of them with a single `'\n'`** before your code sees the characters. That is true for `read()` as well as for iteration and `readlines()`, so a file written on Windows and a file written on Linux produce identical strings. It also means the string you get back is not a faithful picture of the file: a 13-byte file can decode to a 12-character string with no error and no warning. The `newline` argument on read accepts exactly five values, and they split into two behaviours: * `None` — recognise all three, translate them all to `'\n'`. * `''` — still recognise all three as line boundaries (so iterating the file splits in the right places), but **translate nothing**; the original characters stay in the string. * `'\n'`, `'\r'` or `'\r\n'` — only that exact string terminates a line, and again nothing is translated. Any other value raises `ValueError`. The distinction between `None` and `''` is the one that matters in practice: both let you iterate lines sensibly, but only `''` preserves the bytes' shape. ## Writing The write side mirrors it. With `newline=None`, every `'\n'` in the string you pass to `write()` is replaced by `os.linesep` — `'\r\n'` on Windows, `'\n'` on Linux and macOS. This is why the same script emits CRLF files on one machine and LF files on another without a line of platform code. With `newline=''` or `newline='\n'` no translation happens at all and your `'\n'` reaches the disk as one byte. With `newline='\r'` or `newline='\r\n'` each `'\n'` becomes that sequence, on every platform — which is how you deliberately produce CRLF output from Linux. Note what the write side does *not* do: it only ever looks at `'\n'`. A `'\r'` you put in the string yourself is passed through untouched, which is exactly how files end up with `'\r\r\n'` — you supplied the `'\r'`, and text mode added its own after the `'\n'`. ## Binary mode opts out entirely `open(path, 'rb')` and `open(path, 'wb')` return byte streams with no wrapper, no codec and no newline translation; the bytes on disk are the bytes you get. Because there is nothing to translate, combining a binary mode with the argument is rejected outright: `open(path, 'rb', newline='')` raises `ValueError: binary mode doesn't take a newline argument`. When the exact bytes matter — hashing, signing, copying, uploading, or any format that is not line-oriented text — binary mode is the correct tool, not a cleverer `newline=` value. ## A neighbouring trap: splitlines Universal-newline mode recognises three sequences. `str.splitlines()` recognises far more: besides `'\n'`, `'\r'` and `'\r\n'` it also splits on vertical tab, form feed, the file/group/record separators, and the Unicode characters NEL, LINE SEPARATOR and PARAGRAPH SEPARATOR. So `text.splitlines()` can produce more lines than iterating the same file did. When you need exactly the file's notion of a line, iterate the file object or use `str.split('\n')`; reach for `splitlines()` when you want the generous Unicode definition. ## What to actually do For ordinary text — reading a config file, writing a report — leave `newline` alone and enjoy the normalization. Reach for `newline=''` when another layer already emits its own terminators or when embedded carriage returns are data. Reach for a binary mode when the file is not text at all. And when a file surprises you, look at `pathlib.Path(p).read_bytes()` rather than `read_text()`: the translation is invisible in the string and obvious in the bytes.
- How do you read a file so the original '\r\n' survives into the string?Pass `newline=''` to `open()`. Line boundaries are still recognised, so iterating the file yields sensible lines, but no characters are replaced and each line keeps its `'\r\n'` ending. `newline='\n'` also avoids translation but changes what counts as a terminator, so a bare `'\r'` no longer ends a line. Binary mode gives you the bytes with no decoding at all.
- What happens if you pass newline='' together with a binary mode?`ValueError: binary mode doesn't take a newline argument`. Binary streams neither decode nor translate, so the argument is meaningless there; the same applies to `encoding=` and `errors=`. If you find yourself wanting it, you either want text mode with `newline=''` or you want plain bytes.
- Does str.splitlines() split on the same set of endings as universal-newline reading?No — it splits on a wider set. Universal-newline mode recognises `'\n'`, `'\r'` and `'\r\n'`; `str.splitlines()` additionally splits on vertical tab, form feed, the file/group/record separators and the Unicode NEL, LINE SEPARATOR and PARAGRAPH SEPARATOR characters. A string containing one of those yields more lines than the file's own iteration produced.
Text mode is a customs desk: whatever line-ending currency the file arrives in, you are handed the local one, and on the way out it is converted back to whatever the destination platform spends.
saying these in an interview costs you the question
- Says Python only translates line endings on Windows
- Believes text-mode reading preserves the file's exact bytes
- Thinks newline=None means no translation at all
- Assumes newline='' turns off line splitting as well
- Passes newline='' alongside a binary mode like 'rb'
- Treats str.splitlines() as identical to iterating the file