skip to content

What does open()'s newline= argument control in text mode?

level: middleimportance: should knowfreq 42%

answer

  1. Three platforms, three line endings
  2. Reading normalises, writing expands
  3. The empty string turns the translation off
  4. A comma-separated format wants raw endings
  5. Binary mode rejects the argument outright

basics

~20 s

It controls line-ending translation. The default newline=None turns any of \r\n, \r or \n into \n when reading and turns \n into the platform's line ending when writing; newline='' disables translation while still splitting lines.

solid answer

~40 s

Text mode does universal-newline handling. With the default `newline=None`, reading translates `\r\n` and a lone `\r` into `\n`, so your code only ever sees `\n`; writing translates every `\n` you write into the platform's line separator, which is `\r\n` on Windows. Passing `newline=""` turns translation off in both directions while still recognising all three sequences as line breaks for the purpose of splitting lines — that is why the `csv` module asks for it, since a quoted field may legitimately contain an embedded newline and the writer emits its own `\r\n`. Passing an explicit `"\n"` or `"\r\n"` pins the ending you want, which is how you produce a file with a fixed format regardless of platform. Binary mode never translates and rejects `newline=` with `ValueError`.

code

python · 11 lines
python
with open("crlf.txt", "w", encoding="utf-8", newline="\r\n") as f:
    f.write("a\nb\n")

with open("crlf.txt", "rb") as f:
    print("raw:", f.read())

with open("crlf.txt", encoding="utf-8", newline="") as f:
    print("untranslated:", f.readlines())

with open("crlf.txt", encoding="utf-8") as f:
    print("universal:", f.readlines())

go deeper

for a junior

Know that text mode normalises line endings so your code sees \n whatever platform wrote the file, and that you should write a plain \n rather than assembling \r\n yourself.

for a middle

Explain the read and write behaviour of newline=None, '' and an explicit terminator, and why a comma-separated-values parser wants translation switched off in both directions.

for a senior

Diagnose real artefacts: doubled \r\r\n from hand-written line separators, a format parser mis-splitting quoted fields, and a size assertion that only fails on one platform. Know to inspect the raw bytes before theorising.

for a principal

Decide the convention for files that cross team or system boundaries: a pinned line ending in the format spec beats per-platform defaults, and it should be enforced where the file is produced rather than patched by every consumer.

Different systems ended lines differently: Unix with a single `\n`, Windows with the pair `\r\n`, older Mac systems with a lone `\r`. Python's text mode hides that from you by default, and `newline=` is the knob that says how much hiding you want. ## Reading * `newline=None` (the default) — *universal newlines*. Any of `\r\n`, `\r` or `\n` in the file is recognised as a line break, and every one of them is translated to `\n` in the `str` you get back. Your parsing code can then assume a single line-ending convention no matter who produced the file. This is why `line.rstrip("\n")` is usually enough even for a file that came from Windows. * `newline=""` — all three sequences are still recognised as line breaks, so iteration and `readlines()` split in the same places, but **no translation happens**: the line you get back still ends in the exact bytes the file had. * `newline="\n"`, `"\r"` or `"\r\n"` — only that sequence counts as a line break, and nothing is translated. A file with the wrong ending then reads as one enormous line, which is occasionally what you want when you are validating a format. ## Writing * `newline=None` — every `\n` you write is translated to the platform's line separator. On Windows that means the four characters `"a\nb\n"` become six bytes. This is the source of the classic "my file has doubled line endings" bug: code that writes `os.linesep` itself in text mode ends up emitting `\r\r\n` on Windows, because the `\n` inside `os.linesep` is translated again. * `newline=""` or `newline="\n"` — no translation; the characters you write are the characters that are encoded. * `newline="\r\n"` — every `\n` becomes `\r\n` regardless of platform, which is how you deliberately produce a CRLF file on Linux. The rule to remember for writing: **write `\n` and let `newline=` decide the bytes**. Never hand-roll platform line endings into text you write in text mode. ## Why the csv module insists on newline='' The `csv` writer emits `\r\n` at the end of each record by default, because that is what the format specifies. If the file were opened with the default translation, that `\n` would be translated *again* on Windows, giving `\r\r\n` and a file other readers mis-parse. On the reading side, a quoted CSV field may legally contain a newline inside it, and the reader has its own logic for telling an embedded newline apart from a record separator. Handing it pre-translated text destroys the distinction it needs. So both `csv.reader` and `csv.writer` want the file opened with `newline=""`, and the codec still comes from `encoding=` as usual. The general shape of that lesson goes beyond CSV: whenever a *parser* owns the line-ending semantics, give it untranslated text and stay out of its way. ## Binary mode Binary mode does nothing to line endings — `\r\n` in the file is `b"\r\n"` in your `bytes`, always, on every platform. Passing `newline=` to a binary open raises `ValueError`, since translation is defined only for text. That makes binary mode the honest choice when the exact byte layout matters: checksums, protocol framing, or copying a file through unchanged. Note the interaction with a common review habit: reading a file `"rb"` and then decoding by hand is *not* equivalent to text mode, because you have skipped the newline translation as well as the codec. ## Practical guidance For ordinary line-oriented data that you own on both ends, the default is fine and desirable: your code sees `\n` and files from any platform just work. Reach for `newline=""` when a format parser wants raw endings, and for an explicit `"\n"` or `"\r\n"` when the output format specifies one and must be identical regardless of where the job runs — a fixed-format export consumed by another team is the usual case. If you are ever unsure what a file actually contains, open it `"rb"` and look at the raw bytes; text mode is precisely the layer that would hide the answer from you. One last trap worth naming: translation happens on the *text* layer, so it is invisible in the length of the string you wrote but visible in the size of the file on disk. A test that asserts on the length of the `str` and a test that asserts on the file size can disagree on Windows purely because of `newline=`, and that is not a bug in either assertion.

  • Why does writing os.linesep yourself in text mode produce broken output on Windows?
    On Windows `os.linesep` is `"\r\n"`, and text mode with the default `newline=None` translates the `\n` inside it as well, so the file ends up with `\r\r\n`. The correct habit is to write a bare `\n` and let the stream's `newline=` setting produce the platform's bytes, or to pin the ending explicitly with `newline="\r\n"` when the output format requires CRLF everywhere.
  • Does newline= have any effect on a file opened in binary mode?
    No — binary mode never translates line endings, and passing `newline=` there raises `ValueError`. Reading `"rb"` gives you the exact bytes, `\r\n` included. This is worth remembering when someone replaces a text-mode read with a binary read plus a manual decode: they have silently dropped the newline translation too, not just moved the codec.
  • With newline='', what exactly still counts as the end of a line?
    All three of `\r\n`, `\r` and `\n` are still recognised as line breaks, so iteration and `readlines()` split at the same places as with the default. The only difference is that no translation happens, so each returned line keeps the terminator the file actually contained. `newline=""` disables rewriting, not line splitting.

saying these in an interview costs you the question

  • Thinks Python never touches line endings in text mode
  • Reads newline='' as stripping newlines from lines
  • Writes os.linesep explicitly into a text-mode file
  • Believes binary mode still normalises \r\n
  • Thinks translation happens only on Windows machines

context