skip to content

Universal Newline Translation

Text mode quietly rewrites line endings in both directions, which is why the csv module insists on newline='' and why a file can grow doubled carriage returns. Binary mode is how you opt out.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

Why must a file handed to Python's csv.writer be opened with newline=''?

level: middleimportance: must knowfreq 45%

answer

  1. Two layers, one line ending
  2. The writer already ends its rows
  3. Something translates the '\n' a second time
  4. Blank rows appear only on one platform
  5. Quoted fields can contain real line breaks

basics

~20 s

Because csv.writer emits its own '\r\n' terminator. Left at newline=None, text mode translates the '\n' inside it again — on Windows that yields '\r\r\n' and a blank row between every record. newline='' switches that translation off.

solid answer

~40 s

The `csv` module writes complete records, terminator included: its default dialect ends each row with `'\r\n'`. A file opened in text mode with the default `newline=None` then applies its own translation to every `'\n'` it sees, replacing it with `os.linesep`. On Windows that turns the writer's `'\r\n'` into `'\r\r\n'`, and re-reading the file shows a blank line between every record. Passing `newline=''` disables translation so the writer's bytes reach the disk unchanged. The same argument matters on the read side: a quoted field may legitimately contain `'\r\n'`, and without `newline=''` the text layer rewrites those characters inside your data before the parser ever sees them. Binary mode is not an alternative — the module works with `str`, so it needs a text file object with `newline=''`.

code

python · 10 lines
python
import csv
import os
import tempfile
from pathlib import Path

path = os.path.join(tempfile.mkdtemp(), "rows.csv")
with open(path, "w", encoding="utf-8", newline="") as f:
    csv.writer(f).writerow(["id", "title"])

print(Path(path).read_bytes())   # b'id,title\r\n' - the writer's own terminator, untouched

go deeper

for a junior

Recall the rule and apply it: files given to a CSV writer or reader are opened with newline='' and an explicit encoding. Knowing that the symptom is blank rows between records is enough here.

for a middle

Explain the mechanism end to end: the writer emits '\r\n', text mode substitutes os.linesep for the '\n', and the result is '\r\r\n' wherever os.linesep is CRLF. Say why binary mode is not the fix.

for a senior

Demonstrate the diagnosis — inspect bytes, not text — and explain why a Linux-only pipeline never reproduces it. Be ready to generalise the rule to any layer that writes its own terminators.

for a principal

Make it unreproducible: a house rule that every text open() states encoding and newline, plus a fixture or lint check, beats relying on each engineer to remember a per-module footnote across a mixed-platform team.

## Two layers both think they own the line ending A CSV file is line-oriented, and the `csv` module takes that seriously: it writes the record terminator itself. Its default dialect ends every row with `'\r\n'`, the sequence the format's specification calls for. That is a deliberate choice — a CSV file is an interchange format, and CRLF is what most consumers of it expect. Underneath that sits the file object. A file opened in text mode with the default `newline=None` also believes it owns line endings: on write it replaces every `'\n'` in the outgoing string with `os.linesep`. Neither layer knows about the other, so both act. Follow the characters. The writer produces `id,title\r\n`. The text layer scans that string, finds the `'\n'`, and substitutes `os.linesep`. On Linux and macOS `os.linesep` is `'\n'`, so the substitution is a no-op and the file is correct. On Windows `os.linesep` is `'\r\n'`, so the `'\n'` becomes `'\r\n'` and the byte sequence on disk is `\r\r\n`. Note the `'\r'` the writer supplied was never touched — text mode only ever rewrites `'\n'` — it simply survives in front of the newly doubled ending. When something reads that file back in universal-newline mode, `'\r'` and `'\r\n'` are both line terminators, so `\r\r\n` reads as two line breaks. Every record is followed by an empty one. That is the famous "blank rows between every row" symptom, and it is entirely a product of the two layers stacking. ## Why the bug hides The most instructive part is where the bug does *not* appear. On Linux the double translation is invisible, because substituting `'\n'` for `'\n'` changes nothing. So a Linux developer writes the export, the Linux test suite reads it back and passes, the Linux build agent stays green — and the file is broken only for whoever runs the code on Windows. On a mixed team this reaches production as a report of "the export looks double-spaced in my spreadsheet" from the two people whose laptops differ from the build agent. `newline=''` costs nothing and removes the whole class, which is why the correct habit is to write it unconditionally rather than to reason about it per file. ## The read side is the more dangerous half Writing produces cosmetic damage you will eventually notice. Reading produces silent data corruption. CSV allows a field to contain a line break as long as the field is quoted, and such a field routinely carries `'\r\n'` — a description pasted from a Windows editor, for instance. With `newline=None`, the text layer translates those characters *inside the quoted field* to `'\n'` before the parser ever sees them. The parse succeeds, the row count is right, and the field's content has quietly changed. Nothing raises. With `newline=''` the parser receives the characters as written and handles embedded line breaks itself, which is exactly what it is built to do. ## Why not just open it in binary Binary mode would certainly stop the translation, but in Python 3 the module operates on `str`: the writer calls `write()` with text and the reader iterates a text file object. Handing it a binary stream fails. So `newline=''` is not a workaround for the absence of binary mode — it is the supported way to say "give me a text stream that decodes but does not touch line endings". ## The habit to build Open every file destined for a record-oriented writer or parser as: ```python with open(path, "w", encoding="utf-8", newline="") as f: ... ``` Three explicit decisions in one line: the mode, the codec, and who owns the line ending. Do the same on read. If a consumer insists on LF-terminated rows, do not strip the `'\r'` afterwards and do not rely on the platform — configure the writer's terminator, and keep `newline=''` so that what you configured is what lands. The generalisable lesson is bigger than one module: **whenever a layer above the file emits its own terminators, the file must be opened with `newline=''`.** The same reasoning applies to any code that constructs protocol-shaped text — HTTP-style headers, fixed-format exports, anything that must end lines with CRLF by specification rather than by platform. When a file looks wrong, confirm with bytes rather than characters: `pathlib.Path(p).read_bytes()` shows `b'...\r\r\n'` immediately, while `read_text()` hides the very thing you are hunting.

  • Why does this bug survive a green test suite on Linux?
    Because the double translation is a no-op there. With `newline=None` the write side substitutes `os.linesep` for `'\n'`, and on Linux and macOS `os.linesep` *is* `'\n'`, so nothing changes. Only where `os.linesep` is `'\r\n'` does the writer's `'\r\n'` become `'\r\r\n'`. A build agent that matches the developers' platform can never reproduce it.
  • What goes wrong on the reading side if you omit newline=''?
    Quoted fields containing real line breaks get rewritten. A field holding `'line one\r\nline two'` is handed to the parser as `'line one\nline two'`, because the text layer translated the characters inside the field before parsing. The row count and the parse both look fine, so the corruption is silent — worse than the visible blank rows on the write side.
  • Could you open the file in binary mode instead?
    No. In Python 3 the module reads and writes `str`, not `bytes`: the writer calls `write()` with text and the reader iterates a text file object, so a binary stream fails. `newline=''` exists precisely for this shape of problem — a text stream that still decodes with your chosen codec but performs no line-ending translation.
  • How do you confirm the doubled ending is really on disk?
    Read the file as bytes and look: `pathlib.Path(p).read_bytes()` shows `b'id,title\r\r\n'` outright. Reading it as text cannot show you, because the same translation that caused the problem also hides it — the extra `'\r'` is consumed as a line boundary on the way back in.

saying these in an interview costs you the question

  • Says newline='' controls the delimiter or the dialect
  • Strips '\r' from each row instead of fixing the open() call
  • Claims the blank rows come from writerow adding an extra row
  • Thinks opening in binary mode is the fix
  • Believes newline='' is only needed on Windows machines
  • Omits newline='' on read, assuming only writing is affected

context

open as a page

What does Python's open() do to line endings when newline is left at its default?

level: juniorimportance: should knowfreq 38%

basics

~10 s

With the default newline=None, open() reads in universal-newline mode: '\r\n', '\r' and lone '\n' all arrive in your string as '\n'. On write, every '\n' you emit is translated to os.linesep.

open as a page

Why can Python's text-mode open() change a file's bytes, and when must you use binary mode?

level: seniorimportance: should knowfreq 28%

basics

~20 s

Text mode both decodes and rewrites line endings, so the characters you read are not the bytes on disk. Whenever the exact bytes matter — checksums, signatures, byte-for-byte copies, non-text formats — open in binary mode.

open as a page

When should you write os.linesep into a Python text file, and why is it usually wrong?

level: middleimportance: nice to knowfreq 20%

basics

~10 s

Almost never in text mode. Text mode already substitutes os.linesep for each '\n' you write, so writing os.linesep yourself gives '\r\r\n' on Windows. It belongs in binary-mode writes, where nothing translates for you.

open as a page