Why must a file passed to csv.writer or csv.reader be opened with newline=''?
answer
- Two layers both rewrite line endings
- The writer emits its own terminator
- Universal newline translation is the culprit
- On Windows you get a blank row
- Embedded newlines in quoted fields corrupt on read
basics
~20 sThe csv module handles line endings itself. Opening with newline='' switches off the text layer's newline translation so the writer's own terminator is not rewritten and a newline embedded in a quoted field is not corrupted on read.
solid answer
~40 s`csv.writer` emits its dialect's line terminator - `\r\n` for the default `csv.excel` dialect - as part of the record it writes. If the file was opened in text mode with the default `newline=None`, the text layer translates every `\n` it sees into `os.linesep`, so on Windows you get `\r\r\n` and a blank line between every row. Opening with `newline=''` disables that translation and lets the writer's bytes through untouched. On the reading side the same argument matters for a different reason: with translation on, a `\r\n` sitting *inside* a quoted field is rewritten before the parser sees it, silently altering the data. `newline=''` is therefore required in both directions, and it is the single most commonly omitted argument in csv code. The same applies to `io.StringIO(newline='')`.
code
python · 11 linesimport csv
import os
import tempfile
path = os.path.join(tempfile.mkdtemp(), "out.csv")
with open(path, "w", newline="") as f:
csv.writer(f).writerow(["id", "city"])
csv.writer(f).writerow([1, "Paris, FR"])
with open(path, "rb") as f:
print(f.read())go deeper
Learn the canonical line by heart: open the file with newline='' and an explicit encoding whenever csv is involved. Being able to state that the csv module writes its own line endings is enough at this level.
Explain both directions: on write the text layer would translate the writer's \r\n again, and on read it would rewrite a newline sitting inside a quoted field before the parser sees it. Name os.linesep and the newline=None default.
Treat it as a portability defect that hides on your own machine. Show how you would catch it - a round-trip test on the exact bytes, or a byte-level assertion - rather than trusting that output that looks fine locally is fine everywhere.
Frame it as a layering question: encoding and line-ending policy belong to the I/O layer, record structure to the format library, and bugs appear wherever the two overlap. Decide where that convention is enforced so every writer in a codebase does not relearn it.
### Two layers, each with an opinion about line endings When you write CSV there are two independent pieces of machinery that both think line endings are their business, and the whole point of `newline=''` is to make one of them stand down. The first is the csv module. Its dialect carries a `lineterminator`, and for `csv.excel` - the default - that terminator is `\r\n`. `csv.writer(f).writerow(row)` formats the fields, applies quoting, appends the terminator and writes the whole record as one string. The module considers line endings part of the record format it is responsible for. The second is Python's text I/O layer, configured by the `newline=` argument to `open`. Its default, `newline=None`, means *universal newlines*: on input, `\r`, `\n` and `\r\n` are all translated to `\n` before your code sees them; on output, every `\n` you write is translated to `os.linesep`, which is `\r\n` on Windows and `\n` elsewhere. Run both and they fight. ### What goes wrong on write The writer emits `a,b\r\n`. With `newline=None` on Windows the text layer sees the `\n`, translates it to `\r\n`, and the bytes that reach disk are `a,b\r\r\n`. Read that file in a spreadsheet or with another CSV parser and you get a blank row between every real row. The bug is platform-specific, which is why it so often escapes a developer working on macOS or Linux and then surfaces in production or on a colleague's machine. Passing `newline=''` turns translation off: whatever string the writer produces is the string that lands in the file. ### What goes wrong on read The reading case is subtler and, when it bites, worse. A quoted field may contain a line break, and that break is *data*. With `newline=None` the text layer performs universal-newline translation before the csv parser ever runs, so a field containing `\r\n` is handed to the parser as `\n`. The record still parses - nothing raises - but the value you get back is not the value that was written. If that field is round-tripped through your pipeline you have silently rewritten someone's data. There is a second, more visible reading failure. The csv parser needs to see records as the file actually holds them so that it can decide when a quoted field continues onto the next line. Handing it pre-translated text takes that decision away from the parser and gives it to a layer that knows nothing about quoting. `newline=''` keeps the parser in charge, which is the correct division of labour: the text layer does encoding, the csv module does record structure. ### The canonical shape ```python with open(path, "w", newline="", encoding="utf-8") as f: writer = csv.writer(f) writer.writerow(header) with open(path, newline="", encoding="utf-8") as f: for row in csv.reader(f): ... ``` Note `encoding=` alongside it. The two arguments answer different questions - encoding is bytes-to-text, `newline` is line-ending policy - and both belong in the call. Omitting `encoding=` falls back to the platform default, which is how a file written on one machine becomes unreadable on another, so naming it is the portable habit. ### What newline='' is not It is not a request to open the file in binary mode. The csv module works on text: it consumes an iterator of `str` and writes `str`. Handing `csv.reader` a file opened with `"rb"` fails, because the parser gets `bytes` objects it cannot process. If your source really is binary - a byte stream from a network or an archive - wrap it in `io.TextIOWrapper` with `newline=''` and the right encoding, and give the wrapper to the reader. It is also not the only way to control terminators. If you specifically want `\n`-terminated output - common when the file is consumed by Unix tooling - set the dialect's terminator with `lineterminator="\n"` on the writer, or use `csv.unix_dialect`. That is a decision about the file format; `newline=''` remains necessary regardless, because it is what stops the text layer from second-guessing whichever terminator you chose. ### In-memory equivalents The same rule applies to `io.StringIO`, whose own `newline` argument defaults to `'\n'`. When you build CSV in memory - to return from a function, to hand to a network client, to feed a test - construct it as `io.StringIO(newline="")` for exactly the reasons above. Forgetting it in a test is a good way to make the test pass while production is broken.
- What exactly goes wrong on Windows if you omit newline='' when writing?The writer appends `\r\n`, then the text layer translates the `\n` to `os.linesep`, which on Windows is `\r\n`. The file gets `\r\r\n` per record, so consumers show a blank row between every row of data. Nothing raises, and the defect is invisible on macOS or Linux, where `os.linesep` is already `\n`.
- Can you just open the file in binary mode instead?No. The csv module reads and writes `str`, so a file opened with `"rb"` or `"wb"` fails. If the underlying stream really is binary, wrap it in `io.TextIOWrapper` with `newline=''` and an explicit encoding, and pass the wrapper to `csv.reader` or `csv.writer`.
- How do you emit \n-terminated rows rather than \r\n?Set `lineterminator="\n"` on the writer, or use `csv.unix_dialect`, which pairs that terminator with `csv.QUOTE_ALL`. That choice is about the file format you are producing. It does not replace `newline=''` - you still need that so the text layer leaves your chosen terminator alone.
It is like a courier who already sealed and addressed the envelope handing it to a second clerk who insists on re-addressing everything that passes - the fix is to tell the clerk this one is done.
saying these in an interview costs you the question
- Calls newline='' cargo-cult boilerplate with no reason
- Says it only matters on Windows
- Suggests opening the file in binary mode instead
- Thinks newline='' controls the encoding
- Believes it is only needed for writing, not reading
- Omits it for io.StringIO in tests