skip to content

When should you write os.linesep into a Python text file, and why is it usually wrong?

level: middleimportance: nice to knowfreq 20%

answer

  1. A constant describing the host, not a setting
  2. Something already substitutes it for you
  3. Adding it yourself doubles a character
  4. Only safe where nothing translates
  5. The newline argument is the real knob

basics

~10 s

Almost never in text mode. Text mode already substitutes os.linesep for each '\n' you write, so writing os.linesep yourself gives '\r\r\n' on Windows. It belongs in binary-mode writes, where nothing translates for you.

solid answer

~40 s

`os.linesep` is a string describing the platform's native line separator: `'\n'` on Linux and macOS, `'\r\n'` on Windows. The trap is that text mode has already used it for you — with the default `newline=None`, every `'\n'` you write is replaced by `os.linesep` on the way out. So `f.write('x' + os.linesep)` on Windows emits `'\r'` untouched followed by a translated `'\n'`, i.e. `'\r\r\n'`. In text mode you write plain `'\n'` and let the layer decide, or you take control explicitly with `newline='\r\n'` when a format demands CRLF regardless of platform. `os.linesep` is genuinely useful only where no translation happens: assembling bytes for a binary-mode write, or interpreting output from another process. It is also the wrong thing to split on when reading, since universal-newline mode has already normalised everything to `'\n'`.

code

python · 15 lines
python
import os
import tempfile
from pathlib import Path

path = os.path.join(tempfile.mkdtemp(), "log.txt")

# Wrong: on Windows text mode turns the '\n' inside os.linesep into '\r\n' again.
with open(path, "w", encoding="utf-8") as f:
    f.write("ready" + os.linesep)

# Right: ask open() for the ending you want and keep writing plain '\n'.
with open(path, "w", encoding="utf-8", newline="\r\n") as f:
    f.write("ready\n")

print(repr(os.linesep), Path(path).read_bytes())

go deeper

for a junior

Remember the simple habit: in text mode you write '\n' and let the file object handle the platform. Recognise os.linesep as a description of the host rather than something you paste into output.

for a middle

Explain the doubling: text mode substitutes os.linesep for your '\n' while leaving your '\r' alone, so supplying it yourself yields '\r\r\n' on Windows. Name newline='\r\n' as the correct explicit control.

for a senior

Recognise the pattern in review, and note that a same-platform test suite cannot catch it. Be able to say where the constant is legitimate — binary writes and statements about the host — and why splitting on it is always wrong.

for a principal

Decide the policy: formats specify their line ending, hosts do not. Encoding that as a convention plus a review or lint rule removes a whole family of platform-only defects that reproduce for only part of the team.

## What the constant actually is `os.linesep` is a plain string holding the line separator the current platform uses natively: `'\n'` on Linux and macOS, `'\r\n'` on Windows. It is a fact about the host, evaluated at import time, and nothing more. It is not a setting, it does not influence `open()`, and changing it would change nothing about how files behave. Its name invites a natural but wrong inference: that portable code should write `os.linesep` instead of hard-coding `'\n'`. In text mode that reasoning is exactly backwards, and the constant's own documentation says so. ## Why writing it in text mode doubles up Recall the write-side rule for a text-mode file with the default `newline=None`: every `'\n'` in the string you pass to `write()` is replaced with `os.linesep`. Nothing else is touched — a `'\r'` you supply passes straight through. Now write `'ready' + os.linesep` on Windows. The string you hand over is `'ready\r\n'`. The layer leaves the `'\r'` alone and expands the `'\n'` to `'\r\n'`. What lands on disk is `ready\r\r\n`. Read that back in universal-newline mode and the doubled carriage return reads as an extra line break, so the file appears double-spaced. On Linux the same code is harmless, because `os.linesep` is `'\n'` and the expansion is a no-op — which is precisely why the mistake ships: it is invisible to whoever wrote it and to a same-platform test suite. The correct text-mode idiom is to write `'\n'` and nothing else. That is what the abstraction is for: `'\n'` is the in-memory representation of "end of line", and the file object converts at the boundary. ## When you do want to be explicit Sometimes the platform's preference is irrelevant because a *format* dictates the ending — an interchange file that must be CRLF-terminated for its consumers, or protocol-shaped text. The knob for that is the `newline` argument, not the constant: ```python with open(path, "w", encoding="utf-8", newline="\r\n") as f: f.write("first\nsecond\n") # b'first\r\nsecond\r\n' everywhere ``` You keep writing `'\n'` in the code and state the wire format once, at the boundary, where a reader can see it. The inverse — `newline=''` — says "I have already put the exact endings in my string, do not touch them", which is the right choice when a layer above you emits its own terminators. ## Where os.linesep earns its keep There are legitimate uses, and they share a property: no translation layer is present. * **Binary-mode writes.** `open(path, 'wb')` does nothing for you, so if you are assembling bytes that must end lines the native way, `os.linesep.encode()` is the honest expression of that intent. * **Talking about the host.** Code that reports or reasons about the platform — building a message, choosing a default for a config file it will hand to a native tool — can legitimately consult it. Even there, pause: most formats want a specific ending rather than the local one, and "native" is often a weaker requirement than it first sounds. ## Never split on it A related misuse is reading: `text.split(os.linesep)`. In text mode with the default `newline=None`, everything the file contained — `'\r\n'`, `'\r'`, `'\n'` — has already been normalised to `'\n'` in the string, so on Windows you would be splitting on a two-character sequence that no longer exists in the data and getting one giant line. Iterate the file object, or split on `'\n'`, or use `str.splitlines()` if you want the generous Unicode definition of a break. ## Reading is where the constant misleads twice Beyond splitting, people reach for `os.linesep` when comparing: checking whether a line "ends properly", or stripping it before further work. Both are unnecessary and both are wrong for the same reason. In text mode the file object already told you what the ending was by turning it into `'\n'`; the original sequence is no longer in the string to compare against. If you genuinely need to know what a file used on disk, do not consult the host constant at all - open the file in binary and look at the bytes, which is a question about that file rather than about the machine you happen to be running on. This is the deeper point: `os.linesep` describes the host, and almost every real question - what does this format require, what did this file use, what does this consumer expect - is about something other than the host. ## The rule in one line In text mode, write `'\n'` and choose the ending with `newline=`. In binary mode there is no translation, so any ending you want is yours to write — and `os.linesep` is finally the right constant. Treat every appearance of `os.linesep` in a text-mode write as a defect until proven otherwise; in a mixed-platform codebase it is one of the few one-line changes that fixes a bug nobody on the majority platform can reproduce.

  • What does os.linesep contain on each major platform?
    `'\n'` on Linux and on macOS, and `'\r\n'` on Windows. It reflects the host running the code, not the file being written and not the machine that will read it — which is the core reason it is a poor basis for choosing a file format's line ending.
  • How do you guarantee CRLF output regardless of the platform you run on?
    Open the file with `newline='\r\n'` and keep writing plain `'\n'` in your code. The file object performs the substitution on every platform, so the output is identical from Linux, macOS and Windows. Do not concatenate `'\r\n'` into the strings yourself unless you also pass `newline=''` to stop the second translation.
  • Is os.linesep the right separator to split a file's contents on?
    No. In text mode with the default `newline=None` the file's endings were already normalised to `'\n'` in the string, so splitting on `'\r\n'` under Windows matches nothing and returns one enormous line. Iterate the file object, split on `'\n'`, or use `str.splitlines()` when you want the broader Unicode set of breaks.

It is like writing the country code on a letter you have already handed to a service that stamps the country code on for you — harmless where the code is empty, duplicated everywhere else.

saying these in an interview costs you the question

  • Calls writing os.linesep the portable way to end a line
  • Thinks os.linesep configures how open() translates
  • Splits decoded text on os.linesep after a text-mode read
  • Believes os.linesep is '\r' on macOS
  • Cannot say why the doubling is invisible on Linux

context