skip to content

Encoding and Normalization

The trip between code points and bytes and what changes on the way: normalization forms, codec error handlers, the locale default encoding and newline translation. Text bugs land here.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

19

What is the difference between Python's 'utf-8' and 'utf-8-sig' codecs?

level: juniorimportance: must knowfreq 45%

answer

  1. The first field looks right, compares unequal
  2. Spreadsheet exports add an invisible signature
  3. U+FEFF, three bytes in UTF-8
  4. One codec name ends in -sig

basics

~20 s

Python's utf-8 codec treats a leading byte-order mark as an ordinary character, so it decodes to U+FEFF at the front of the string. The utf-8-sig codec strips one leading mark on read and always writes one on encode.

solid answer

~50 s

U+FEFF at the start of a file is the byte-order mark; encoded as UTF-8 it is the three bytes `\xef\xbb\xbf`, which Python exposes as `codecs.BOM_UTF8`. UTF-8 has no byte order to signal, so there the mark is only a *signature* that some spreadsheet and Windows exporters write. The `utf-8` codec is faithful: those three bytes decode to the character U+FEFF and end up glued to your first field, invisibly, so `header[0] == "id"` is `False` while the printed header still looks right. The `utf-8-sig` codec removes exactly one leading signature on decode and decodes identically when there is none — so it is the safe read-side default at any boundary fed by desktop tools. On encode the asymmetry bites: `utf-8-sig` *always* prepends the signature, so write plain `utf-8` unless a specific consumer needs the mark.

code

pycon · 8 lines
pycon
>>> b"\xef\xbb\xbfid,lat".decode("utf-8")
'\ufeffid,lat'
>>> b"\xef\xbb\xbfid,lat".decode("utf-8-sig")
'id,lat'
>>> b"id,lat".decode("utf-8-sig")
'id,lat'
>>> "id,lat".encode("utf-8-sig")
b'\xef\xbb\xbfid,lat'

go deeper

for a junior

Be ready to say what the three bytes at the front of an exported file are and which codec name removes them. Remember that the character is invisible, so repr() is how you see it.

for a middle

Explain the mechanics both ways: decode strips one leading signature and is a no-op without it, encode always adds one. Know that U+FEFF is a format character, so str.strip() will not touch it.

for a senior

An interviewer expects you to place the decision at the boundary: utf-8-sig on ingest, plain utf-8 on output unless a named consumer requires the signature, plus the append trap where a second signature lands mid-file.

for a principal

Own the policy question — which edges of the system are allowed to emit a signature, and how that is enforced. Signature handling belongs in one shared ingest path, not re-decided by each script that reads a partner file.

## The code point and the bytes U+FEFF was originally ZERO WIDTH NO-BREAK SPACE. Unicode then repurposed the same code point as the **byte-order mark**: placed first in a UTF-16 or UTF-32 stream, it lets a reader work out which byte order the encoder used, because those encodings have multi-byte code units and therefore an endianness. UTF-8 has no such ambiguity — its code units are single bytes — so in a UTF-8 file the mark carries no ordering information at all. Some spreadsheet exporters and Windows editors write it anyway, as a **signature** meaning "the bytes that follow are UTF-8". Encoded as UTF-8, U+FEFF is the three bytes `\xef\xbb\xbf`, which Python spells `codecs.BOM_UTF8`. ## What each codec actually does `utf-8` is a faithful UTF-8 codec and nothing more. Decoding `b"\xef\xbb\xbfid,lat"` with it yields the string `'\ufeffid,lat'` — the mark is a valid character, so it decodes to a character. Encoding with `utf-8` never adds one. `utf-8-sig` is the same codec wrapped in signature handling: - **On decode** it removes exactly one leading `codecs.BOM_UTF8` if present, and otherwise decodes byte-for-byte identically to `utf-8`. That makes it *harmless on files that have no mark*. - **On encode** it unconditionally prepends the three bytes. The sentence worth memorising: **`utf-8-sig` is forgiving on input and opinionated on output.** Note the narrowness of the stripping. It removes one signature, only at the very start of the stream. A U+FEFF sitting in the middle of the document stays — correctly, because there it is data, not a signature. ## Why a stray mark hurts so much The character is zero-width. It does not render, most editors do not show it, and a terminal prints `\ufeffid` as something indistinguishable from `id`. Everything downstream that compares or parses that first field then fails in a way that reads like a ghost: - a header row's first name becomes `'\ufeffid'`, so a lookup keyed on `"id"` raises `KeyError` while the printed header looks perfect; - `int(first_field)` raises `ValueError` on what looks like a plain number; - `line.startswith("#")` is `False` for the first line of a file whose first line visibly starts with `#`; - a dict built from the file has a key that prints as `id` and compares unequal to `"id"`; - two files that differ only in whether the exporter wrote a signature produce different keys, so a merge silently drops rows. The diagnosis is mechanical, and knowing it is most of the value of this topic. Print `repr(field)` rather than `field`; check `ord(field[0]) == 0xFEFF`; call `unicodedata.name(field[0])` and see `ZERO WIDTH NO-BREAK SPACE`; or drop to bytes and compare `raw[:3]` against `codecs.BOM_UTF8`. Never diagnose by eye. ## Which codec where **Reading.** At any boundary where files arrive from spreadsheets, desktop editors or partner uploads, decode with `utf-8-sig`. It costs nothing on clean files and removes a whole bug class. Inside your own system, where you produced the file, plain `utf-8` is fine — but `utf-8-sig` is never *wrong* on the read side, which is why it is the better default at an ingest edge. **Writing.** Default to plain `utf-8`. Reach for `utf-8-sig` only where a consumer you do not control uses the signature to choose an encoding — chiefly desktop spreadsheet software, which otherwise opens a UTF-8 export in a legacy code page and shows mojibake. Handing a signed file to line-oriented tooling, to a parser that expects a leading `{` or `[`, or to a script whose first line is a shebang creates the mirror-image bug. One more encode-side trap. `open()` builds an *incremental* encoder that emits the signature once per stream, so appending to an existing file with `encoding="utf-8-sig"` writes a **second** signature into the middle of the file. Likewise, encoding chunk by chunk with `str.encode("utf-8-sig")` and concatenating gives one signature per chunk. Encode the document once, or append with plain `utf-8` after the first write. ## Two facts that settle most interview follow-ups First, the mark is not whitespace: `str.strip()` does not remove it, because U+FEFF is a format character (category `Cf`), not a space. Second, UTF-8 does not *need* a mark for correctness — anything claiming a UTF-8 file must start with one has confused the signature with the byte-order role it plays in UTF-16 and UTF-32.

  • How would you confirm a header field carries a stray U+FEFF rather than something else odd?
    Never by eye — the character is zero width. Print `repr(field)` and look for `\ufeff`, or test `ord(field[0]) == 0xFEFF`, or call `unicodedata.name(field[0])` and expect `ZERO WIDTH NO-BREAK SPACE`. From the file side, open in binary and compare the first three bytes against `codecs.BOM_UTF8`.
  • If a file has no byte-order mark at all, what does decoding it with utf-8-sig cost you?
    Nothing. When the first three bytes are not the signature, `utf-8-sig` decodes byte-for-byte identically to `utf-8`; it only checks and skips those bytes when they match. That is why it is the safe default on the read side. The asymmetry is entirely on encode, where `utf-8-sig` always writes the signature whether or not anyone wants it.
  • When would you deliberately write a file with utf-8-sig?
    When a consumer you do not control uses the signature to pick an encoding — chiefly desktop spreadsheet software, which otherwise opens a UTF-8 export in a legacy code page. Treat it as an interop concession at one specific boundary, not a house style: internal data files, config files and anything read by line-oriented tools should stay plain `utf-8`.

The signature is a sticker on the outside of the envelope saying which alphabet is inside. The utf-8-sig codec peels the sticker off before handing you the letter; the utf-8 codec hands you the letter with the sticker still stuck to the first word.

saying these in an interview costs you the question

  • Claims UTF-8 needs a byte-order mark to state byte order
  • Says the plain utf-8 codec strips the signature automatically
  • Thinks the mark is whitespace that str.strip() removes
  • Assumes utf-8-sig fails or errors on files without a mark
  • Says an invisible character cannot break a key lookup
  • Uses utf-8-sig for every write as a house default

context

open as a page

What does the errors= argument to bytes.decode() control in Python?

level: juniorimportance: must knowfreq 70%

basics

~20 s

It chooses what happens when the input bytes are not valid for the codec. 'strict' is the default and raises UnicodeDecodeError; 'ignore' deletes the bad bytes; 'replace' substitutes U+FFFD; 'backslashreplace' writes an escape for each bad byte.

open as a page

What encoding does Python's open() use when you do not pass encoding=?

level: juniorimportance: must knowfreq 60%

basics

~20 s

A text-mode open() with no encoding= uses the locale encoding that locale.getencoding() reports: usually UTF-8 on Linux and macOS, but the ANSI code page on Windows. It is not guaranteed to be UTF-8, so pass encoding= explicitly.

open as a page

Why can two Python strings that look identical compare unequal with ==?

level: middleimportance: must knowfreq 62%

basics

~20 s

Because == compares code points, not shapes. An accented letter can be one precomposed code point or a base letter plus a combining mark; both render the same but differ. unicodedata.normalize('NFC', s) folds them to one form before comparing.

open as a page

Why must a file handed to Python's csv.writer be opened with newline=''?

level: middleimportance: must knowfreq 45%

basics

~20 s

Because csv.writer emits its own '\r\n' terminator. Left at newline=None, text mode translates the '\n' inside it again — on Windows that yields '\r\r\n' and a blank row between every record. newline='' switches that translation off.

open as a page

What is the difference between str.casefold() and str.lower() in Python?

level: juniorimportance: should knowfreq 42%

basics

~20 s

str.lower() applies simple lowercasing for display. str.casefold() applies Unicode full case folding, which is more aggressive: German ß becomes ss and Greek final sigma becomes sigma. Use casefold for case-insensitive comparison, lower for showing text.

open as a page

What does Python's open() do to line endings when newline is left at its default?

level: juniorimportance: should knowfreq 38%

basics

~10 s

With the default newline=None, open() reads in universal-newline mode: '\r\n', '\r' and lone '\n' all arrive in your string as '\n'. On write, every '\n' you emit is translated to os.linesep.

open as a page

How does Python's surrogateescape handler round-trip undecodable bytes?

level: middleimportance: should knowfreq 35%

basics

~20 s

It maps each undecodable byte to a lone surrogate code point in U+DC80–U+DCFF, so re-encoding that str with errors='surrogateescape' reproduces the original bytes exactly. Python uses it for filenames, argv and environment variables on Unix.

open as a page

Which encoding do sys.stdout and sys.stdin use, and what does PYTHONIOENCODING change?

level: middleimportance: should knowfreq 30%

basics

~10 s

The standard streams are text wrappers that use the locale encoding, or UTF-8 under UTF-8 mode. PYTHONIOENCODING overrides that as encodingname:errorhandler, with either half optional; sys.stdout.reconfigure() changes it from inside a running program.

open as a page

How do you enable Python's UTF-8 mode, and what does it change?

level: middleimportance: should knowfreq 35%

basics

~10 s

Start the interpreter with -X utf8 or set PYTHONUTF8=1. UTF-8 mode ignores the locale and makes UTF-8 the default encoding for open(), the standard streams and the filesystem encoding. Explicit encoding= arguments are unaffected.

open as a page

Stripping a byte-order mark with text[1:] truncated real data in a 6,800-row geocoding batch — how should a Python ingest remove a BOM safely?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Do not slice blindly. Decode with the utf-8-sig codec, which removes a leading mark only when one is there, or guard the removal with str.removeprefix('\ufeff'). Handle it once, at the decode boundary, and keep a no-mark fixture in the tests.

open as a page

An email-digest sender intermittently raises UnicodeEncodeError on some subject lines — how do you diagnose and fix it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Catch the exception and read encoding, reason, start and object[start:end]: that names the exact character and the target codec. Intermittent means data-dependent, not load-dependent, so reproduce with those characters rather than retrying at higher volume.

open as a page

Why does len() on a Python str mislead offset math that slices labels at fixed positions?

level: seniorimportance: should knowfreq 33%

basics

~20 s

len() counts code points, not the characters a reader sees. One visible mark can be several code points, so any offset computed from len() lands mid-character and slicing splits a combining mark from its base letter, shifting every boundary after it.

open as a page

Why can Python's text-mode open() change a file's bytes, and when must you use binary mode?

level: seniorimportance: should knowfreq 28%

basics

~20 s

Text mode both decodes and rewrites line endings, so the characters you read are not the bytes on disk. Whenever the exact bytes matter — checksums, signatures, byte-for-byte copies, non-text formats — open in binary mode.

open as a page

When encoding with Python's 'utf-16' codec, what does it write that 'utf-16-le' does not?

level: middleimportance: nice to knowfreq 18%

basics

~10 s

The 'utf-16' codec writes a two-byte byte-order mark first, in the host's native order, and consumes one when decoding. The endian-suffixed 'utf-16-le' and 'utf-16-be' codecs neither write nor consume a mark.

open as a page

When should you write os.linesep into a Python text file, and why is it usually wrong?

level: middleimportance: nice to knowfreq 20%

basics

~10 s

Almost never in text mode. Text mode already substitutes os.linesep for each '\n' you write, so writing os.linesep yourself gives '\r\r\n' on Windows. It belongs in binary-mode writes, where nothing translates for you.

open as a page

When is unicodedata.normalize with NFKC the right choice, and what does it destroy?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

NFKC applies compatibility folding on top of NFC: ligatures split into letters, superscripts become plain digits, full-width forms become ASCII, no-break space becomes a space. Use it for search and matching keys only — it is lossy and must never overwrite stored text.

open as a page

A fraud-scoring service reads a cached rules file with open() and no encoding=, and one host decodes it as mojibake. How would you find every such call site before the next release?

level: seniorimportance: nice to knowfreq 18%

basics

~10 s

Run the test suite under -X warn_default_encoding (or PYTHONWARNDEFAULTENCODING=1). Every open() or io.TextIOWrapper built without encoding= then raises an EncodingWarning at its own call site; escalate it to an error with -W error::EncodingWarning.

open as a page

How would you set a Python-wide policy for codec error handlers at a system's boundaries?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

Assign a handler per boundary by audience: 'strict' wherever you control the producer, 'surrogateescape' only for OS-supplied values that must round-trip, lossy handlers only on human-facing output, and 'ignore' nowhere. Register a custom handler when you need lossiness counted.

open as a page