skip to content

questions

3

What is the difference between Python's 'utf-8' and 'utf-8-sig' codecs?

level: juniorimportance: must knowfreq 45%

answer

  1. The first field looks right, compares unequal
  2. Spreadsheet exports add an invisible signature
  3. U+FEFF, three bytes in UTF-8
  4. One codec name ends in -sig

basics

~20 s

Python's utf-8 codec treats a leading byte-order mark as an ordinary character, so it decodes to U+FEFF at the front of the string. The utf-8-sig codec strips one leading mark on read and always writes one on encode.

solid answer

~50 s

U+FEFF at the start of a file is the byte-order mark; encoded as UTF-8 it is the three bytes `\xef\xbb\xbf`, which Python exposes as `codecs.BOM_UTF8`. UTF-8 has no byte order to signal, so there the mark is only a *signature* that some spreadsheet and Windows exporters write. The `utf-8` codec is faithful: those three bytes decode to the character U+FEFF and end up glued to your first field, invisibly, so `header[0] == "id"` is `False` while the printed header still looks right. The `utf-8-sig` codec removes exactly one leading signature on decode and decodes identically when there is none — so it is the safe read-side default at any boundary fed by desktop tools. On encode the asymmetry bites: `utf-8-sig` *always* prepends the signature, so write plain `utf-8` unless a specific consumer needs the mark.

code

pycon · 8 lines
pycon
>>> b"\xef\xbb\xbfid,lat".decode("utf-8")
'\ufeffid,lat'
>>> b"\xef\xbb\xbfid,lat".decode("utf-8-sig")
'id,lat'
>>> b"id,lat".decode("utf-8-sig")
'id,lat'
>>> "id,lat".encode("utf-8-sig")
b'\xef\xbb\xbfid,lat'

go deeper

for a junior

Be ready to say what the three bytes at the front of an exported file are and which codec name removes them. Remember that the character is invisible, so repr() is how you see it.

for a middle

Explain the mechanics both ways: decode strips one leading signature and is a no-op without it, encode always adds one. Know that U+FEFF is a format character, so str.strip() will not touch it.

for a senior

An interviewer expects you to place the decision at the boundary: utf-8-sig on ingest, plain utf-8 on output unless a named consumer requires the signature, plus the append trap where a second signature lands mid-file.

for a principal

Own the policy question — which edges of the system are allowed to emit a signature, and how that is enforced. Signature handling belongs in one shared ingest path, not re-decided by each script that reads a partner file.

## The code point and the bytes U+FEFF was originally ZERO WIDTH NO-BREAK SPACE. Unicode then repurposed the same code point as the **byte-order mark**: placed first in a UTF-16 or UTF-32 stream, it lets a reader work out which byte order the encoder used, because those encodings have multi-byte code units and therefore an endianness. UTF-8 has no such ambiguity — its code units are single bytes — so in a UTF-8 file the mark carries no ordering information at all. Some spreadsheet exporters and Windows editors write it anyway, as a **signature** meaning "the bytes that follow are UTF-8". Encoded as UTF-8, U+FEFF is the three bytes `\xef\xbb\xbf`, which Python spells `codecs.BOM_UTF8`. ## What each codec actually does `utf-8` is a faithful UTF-8 codec and nothing more. Decoding `b"\xef\xbb\xbfid,lat"` with it yields the string `'\ufeffid,lat'` — the mark is a valid character, so it decodes to a character. Encoding with `utf-8` never adds one. `utf-8-sig` is the same codec wrapped in signature handling: - **On decode** it removes exactly one leading `codecs.BOM_UTF8` if present, and otherwise decodes byte-for-byte identically to `utf-8`. That makes it *harmless on files that have no mark*. - **On encode** it unconditionally prepends the three bytes. The sentence worth memorising: **`utf-8-sig` is forgiving on input and opinionated on output.** Note the narrowness of the stripping. It removes one signature, only at the very start of the stream. A U+FEFF sitting in the middle of the document stays — correctly, because there it is data, not a signature. ## Why a stray mark hurts so much The character is zero-width. It does not render, most editors do not show it, and a terminal prints `\ufeffid` as something indistinguishable from `id`. Everything downstream that compares or parses that first field then fails in a way that reads like a ghost: - a header row's first name becomes `'\ufeffid'`, so a lookup keyed on `"id"` raises `KeyError` while the printed header looks perfect; - `int(first_field)` raises `ValueError` on what looks like a plain number; - `line.startswith("#")` is `False` for the first line of a file whose first line visibly starts with `#`; - a dict built from the file has a key that prints as `id` and compares unequal to `"id"`; - two files that differ only in whether the exporter wrote a signature produce different keys, so a merge silently drops rows. The diagnosis is mechanical, and knowing it is most of the value of this topic. Print `repr(field)` rather than `field`; check `ord(field[0]) == 0xFEFF`; call `unicodedata.name(field[0])` and see `ZERO WIDTH NO-BREAK SPACE`; or drop to bytes and compare `raw[:3]` against `codecs.BOM_UTF8`. Never diagnose by eye. ## Which codec where **Reading.** At any boundary where files arrive from spreadsheets, desktop editors or partner uploads, decode with `utf-8-sig`. It costs nothing on clean files and removes a whole bug class. Inside your own system, where you produced the file, plain `utf-8` is fine — but `utf-8-sig` is never *wrong* on the read side, which is why it is the better default at an ingest edge. **Writing.** Default to plain `utf-8`. Reach for `utf-8-sig` only where a consumer you do not control uses the signature to choose an encoding — chiefly desktop spreadsheet software, which otherwise opens a UTF-8 export in a legacy code page and shows mojibake. Handing a signed file to line-oriented tooling, to a parser that expects a leading `{` or `[`, or to a script whose first line is a shebang creates the mirror-image bug. One more encode-side trap. `open()` builds an *incremental* encoder that emits the signature once per stream, so appending to an existing file with `encoding="utf-8-sig"` writes a **second** signature into the middle of the file. Likewise, encoding chunk by chunk with `str.encode("utf-8-sig")` and concatenating gives one signature per chunk. Encode the document once, or append with plain `utf-8` after the first write. ## Two facts that settle most interview follow-ups First, the mark is not whitespace: `str.strip()` does not remove it, because U+FEFF is a format character (category `Cf`), not a space. Second, UTF-8 does not *need* a mark for correctness — anything claiming a UTF-8 file must start with one has confused the signature with the byte-order role it plays in UTF-16 and UTF-32.

  • How would you confirm a header field carries a stray U+FEFF rather than something else odd?
    Never by eye — the character is zero width. Print `repr(field)` and look for `\ufeff`, or test `ord(field[0]) == 0xFEFF`, or call `unicodedata.name(field[0])` and expect `ZERO WIDTH NO-BREAK SPACE`. From the file side, open in binary and compare the first three bytes against `codecs.BOM_UTF8`.
  • If a file has no byte-order mark at all, what does decoding it with utf-8-sig cost you?
    Nothing. When the first three bytes are not the signature, `utf-8-sig` decodes byte-for-byte identically to `utf-8`; it only checks and skips those bytes when they match. That is why it is the safe default on the read side. The asymmetry is entirely on encode, where `utf-8-sig` always writes the signature whether or not anyone wants it.
  • When would you deliberately write a file with utf-8-sig?
    When a consumer you do not control uses the signature to pick an encoding — chiefly desktop spreadsheet software, which otherwise opens a UTF-8 export in a legacy code page. Treat it as an interop concession at one specific boundary, not a house style: internal data files, config files and anything read by line-oriented tools should stay plain `utf-8`.

The signature is a sticker on the outside of the envelope saying which alphabet is inside. The utf-8-sig codec peels the sticker off before handing you the letter; the utf-8 codec hands you the letter with the sticker still stuck to the first word.

saying these in an interview costs you the question

  • Claims UTF-8 needs a byte-order mark to state byte order
  • Says the plain utf-8 codec strips the signature automatically
  • Thinks the mark is whitespace that str.strip() removes
  • Assumes utf-8-sig fails or errors on files without a mark
  • Says an invisible character cannot break a key lookup
  • Uses utf-8-sig for every write as a house default

context

open as a page

Stripping a byte-order mark with text[1:] truncated real data in a 6,800-row geocoding batch — how should a Python ingest remove a BOM safely?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Do not slice blindly. Decode with the utf-8-sig codec, which removes a leading mark only when one is there, or guard the removal with str.removeprefix('\ufeff'). Handle it once, at the decode boundary, and keep a no-mark fixture in the tests.

open as a page

When encoding with Python's 'utf-16' codec, what does it write that 'utf-16-le' does not?

level: middleimportance: nice to knowfreq 18%

basics

~10 s

The 'utf-16' codec writes a two-byte byte-order mark first, in the host's native order, and consumes one when decoding. The endian-suffixed 'utf-16-le' and 'utf-16-be' codecs neither write nor consume a mark.

open as a page