What is the difference between Python's 'utf-8' and 'utf-8-sig' codecs?
answer
- The first field looks right, compares unequal
- Spreadsheet exports add an invisible signature
- U+FEFF, three bytes in UTF-8
- One codec name ends in -sig
basics
~20 sPython's utf-8 codec treats a leading byte-order mark as an ordinary character, so it decodes to U+FEFF at the front of the string. The utf-8-sig codec strips one leading mark on read and always writes one on encode.
solid answer
~50 sU+FEFF at the start of a file is the byte-order mark; encoded as UTF-8 it is the three bytes `\xef\xbb\xbf`, which Python exposes as `codecs.BOM_UTF8`. UTF-8 has no byte order to signal, so there the mark is only a *signature* that some spreadsheet and Windows exporters write. The `utf-8` codec is faithful: those three bytes decode to the character U+FEFF and end up glued to your first field, invisibly, so `header[0] == "id"` is `False` while the printed header still looks right. The `utf-8-sig` codec removes exactly one leading signature on decode and decodes identically when there is none — so it is the safe read-side default at any boundary fed by desktop tools. On encode the asymmetry bites: `utf-8-sig` *always* prepends the signature, so write plain `utf-8` unless a specific consumer needs the mark.
code
pycon · 8 lines>>> b"\xef\xbb\xbfid,lat".decode("utf-8")
'\ufeffid,lat'
>>> b"\xef\xbb\xbfid,lat".decode("utf-8-sig")
'id,lat'
>>> b"id,lat".decode("utf-8-sig")
'id,lat'
>>> "id,lat".encode("utf-8-sig")
b'\xef\xbb\xbfid,lat'go deeper
Be ready to say what the three bytes at the front of an exported file are and which codec name removes them. Remember that the character is invisible, so repr() is how you see it.
Explain the mechanics both ways: decode strips one leading signature and is a no-op without it, encode always adds one. Know that U+FEFF is a format character, so str.strip() will not touch it.
An interviewer expects you to place the decision at the boundary: utf-8-sig on ingest, plain utf-8 on output unless a named consumer requires the signature, plus the append trap where a second signature lands mid-file.
Own the policy question — which edges of the system are allowed to emit a signature, and how that is enforced. Signature handling belongs in one shared ingest path, not re-decided by each script that reads a partner file.
## The code point and the bytes U+FEFF was originally ZERO WIDTH NO-BREAK SPACE. Unicode then repurposed the same code point as the **byte-order mark**: placed first in a UTF-16 or UTF-32 stream, it lets a reader work out which byte order the encoder used, because those encodings have multi-byte code units and therefore an endianness. UTF-8 has no such ambiguity — its code units are single bytes — so in a UTF-8 file the mark carries no ordering information at all. Some spreadsheet exporters and Windows editors write it anyway, as a **signature** meaning "the bytes that follow are UTF-8". Encoded as UTF-8, U+FEFF is the three bytes `\xef\xbb\xbf`, which Python spells `codecs.BOM_UTF8`. ## What each codec actually does `utf-8` is a faithful UTF-8 codec and nothing more. Decoding `b"\xef\xbb\xbfid,lat"` with it yields the string `'\ufeffid,lat'` — the mark is a valid character, so it decodes to a character. Encoding with `utf-8` never adds one. `utf-8-sig` is the same codec wrapped in signature handling: - **On decode** it removes exactly one leading `codecs.BOM_UTF8` if present, and otherwise decodes byte-for-byte identically to `utf-8`. That makes it *harmless on files that have no mark*. - **On encode** it unconditionally prepends the three bytes. The sentence worth memorising: **`utf-8-sig` is forgiving on input and opinionated on output.** Note the narrowness of the stripping. It removes one signature, only at the very start of the stream. A U+FEFF sitting in the middle of the document stays — correctly, because there it is data, not a signature. ## Why a stray mark hurts so much The character is zero-width. It does not render, most editors do not show it, and a terminal prints `\ufeffid` as something indistinguishable from `id`. Everything downstream that compares or parses that first field then fails in a way that reads like a ghost: - a header row's first name becomes `'\ufeffid'`, so a lookup keyed on `"id"` raises `KeyError` while the printed header looks perfect; - `int(first_field)` raises `ValueError` on what looks like a plain number; - `line.startswith("#")` is `False` for the first line of a file whose first line visibly starts with `#`; - a dict built from the file has a key that prints as `id` and compares unequal to `"id"`; - two files that differ only in whether the exporter wrote a signature produce different keys, so a merge silently drops rows. The diagnosis is mechanical, and knowing it is most of the value of this topic. Print `repr(field)` rather than `field`; check `ord(field[0]) == 0xFEFF`; call `unicodedata.name(field[0])` and see `ZERO WIDTH NO-BREAK SPACE`; or drop to bytes and compare `raw[:3]` against `codecs.BOM_UTF8`. Never diagnose by eye. ## Which codec where **Reading.** At any boundary where files arrive from spreadsheets, desktop editors or partner uploads, decode with `utf-8-sig`. It costs nothing on clean files and removes a whole bug class. Inside your own system, where you produced the file, plain `utf-8` is fine — but `utf-8-sig` is never *wrong* on the read side, which is why it is the better default at an ingest edge. **Writing.** Default to plain `utf-8`. Reach for `utf-8-sig` only where a consumer you do not control uses the signature to choose an encoding — chiefly desktop spreadsheet software, which otherwise opens a UTF-8 export in a legacy code page and shows mojibake. Handing a signed file to line-oriented tooling, to a parser that expects a leading `{` or `[`, or to a script whose first line is a shebang creates the mirror-image bug. One more encode-side trap. `open()` builds an *incremental* encoder that emits the signature once per stream, so appending to an existing file with `encoding="utf-8-sig"` writes a **second** signature into the middle of the file. Likewise, encoding chunk by chunk with `str.encode("utf-8-sig")` and concatenating gives one signature per chunk. Encode the document once, or append with plain `utf-8` after the first write. ## Two facts that settle most interview follow-ups First, the mark is not whitespace: `str.strip()` does not remove it, because U+FEFF is a format character (category `Cf`), not a space. Second, UTF-8 does not *need* a mark for correctness — anything claiming a UTF-8 file must start with one has confused the signature with the byte-order role it plays in UTF-16 and UTF-32.
- How would you confirm a header field carries a stray U+FEFF rather than something else odd?Never by eye — the character is zero width. Print `repr(field)` and look for `\ufeff`, or test `ord(field[0]) == 0xFEFF`, or call `unicodedata.name(field[0])` and expect `ZERO WIDTH NO-BREAK SPACE`. From the file side, open in binary and compare the first three bytes against `codecs.BOM_UTF8`.
- If a file has no byte-order mark at all, what does decoding it with utf-8-sig cost you?Nothing. When the first three bytes are not the signature, `utf-8-sig` decodes byte-for-byte identically to `utf-8`; it only checks and skips those bytes when they match. That is why it is the safe default on the read side. The asymmetry is entirely on encode, where `utf-8-sig` always writes the signature whether or not anyone wants it.
- When would you deliberately write a file with utf-8-sig?When a consumer you do not control uses the signature to pick an encoding — chiefly desktop spreadsheet software, which otherwise opens a UTF-8 export in a legacy code page. Treat it as an interop concession at one specific boundary, not a house style: internal data files, config files and anything read by line-oriented tools should stay plain `utf-8`.
The signature is a sticker on the outside of the envelope saying which alphabet is inside. The utf-8-sig codec peels the sticker off before handing you the letter; the utf-8 codec hands you the letter with the sticker still stuck to the first word.
saying these in an interview costs you the question
- Claims UTF-8 needs a byte-order mark to state byte order
- Says the plain utf-8 codec strips the signature automatically
- Thinks the mark is whitespace that str.strip() removes
- Assumes utf-8-sig fails or errors on files without a mark
- Says an invisible character cannot break a key lookup
- Uses utf-8-sig for every write as a house default