skip to content

When encoding with Python's 'utf-16' codec, what does it write that 'utf-16-le' does not?

level: middleimportance: nice to knowfreq 18%

answer

  1. Two-byte code units need an order
  2. The endian-suffixed names stay silent
  3. codecs.BOM_UTF16 follows sys.byteorder
  4. Append twice, get two marks

basics

~10 s

The 'utf-16' codec writes a two-byte byte-order mark first, in the host's native order, and consumes one when decoding. The endian-suffixed 'utf-16-le' and 'utf-16-be' codecs neither write nor consume a mark.

solid answer

~40 s

UTF-16 has two-byte code units, so a reader needs to know the byte order. Python's `utf-16` codec handles that for you: on encode it emits U+FEFF first, in the machine's native order (`sys.byteorder`), then the text; on decode it reads a leading mark, uses it to pick the order, and does not hand the mark back to you. The suffixed `utf-16-le` and `utf-16-be` codecs fix the order in the name and do neither job, so decoding a marked stream with `utf-16-le` leaves a stray `'\ufeff'` at the front of your string. `codecs.BOM_UTF16` is the two-byte native mark; `codecs.BOM_UTF16_LE` and `codecs.BOM_UTF16_BE` are the fixed forms. UTF-32 splits the same way, with a four-byte mark.

code

python · 11 lines
python
import codecs
import sys

print(sys.byteorder)
print("ok".encode("utf-16"))     # mark + code units, native order
print("ok".encode("utf-16-le"))  # no mark

print(codecs.BOM_UTF16 == codecs.BOM_UTF16_LE)  # True on a little-endian host

print(repr(b"\xff\xfeo\x00k\x00".decode("utf-16")))     # 'ok'
print(repr(b"\xff\xfeo\x00k\x00".decode("utf-16-le")))  # '\ufeffok'

go deeper

for a junior

Recall that the codec name tells you the behaviour: the plain name manages the mark, the -le and -be names do not. That alone explains a stray invisible character at the front of decoded text.

for a middle

Be ready to explain both directions — write native-order mark then units, read the mark to pick the order and drop it — and why UTF-8 has no equivalent byte-order problem.

for a senior

Show the operational traps: appending with the unsuffixed codec plants a mark mid-file, a missing mark decodes silently in native order, and a byte sniffer must test the four-byte marks before the two-byte ones.

for a principal

The judgement call is whether the file format should be self-describing at all. Pinning byte order in the contract and using the suffixed codecs removes a class of host-dependent output; carrying the mark buys interoperability with tools you do not own.

## Why UTF-16 needs a mark and UTF-8 does not UTF-16 stores text in two-byte code units, and UTF-32 in four-byte units. A multi-byte unit can be laid down low byte first (little-endian) or high byte first (big-endian), and the bytes alone do not say which. That is the problem the **byte-order mark** solves: encode U+FEFF at the very start, and a reader that sees `\xff\xfe` knows the stream is little-endian while one that sees `\xfe\xff` knows it is big-endian. The choice is unambiguous because U+FFFE is not a valid character, so the reversed reading is never a legitimate document. UTF-8 has single-byte code units and therefore no endianness at all, which is why the same three bytes there are only a *signature*, not a byte-order mark. ## The two families of codec name Python exposes both roles, and the name tells you which you get. **`utf-16` (unsuffixed) manages the mark for you.** - On encode it writes U+FEFF first, in the host's native order — `sys.byteorder` decides — and then the code units in that same order. - On decode it looks at the first two bytes, uses a mark to select the byte order, and strips it, so the mark never reaches your string. If there is no mark, CPython falls back to the host's native order, which is exactly why a file produced on the other endianness and stripped of its mark decodes to mojibake rather than raising. **`utf-16-le` and `utf-16-be` fix the order in the name and do nothing else.** - On encode they write no mark at all. - On decode they consume no mark, so `b"\xff\xfeo\x00k\x00".decode("utf-16-le")` gives `'\ufeffok'`: a U+FEFF glued to the front of your data, the same invisible passenger the UTF-8 signature produces. `utf-32`, `utf-32-le` and `utf-32-be` split identically, with a four-byte mark. ## The constants `codecs` names all of them. `codecs.BOM_UTF16_LE` is `b"\xff\xfe"`, `codecs.BOM_UTF16_BE` is `b"\xfe\xff"`, and `codecs.BOM_UTF16` is whichever of the two matches `sys.byteorder`. `codecs.BOM_UTF32_LE` is `b"\xff\xfe\x00\x00"`, `codecs.BOM_UTF32_BE` is `b"\x00\x00\xfe\xff"`, and `codecs.BOM_UTF32` is again the native one. One consequence matters when you sniff a file's first bytes: the UTF-32 little-endian mark *begins with* the UTF-16 little-endian mark, so a sniffer must test the four-byte forms before the two-byte forms or it will call every UTF-32 file UTF-16. ## Where this bites in practice **Appending.** `open()` creates a fresh incremental encoder per stream, and the unsuffixed codec writes its mark at the start of that stream. Open a file in append mode with `encoding="utf-16"` on a second run and you get a second mark part-way through the file, which the reader hands back as a stray U+FEFF in the middle of a line rather than as ordering information. The fix is to write the file once with `utf-16`, or to append with `utf-16-le`/`utf-16-be` matching what the first write produced. **Chunked encoding.** The same asymmetry appears if you call `str.encode("utf-16")` on each chunk and concatenate: one mark per chunk. Encode the whole document once, or encode chunks with the suffixed codec after emitting the mark yourself. **Cross-platform files.** Because the unsuffixed encoder follows the *encoding* machine's native order, the same code produces different bytes on different hosts. That is fine as long as the mark travels with the file; it stops being fine the moment some tool strips the first two bytes, at which point the file only reads correctly on hosts of the original endianness. **Sizing assumptions.** A common companion misconception is that UTF-16 is two bytes per character. It is two bytes per *code unit*; characters outside the Basic Multilingual Plane occupy a surrogate pair, four bytes. So neither the mark nor the unit size makes byte counting a proxy for character counting. ## The rule of thumb Use the unsuffixed `utf-16` when you are reading or writing a whole self-describing file and want the mark handled for you. Use `utf-16-le` or `utf-16-be` when an external contract fixes the byte order — a wire format, a fixed-layout record, an appended stream — and you want the codec to add nothing you did not ask for. ## Round-tripping is not automatic A detail that catches people: `"ok".encode("utf-16").decode("utf-16")` round-trips, and so does `"ok".encode("utf-16-le").decode("utf-16-le")`, but crossing the families does not. Encode with the unsuffixed codec and decode with the suffixed one and you gain a leading U+FEFF; encode with the suffixed codec and decode with the unsuffixed one and the first two bytes of real text are read as a byte-order mark if they happen to look like one, or silently interpreted in native order if they do not. Neither mismatch raises. So the codec name is part of the file's contract, not an implementation detail one side can change: whoever writes and whoever reads must agree on which of the two families is in play, and the cheapest way to make that agreement visible is to record it beside the format rather than infer it per site.

  • What does Python's 'utf-16' decoder do when the input carries no mark at all?
    It falls back to the host's native byte order rather than raising. So bytes produced on a big-endian machine, with the mark stripped somewhere in transit, decode on a little-endian host into valid-looking but wrong characters. That silence is the argument for either keeping the mark or pinning the order explicitly with `utf-16-le`/`utf-16-be` as part of the format contract.
  • How does UTF-32 differ from UTF-16 here?
    Not at all in structure, only in width. The unsuffixed `utf-32` codec writes a four-byte mark in native order and consumes one on read; `utf-32-le` and `utf-32-be` do neither. `codecs.BOM_UTF32` is the four-byte native form. The one gotcha is that the UTF-32 little-endian mark starts with the UTF-16 little-endian mark, so byte-sniffing must test four bytes before two.

saying these in an interview costs you the question

  • Thinks 'utf-16' and 'utf-16-le' encode identical bytes
  • Believes 'utf-16-le' strips a leading mark on decode
  • Says UTF-8 output needs a mark for byte order
  • Assumes appending re-uses the first stream's mark
  • Treats UTF-16 as a fixed two bytes per character
  • Sniffs the two-byte mark before the four-byte one

context