skip to content

What does the errors= argument to bytes.decode() control in Python?

level: juniorimportance: must knowfreq 70%

answer

  1. Second argument to the decode call
  2. A recovery policy, not a codec
  3. The default one raises
  4. One handler silently shortens the result
  5. One keeps byte values readable in logs

basics

~20 s

It chooses what happens when the input bytes are not valid for the codec. 'strict' is the default and raises UnicodeDecodeError; 'ignore' deletes the bad bytes; 'replace' substitutes U+FFFD; 'backslashreplace' writes an escape for each bad byte.

solid answer

~40 s

Decoding turns `bytes` into `str`, and `errors=` is the recovery policy applied at every position the codec cannot interpret. The default `'strict'` raises `UnicodeDecodeError`, which carries `encoding`, `object`, `start`, `end` and `reason`, so you can point at the exact failing byte. `'ignore'` silently deletes those bytes, shortening the result and destroying the evidence. `'replace'` substitutes U+FFFD, which is still lossy but visible. `'backslashreplace'` writes `\xNN` for each bad byte, so the original byte values stay readable, which makes it the best choice for logs. `'surrogateescape'` is the only handler that can be reversed. The same argument exists on `str.encode()`, where the failure is `UnicodeEncodeError` and `'xmlcharrefreplace'` and `'namereplace'` become available as well.

code

python · 8 lines
python
data = b"caf\xe9s"
print(data.decode("utf-8", errors="ignore"))
print(data.decode("utf-8", errors="replace"))
print(data.decode("utf-8", errors="backslashreplace"))
try:
    data.decode("utf-8")
except UnicodeDecodeError as exc:
    print(exc.encoding, exc.start, exc.end, exc.reason)

go deeper

for a junior

Be ready to name the four common handlers and say which one is the default. Knowing that decoding goes bytes to str, and that 'strict' raises UnicodeDecodeError, already answers most screening versions of this question.

for a middle

Explain the mechanics: what each handler does to the length and content of the result, that the same argument appears on str.encode() with its own extra handlers, and that open() and the text wrappers accept it too.

for a senior

Show judgment about placement. Say which boundary gets 'strict' and which gets a lossy handler, and demonstrate reading UnicodeDecodeError attributes to identify the failing bytes instead of switching handlers until the traceback stops.

for a principal

Own the policy angle: a lossy handler applied where data will be re-parsed converts a loud failure into silent corruption downstream, so the decision belongs in a shared boundary helper rather than scattered across call sites.

Decoding is the step that turns a `bytes` object into a `str`. The codec named in the call — `utf-8`, `latin-1`, `cp1252` — defines which byte sequences are legal, and real input violates that definition constantly: a file written by another program, a header pasted by a human, a payload cut off in the middle of a multi-byte character. The `errors=` argument to `bytes.decode()` names the **recovery policy** the codec applies at each offending position. It never changes which bytes are legal; it changes only what happens at the moment one is not. ## The decode-side handlers ### `'strict'` `'strict'` is the default, so it is what you get whenever you omit the argument. The codec raises `UnicodeDecodeError`, which inherits from `UnicodeError` and therefore from `ValueError`. The instance carries five attributes worth memorising: - `encoding` (the codec name), - `object` (the entire `bytes` input), - `start` and `end` (the slice that failed) - and `reason` (a phrase such as `invalid continuation byte`). Those attributes are the difference between “it crashed on Unicode” and “byte 0xe9 at offset 3 of this exact payload is not valid UTF-8”. Always slice `exc.object[exc.start:exc.end]` before you theorise about a cause. ### `'ignore'` `'ignore'` deletes the offending bytes. Nothing is raised, nothing is logged, and the resulting `str` is shorter than the input. That silence is why it is the handler most often chosen and most often regretted: offsets shift, a hash over the text no longer matches the source, and any validation that runs after the decode is inspecting a string from which the suspicious bytes have already been removed. ### `'replace'` `'replace'` substitutes U+FFFD REPLACEMENT CHARACTER for each bad sequence. It is equally lossy, but the damage is visible: a person reading the output can see where it happened, and code can count the U+FFFDs to measure how bad an input was. It is a reasonable default for text you are about to show to a human and never parse again. ### `'backslashreplace'` `'backslashreplace'` writes `\xNN` for each undecodable byte, so `b"caf\xe9s"` decodes to the eight-character string `caf\xe9s`. The result is still not the intended text, but it is diagnostic: the byte values survive as readable ASCII, so a log line tells you what the input actually contained. For logs and error messages this beats both `'ignore'` and `'replace'`. It only became usable on the decoding side in Python 3.5; before that it was an encode-only handler. ### `'surrogateescape'` `'surrogateescape'` is the one handler that is **reversible** — it maps each undecodable byte to a lone surrogate code point so that re-encoding with the same handler reproduces the original bytes. ## The encode direction The same `errors=` argument travels in the opposite direction on `str.encode()`, where the failure type is `UnicodeEncodeError` and the failing item is a character the target codec cannot represent — an em dash headed for `ascii`, say. Two handlers exist only on that side: - `'xmlcharrefreplace'` writes a numeric character reference like `—`, - and `'namereplace'` (added in 3.5) writes `\N{EM DASH}`. Handler names are resolved through a process-wide registry, and the built-ins are also exposed as plain functions such as `codecs.replace_errors`, which is occasionally handy when you want to delegate to one from your own handler. ## Every text boundary takes the same argument The argument is not confined to manual `.decode()` calls. `open()`, `io.TextIOWrapper`, `codecs.open` and the text-mode wrappers around subprocesses and sockets all accept the same `errors=` string and pass it straight to the codec, so the choice is available at every text boundary in the program. Those APIs also default to `'strict'`, which is why a program that never mentions `errors=` anywhere still fails loudly the first time it meets a byte its codec rejects. ## Lengths and offsets One consequence worth internalising is that only `'strict'` preserves a **one-to-one relationship** between input and output. - `'ignore'` shortens the result, - `'replace'` collapses a multi-byte sequence into a single U+FFFD, - and `'backslashreplace'` lengthens it to four characters per bad byte. Any code that computes offsets, slices by index, or compares a length against a limit before and after a decode has to know which handler ran. That is also how you measure damage after the fact: counting occurrences of U+FFFD in a decoded string tells you how many sequences the codec could not interpret, information `'ignore'` destroys entirely. ## Choosing well Two habits separate a good answer from a weak one. 1. First, the handler is chosen **per call site**, and the right one depends on the direction of travel and the audience: `'strict'` where you control the producer and want the bug reported, a lossy handler only where the output is for human eyes. 2. Second, `errors=` is **not a fix for the wrong codec**. If you are decoding cp1252 bytes as UTF-8, every handler produces wrong text; `'replace'` merely makes the wrongness quiet. That is the classic interview trap — a candidate who reaches for `errors='ignore'` to make a traceback disappear has usually hidden a mis-declared encoding rather than handled bad input.

  • Which attributes of a caught UnicodeDecodeError do you actually read when debugging?
    `encoding` names the codec that failed, `object` is the whole input, `start` and `end` bound the failing slice, and `reason` explains it in words. The useful move is `exc.object[exc.start:exc.end]`, which hands you the exact bytes; printing them with `ascii()` or hex makes the offending sequence unambiguous instead of a mangled terminal glyph.
  • Why is errors='ignore' a poor default even when it stops the crash?
    It deletes bytes with no signal at all: no exception, no log line, and a shorter result. Offsets and lengths shift, a checksum over the text stops matching, and anything validating the text afterwards is inspecting a string the suspicious bytes have already been removed from. `'replace'` at least leaves a visible marker you can count.
  • Does errors= help if you decoded with the wrong codec entirely?
    No. The handler only decides what to do at bytes the codec rejects; it cannot recover the intended characters. Decoding cp1252 bytes as UTF-8 yields wrong text in every mode, and `'replace'` simply makes it silent. When most of a payload comes out as U+FFFD, suspect the encoding choice rather than tuning the handler.

It is the instruction you leave a proofreader for smudged words: stop and ask, cross them out, write a box, or copy the smudge shape faithfully into the margin.

saying these in an interview costs you the question

  • Claiming errors='ignore' is the safe default for user input
  • Thinking errors= can recover text after choosing the wrong codec
  • Believing 'replace' preserves the original bytes
  • Not knowing 'strict' is the default and raises
  • Confusing UnicodeDecodeError with UnicodeEncodeError
  • Assuming the decoded string is always the same length as the input

context