skip to content

When does Python raise UnicodeDecodeError rather than UnicodeEncodeError?

level: middleimportance: must knowfreq 60%

answer

  1. The verb in the name is the operation
  2. One direction reads octets, one writes them
  3. Both descend from ValueError
  4. Coverage fails one way, validity the other
  5. One codec accepts every possible octet

basics

~20 s

UnicodeDecodeError comes from bytes.decode(): the octets are not valid under the codec you named. UnicodeEncodeError comes from str.encode(): the codec has no representation for a character you have. Decode is octets to text; encode is text to octets.

solid answer

~40 s

Both are subclasses of `UnicodeError`, which is a `ValueError`, and the half of the name before `Error` tells you which direction failed. `bytes.decode("utf-8")` raises `UnicodeDecodeError` when the octets are not a well-formed UTF-8 sequence — an invalid start byte, a truncated multi-octet sequence, a continuation byte where one is not expected. `str.encode("ascii")` raises `UnicodeEncodeError` when the target codec simply cannot express a character you have, such as a euro sign in ASCII. The message carries the codec name, the offending position and a reason, which is usually enough to identify the real encoding. Note the asymmetry: latin-1 decodes *any* octet sequence without complaint because all 256 values map to code points, so a wrong-codec bug there produces mojibake instead of an exception.

code

pycon · 10 lines
pycon
>>> b"\xff\xfe".decode("utf-8")
Traceback (most recent call last):
  ...
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte
>>> b"\xff\xfe".decode("latin-1")
'ÿþ'
>>> "€".encode("ascii")
Traceback (most recent call last):
  ...
UnicodeEncodeError: 'ascii' codec can't encode character '\u20ac' in position 0: ordinal not in range(128)

go deeper

for a junior

Learn the direction cold: decode turns octets into text and fails when the octets are invalid, encode turns text into octets and fails when the target codec cannot represent a character. Being able to say which of the two you just saw is most of the answer.

for a middle

Explain the mechanics: what makes a UTF-8 sequence invalid, why ASCII fails on a euro sign, and why latin-1 never raises on decode. Mention that both errors subclass ValueError and that the message carries the codec, position and reason.

for a senior

Show that you treat the exception as evidence about the source. Walk through identifying which boundary produced the octets, whether the encoding is declared out of band, and why choosing a strategy for malformed input before you know the real encoding turns a loud failure into silent corruption.

for a principal

Own the policy: what your systems do when input does not match its declared encoding, whether that is rejection, quarantine or repair, and who is accountable for the producer. The tradeoff is availability against silently accepting corrupted text into long-lived storage.

## The name tells you the direction Python has exactly two codec failures worth memorising, and the verb in each name is the operation that failed. - `UnicodeDecodeError` is raised by **decode**: `bytes` in, `str` out. It means the octets you were handed are not valid input for the codec you named. - `UnicodeEncodeError` is raised by **encode**: `str` in, `bytes` out. It means the codec you named has no way to represent a code point you gave it. Both inherit from `UnicodeError`, which inherits from `ValueError`, so a broad `except ValueError` will swallow them — usually to your cost, because the message is the diagnostic. ## Decode failures UTF-8 is a self-validating format: not every octet sequence is legal. Decoding fails on - an **invalid start byte** — an octet such as `0xFF` that begins no valid sequence; - **unexpected end of data** — a multi-octet sequence cut short, which is what you get when a stream is chunked and you decode a partial buffer; - an **invalid continuation byte** — the second octet of a sequence not having the expected high bits, which is the usual symptom of text that is actually in some single-octet legacy encoding. ```python b"\xff\xfe".decode("utf-8") # UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte ``` The exception message names the codec, the offending octet, its position and the reason. In production that position is the lead: if it is 0, you probably have the wrong codec entirely or a byte-order mark you did not expect; if it is deep into a large payload, you more likely have one bad record in an otherwise clean stream. The important asymmetry is that **some codecs cannot fail on decode at all**. latin-1 maps each of the 256 possible octet values onto the code points U+0000 to U+00FF, so every byte string decodes successfully. That is why guessing latin-1 never raises and never helps: you get plausible-looking garbage — the *mojibake* where `café` shows up as two characters instead of one — instead of an exception pointing at the real problem. It is also why latin-1 is occasionally used deliberately as a lossless octet-preserving round trip when a layer insists on handing you text. ## Encode failures Going the other way, the failure is about coverage. ASCII covers 128 code points; latin-1 covers 256. Ask either to represent a euro sign, a CJK character or an emoji and it cannot: ```python "€".encode("ascii") # UnicodeEncodeError: 'ascii' codec can't encode character '\u20ac' in position 0: ordinal not in range(128) ``` UTF-8 covers the whole Unicode range, so an encode to UTF-8 almost never fails — with one real exception. **Lone surrogates**, code points in U+D800-U+DFFF, exist only as a UTF-16 encoding mechanism and are not valid characters; UTF-8 refuses them. You can end up holding one when text arrived through a layer that decoded UTF-16 sloppily, or through a filesystem API that smuggles undecodable octets into a `str`. Encoding such a value to UTF-8 raises `UnicodeEncodeError` with the reason "surrogates not allowed", and it is one of the genuinely puzzling production failures on this topic because most people believe UTF-8 encoding cannot fail. ## Diagnosing, not suppressing A codec exception is almost always a **wrong assumption about the source**, not a defect in the data. Work the question backwards: 1. Which boundary produced these octets — a socket, a file opened in binary mode, a subprocess pipe, a database driver? 2. Does that source *declare* its encoding out of band, in an HTTP header, an XML declaration or a file-format header? If yes, use the declaration rather than a guess. 3. If nothing declares it, whose encoding is it in practice? A single upstream producer with a legacy default is the common answer, and once identified the fix is one explicit codec name at the boundary. 4. Only then decide what to do about genuinely malformed input. The anti-pattern is turning the exception off and shipping. Whatever the strategy for bad input, choosing it before you know the source's real encoding converts a loud failure with an exact byte offset into a silent data-corruption bug that surfaces weeks later in someone else's report. ## What an interviewer is checking That you can say the direction of each error without hesitating, that you know the message contains the position and reason, and that you treat the exception as evidence about the source rather than an obstacle. The candidate who immediately proposes suppressing the error has told you they will corrupt data quietly.

  • Why does decoding arbitrary octets with the latin-1 codec never raise?
    latin-1 maps every one of the 256 possible octet values to a code point in U+0000-U+00FF, so no octet sequence is invalid. That makes it useless as a guess — you get mojibake rather than a diagnosis — but useful as a deliberate lossless round trip when a layer insists on handing you text and you need the original octets back.
  • Can encoding a str to UTF-8 ever fail?
    Yes, on lone surrogates in U+D800-U+DFFF. They are not real characters, only a UTF-16 mechanism, and UTF-8 refuses them with `UnicodeEncodeError: ... surrogates not allowed`. You acquire one from a layer that decoded UTF-16 badly or from a filesystem API smuggling undecodable octets into a str, and it is the one case where the usual assumption that UTF-8 covers everything breaks.
  • What does the position in the error message tell you?
    Where in the input the codec gave up. Position 0 usually means the whole payload is in a different encoding, or begins with an unexpected byte-order mark. A position deep inside a large payload more often means one bad record in an otherwise clean stream, which points at a specific producer rather than a wrong assumption about the whole source.

saying these in an interview costs you the question

  • Cannot say which direction each error name refers to
  • Claims encoding to UTF-8 can never fail
  • Thinks decoding as latin-1 fixes the encoding
  • Silently discards bad octets instead of finding the real encoding
  • Catches ValueError broadly and loses the diagnostic message
  • Assumes the data is corrupt rather than the assumed codec wrong

context