skip to content

str vs bytes and Encodings

str holds Unicode code points and bytes holds raw octets, and every conversion between them names an encoding. UnicodeDecodeError in production almost always traces to an assumed default.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

What is the difference between Python's str and bytes types?

level: juniorimportance: must knowfreq 85%

answer

  1. Two sequences that never mix
  2. One means text, one means data
  3. Conversion always names an encoding
  4. encode goes text to octets
  5. decode goes octets to text

basics

~20 s

str holds Unicode code points, meaning text; bytes holds raw octets valued 0 to 255, meaning data. Python 3 never converts between them implicitly: str.encode() produces bytes and bytes.decode() produces text, and each names an encoding.

solid answer

~40 s

`str` is an immutable sequence of Unicode code points — abstract characters with no storage layout. `bytes` is an immutable sequence of octets, each an integer 0-255, with no inherent meaning until you say how to read it; `bytearray` is the mutable version of the same thing. The bridge between them is always a codec: `str.encode(encoding)` turns text into octets, `bytes.decode(encoding)` turns octets back into text. Both default to UTF-8, but naming the encoding explicitly at every boundary is the habit that prevents surprises. Python 3 removed all implicit coercion, so `"a" == b"a"` is `False` and concatenating the two raises `TypeError`. The practical rule is decode at the edge, work in `str` inside, encode on the way out.

code

pycon · 12 lines
pycon
>>> s = "héllo"
>>> len(s)
5
>>> b = s.encode("utf-8")
>>> b
b'h\xc3\xa9llo'
>>> len(b)
6
>>> b.decode("utf-8") == s
True
>>> "a" == b"a"
False

go deeper

for a junior

Be ready to state the distinction in one sentence and name the two methods in the right direction: str.encode() gives bytes, bytes.decode() gives str. Knowing that len() counts code points on one and octets on the other already puts you ahead.

for a middle

Explain the mechanics: why UTF-8 makes a str and its encoding different lengths, why encode and decode exist on only one type each, and why comparing str to bytes returns False instead of raising. Show where you would put the conversion in a real codebase.

for a senior

An interviewer expects you to talk about the boundary as a design rule — decode once at the edge, work in str, encode on the way out — and to describe the bugs that appear when a value is sometimes str and sometimes bytes, which typically only show up on non-ASCII input.

for a principal

Own the convention across services: which layer decodes, how encodings are declared and carried between components, and how you keep the raw octets available for hashing or signing while the rest of the code sees text. The tradeoff is one clear boundary versus per-module improvisation.

## Two different sequences A `str` in Python 3 is an immutable sequence of **Unicode code points**. A code point is a number assigned by the Unicode standard to an abstract character — `ord("A")` is 65, `ord("€")` is 8364. A `str` says nothing about how those numbers are stored; CPython picks a compact internal representation for you and you never see it. Indexing a `str` gives you another one-character `str`, and `len()` counts code points. A `bytes` object is an immutable sequence of **octets** — integers in the range 0-255. It carries no notion of characters at all. `len()` counts octets, and the repr is only *displayed* with ASCII-looking escapes for readability; `b"A"` is the single number 65, not a letter. `bytearray` is the same sequence type with mutation allowed. The crucial design decision of Python 3 is that these two types never mix implicitly. There is no automatic promotion, no shared comparison, no concatenation: ```python "a" == b"a" # False "a" + b"a" # TypeError "a".decode() # AttributeError: str has no decode b"a".encode() # AttributeError: bytes has no encode ``` The last two are worth memorising as a shape: `encode` only exists on text (it produces data), `decode` only exists on data (it produces text). If you reach for the wrong one, the exception is telling you that you were confused about which side of the boundary you were standing on. ## The codec is the bridge An **encoding** (a codec) is a rule mapping code points to octet sequences and back. UTF-8 is the one to choose unless something external forces another: it is variable-width, ASCII-compatible for the first 128 code points, and covers the whole Unicode range. ```python s = "café" len(s) # 4 code points b = s.encode("utf-8") # b'caf\xc3\xa9' len(b) # 5 octets b.decode("utf-8") == s # True ``` That length difference is the single most useful thing to internalise. Anything non-ASCII costs more than one octet in UTF-8, so a `str` length and the length of its encoding are different numbers. Column limits, buffer sizes, database `VARCHAR` lengths and progress counters all have to say which of the two they mean. Both `str.encode()` and `bytes.decode()` default to UTF-8 when you omit the argument, and that default does not depend on the platform. Writing the encoding out anyway is still worth it: it documents the assumption at the point where it is made, and it survives the code being moved into a helper that reads from somewhere else. ## Why anything hands you bytes at all Octets are what actually travel and what actually persists. A socket read gives you `bytes` because the network moves octets and the peer's charset is not carried in the data itself — at best it is declared out of band, in an HTTP header or a file format's own header. A file opened with mode `"rb"` gives you `bytes` for the same reason: the disk stores octets. Hashing, compression and cryptographic signing all operate on octets, so those APIs take and return `bytes`. So every real program has a boundary, and the discipline is to make it explicit and thin: 1. Receive `bytes` at the edge. 2. Decode once, with an encoding you obtained or chose deliberately. 3. Do all business logic in `str`. 4. Encode once on the way out. The alternative — letting the two types travel together through the middle of an application — is where the classic bugs live: a value that is sometimes `str` and sometimes `bytes` depending on which code path produced it, and a `TypeError` or a mangled output that only appears for non-ASCII input, which is exactly the input your test fixtures never contain. ## Constructing each type `bytes` literals use a `b` prefix and may contain only ASCII characters plus `\xNN` escapes. You can build one from a list of integers with `bytes([104, 105])`, and `bytes(3)` gives three zero octets rather than the digit. `str(b"hi")` does *not* decode — it produces the string `"b'hi'"`, which is almost never what anyone wants; `b"hi".decode("utf-8")` is the real conversion. ## What an interviewer is checking They want to hear that you know the two types are not interchangeable, that conversion is always a codec operation, and that you have a rule for where in your program the conversion happens. Candidates who came from Python 2 sometimes still describe `str` as "bytes with characters in it"; that model was true in Python 2, where `str` was octets and `unicode` was text with implicit ASCII coercion between them, and it is the source of most confusion about this topic today.

  • What encoding does str.encode() use when you omit the argument?
    UTF-8, on every platform — the default of `str.encode()` and `bytes.decode()` is fixed, not taken from the environment. Naming it explicitly is still the better habit, because it documents the assumption where it is made and keeps the call correct if the surrounding code is later reused for a source that is not UTF-8.
  • Why does "a" == b"a" evaluate to False rather than raising?
    Equality between unrelated types is defined to return `False` rather than to fail, so the comparison is legal but never true. That makes it a quiet bug: a dictionary keyed by `str` will silently miss every `bytes` lookup. Ordering the two, or concatenating them, does raise `TypeError`, which is why mixed-type bugs often surface only on the branch that concatenates.
  • Where in an application should the decode happen?
    At the edge, once. Read `bytes` from the socket or the binary file, decode with an encoding you obtained from a header or chose deliberately, then keep everything internal as `str` and encode again only when writing out. A value that is sometimes `str` and sometimes `bytes` in the middle of the code is the shape most encoding bugs take.

A str is the sentence you mean; bytes is the ink on the page. Going between them requires an alphabet — the encoding — and using the wrong alphabet gives you nonsense rather than an error.

saying these in an interview costs you the question

  • Describes bytes as just a string of ASCII characters
  • Thinks Python converts str to bytes automatically when needed
  • Assumes len() of a str equals its size in octets
  • Calls encode() on bytes or decode() on str
  • Uses str(some_bytes) expecting it to decode
  • Believes a str can compare equal to a bytes literal

context

open as a page

When does Python raise UnicodeDecodeError rather than UnicodeEncodeError?

level: middleimportance: must knowfreq 60%

basics

~20 s

UnicodeDecodeError comes from bytes.decode(): the octets are not valid under the codec you named. UnicodeEncodeError comes from str.encode(): the codec has no representation for a character you have. Decode is octets to text; encode is text to octets.

open as a page

Why can verifying a webhook signature after decoding the body to str fail?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A signature is computed over octets, and decoding a body to str then re-encoding it is not guaranteed to reproduce the exact octets that arrived. Verify against the raw bytes the connection delivered, then decode afterwards with an explicitly named encoding.

open as a page

Why does b'abc'[0] give 97 while b'abc'[0:1] gives b'a' in Python 3?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

A bytes object is a sequence of integers 0 to 255, so indexing one element yields an int, while slicing preserves the container type and yields bytes. str is the unusual sequence whose elements are themselves str.

open as a page