What is the difference between Python's str and bytes types?
answer
- Two sequences that never mix
- One means text, one means data
- Conversion always names an encoding
- encode goes text to octets
- decode goes octets to text
basics
~20 sstr holds Unicode code points, meaning text; bytes holds raw octets valued 0 to 255, meaning data. Python 3 never converts between them implicitly: str.encode() produces bytes and bytes.decode() produces text, and each names an encoding.
solid answer
~40 s`str` is an immutable sequence of Unicode code points — abstract characters with no storage layout. `bytes` is an immutable sequence of octets, each an integer 0-255, with no inherent meaning until you say how to read it; `bytearray` is the mutable version of the same thing. The bridge between them is always a codec: `str.encode(encoding)` turns text into octets, `bytes.decode(encoding)` turns octets back into text. Both default to UTF-8, but naming the encoding explicitly at every boundary is the habit that prevents surprises. Python 3 removed all implicit coercion, so `"a" == b"a"` is `False` and concatenating the two raises `TypeError`. The practical rule is decode at the edge, work in `str` inside, encode on the way out.
code
pycon · 12 lines>>> s = "héllo"
>>> len(s)
5
>>> b = s.encode("utf-8")
>>> b
b'h\xc3\xa9llo'
>>> len(b)
6
>>> b.decode("utf-8") == s
True
>>> "a" == b"a"
Falsego deeper
Be ready to state the distinction in one sentence and name the two methods in the right direction: str.encode() gives bytes, bytes.decode() gives str. Knowing that len() counts code points on one and octets on the other already puts you ahead.
Explain the mechanics: why UTF-8 makes a str and its encoding different lengths, why encode and decode exist on only one type each, and why comparing str to bytes returns False instead of raising. Show where you would put the conversion in a real codebase.
An interviewer expects you to talk about the boundary as a design rule — decode once at the edge, work in str, encode on the way out — and to describe the bugs that appear when a value is sometimes str and sometimes bytes, which typically only show up on non-ASCII input.
Own the convention across services: which layer decodes, how encodings are declared and carried between components, and how you keep the raw octets available for hashing or signing while the rest of the code sees text. The tradeoff is one clear boundary versus per-module improvisation.
## Two different sequences A `str` in Python 3 is an immutable sequence of **Unicode code points**. A code point is a number assigned by the Unicode standard to an abstract character — `ord("A")` is 65, `ord("€")` is 8364. A `str` says nothing about how those numbers are stored; CPython picks a compact internal representation for you and you never see it. Indexing a `str` gives you another one-character `str`, and `len()` counts code points. A `bytes` object is an immutable sequence of **octets** — integers in the range 0-255. It carries no notion of characters at all. `len()` counts octets, and the repr is only *displayed* with ASCII-looking escapes for readability; `b"A"` is the single number 65, not a letter. `bytearray` is the same sequence type with mutation allowed. The crucial design decision of Python 3 is that these two types never mix implicitly. There is no automatic promotion, no shared comparison, no concatenation: ```python "a" == b"a" # False "a" + b"a" # TypeError "a".decode() # AttributeError: str has no decode b"a".encode() # AttributeError: bytes has no encode ``` The last two are worth memorising as a shape: `encode` only exists on text (it produces data), `decode` only exists on data (it produces text). If you reach for the wrong one, the exception is telling you that you were confused about which side of the boundary you were standing on. ## The codec is the bridge An **encoding** (a codec) is a rule mapping code points to octet sequences and back. UTF-8 is the one to choose unless something external forces another: it is variable-width, ASCII-compatible for the first 128 code points, and covers the whole Unicode range. ```python s = "café" len(s) # 4 code points b = s.encode("utf-8") # b'caf\xc3\xa9' len(b) # 5 octets b.decode("utf-8") == s # True ``` That length difference is the single most useful thing to internalise. Anything non-ASCII costs more than one octet in UTF-8, so a `str` length and the length of its encoding are different numbers. Column limits, buffer sizes, database `VARCHAR` lengths and progress counters all have to say which of the two they mean. Both `str.encode()` and `bytes.decode()` default to UTF-8 when you omit the argument, and that default does not depend on the platform. Writing the encoding out anyway is still worth it: it documents the assumption at the point where it is made, and it survives the code being moved into a helper that reads from somewhere else. ## Why anything hands you bytes at all Octets are what actually travel and what actually persists. A socket read gives you `bytes` because the network moves octets and the peer's charset is not carried in the data itself — at best it is declared out of band, in an HTTP header or a file format's own header. A file opened with mode `"rb"` gives you `bytes` for the same reason: the disk stores octets. Hashing, compression and cryptographic signing all operate on octets, so those APIs take and return `bytes`. So every real program has a boundary, and the discipline is to make it explicit and thin: 1. Receive `bytes` at the edge. 2. Decode once, with an encoding you obtained or chose deliberately. 3. Do all business logic in `str`. 4. Encode once on the way out. The alternative — letting the two types travel together through the middle of an application — is where the classic bugs live: a value that is sometimes `str` and sometimes `bytes` depending on which code path produced it, and a `TypeError` or a mangled output that only appears for non-ASCII input, which is exactly the input your test fixtures never contain. ## Constructing each type `bytes` literals use a `b` prefix and may contain only ASCII characters plus `\xNN` escapes. You can build one from a list of integers with `bytes([104, 105])`, and `bytes(3)` gives three zero octets rather than the digit. `str(b"hi")` does *not* decode — it produces the string `"b'hi'"`, which is almost never what anyone wants; `b"hi".decode("utf-8")` is the real conversion. ## What an interviewer is checking They want to hear that you know the two types are not interchangeable, that conversion is always a codec operation, and that you have a rule for where in your program the conversion happens. Candidates who came from Python 2 sometimes still describe `str` as "bytes with characters in it"; that model was true in Python 2, where `str` was octets and `unicode` was text with implicit ASCII coercion between them, and it is the source of most confusion about this topic today.
- What encoding does str.encode() use when you omit the argument?UTF-8, on every platform — the default of `str.encode()` and `bytes.decode()` is fixed, not taken from the environment. Naming it explicitly is still the better habit, because it documents the assumption where it is made and keeps the call correct if the surrounding code is later reused for a source that is not UTF-8.
- Why does "a" == b"a" evaluate to False rather than raising?Equality between unrelated types is defined to return `False` rather than to fail, so the comparison is legal but never true. That makes it a quiet bug: a dictionary keyed by `str` will silently miss every `bytes` lookup. Ordering the two, or concatenating them, does raise `TypeError`, which is why mixed-type bugs often surface only on the branch that concatenates.
- Where in an application should the decode happen?At the edge, once. Read `bytes` from the socket or the binary file, decode with an encoding you obtained from a header or chose deliberately, then keep everything internal as `str` and encode again only when writing out. A value that is sometimes `str` and sometimes `bytes` in the middle of the code is the shape most encoding bugs take.
A str is the sentence you mean; bytes is the ink on the page. Going between them requires an alphabet — the encoding — and using the wrong alphabet gives you nonsense rather than an error.
saying these in an interview costs you the question
- Describes bytes as just a string of ASCII characters
- Thinks Python converts str to bytes automatically when needed
- Assumes len() of a str equals its size in octets
- Calls encode() on bytes or decode() on str
- Uses str(some_bytes) expecting it to decode
- Believes a str can compare equal to a bytes literal