skip to content

Strings and Bytes

The text side of Python: the line Python 3 draws between str and bytes, the string methods you reach for daily, and how to build text without paying a quadratic cost. Encoding bugs come up constantly.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

22

What is the difference between Python's str and bytes types?

level: juniorimportance: must knowfreq 85%

answer

  1. Two sequences that never mix
  2. One means text, one means data
  3. Conversion always names an encoding
  4. encode goes text to octets
  5. decode goes octets to text

basics

~20 s

str holds Unicode code points, meaning text; bytes holds raw octets valued 0 to 255, meaning data. Python 3 never converts between them implicitly: str.encode() produces bytes and bytes.decode() produces text, and each names an encoding.

solid answer

~40 s

`str` is an immutable sequence of Unicode code points — abstract characters with no storage layout. `bytes` is an immutable sequence of octets, each an integer 0-255, with no inherent meaning until you say how to read it; `bytearray` is the mutable version of the same thing. The bridge between them is always a codec: `str.encode(encoding)` turns text into octets, `bytes.decode(encoding)` turns octets back into text. Both default to UTF-8, but naming the encoding explicitly at every boundary is the habit that prevents surprises. Python 3 removed all implicit coercion, so `"a" == b"a"` is `False` and concatenating the two raises `TypeError`. The practical rule is decode at the edge, work in `str` inside, encode on the way out.

code

pycon · 12 lines
pycon
>>> s = "héllo"
>>> len(s)
5
>>> b = s.encode("utf-8")
>>> b
b'h\xc3\xa9llo'
>>> len(b)
6
>>> b.decode("utf-8") == s
True
>>> "a" == b"a"
False

go deeper

for a junior

Be ready to state the distinction in one sentence and name the two methods in the right direction: str.encode() gives bytes, bytes.decode() gives str. Knowing that len() counts code points on one and octets on the other already puts you ahead.

for a middle

Explain the mechanics: why UTF-8 makes a str and its encoding different lengths, why encode and decode exist on only one type each, and why comparing str to bytes returns False instead of raising. Show where you would put the conversion in a real codebase.

for a senior

An interviewer expects you to talk about the boundary as a design rule — decode once at the edge, work in str, encode on the way out — and to describe the bugs that appear when a value is sometimes str and sometimes bytes, which typically only show up on non-ASCII input.

for a principal

Own the convention across services: which layer decodes, how encodings are declared and carried between components, and how you keep the raw octets available for hashing or signing while the rest of the code sees text. The tradeoff is one clear boundary versus per-module improvisation.

## Two different sequences A `str` in Python 3 is an immutable sequence of **Unicode code points**. A code point is a number assigned by the Unicode standard to an abstract character — `ord("A")` is 65, `ord("€")` is 8364. A `str` says nothing about how those numbers are stored; CPython picks a compact internal representation for you and you never see it. Indexing a `str` gives you another one-character `str`, and `len()` counts code points. A `bytes` object is an immutable sequence of **octets** — integers in the range 0-255. It carries no notion of characters at all. `len()` counts octets, and the repr is only *displayed* with ASCII-looking escapes for readability; `b"A"` is the single number 65, not a letter. `bytearray` is the same sequence type with mutation allowed. The crucial design decision of Python 3 is that these two types never mix implicitly. There is no automatic promotion, no shared comparison, no concatenation: ```python "a" == b"a" # False "a" + b"a" # TypeError "a".decode() # AttributeError: str has no decode b"a".encode() # AttributeError: bytes has no encode ``` The last two are worth memorising as a shape: `encode` only exists on text (it produces data), `decode` only exists on data (it produces text). If you reach for the wrong one, the exception is telling you that you were confused about which side of the boundary you were standing on. ## The codec is the bridge An **encoding** (a codec) is a rule mapping code points to octet sequences and back. UTF-8 is the one to choose unless something external forces another: it is variable-width, ASCII-compatible for the first 128 code points, and covers the whole Unicode range. ```python s = "café" len(s) # 4 code points b = s.encode("utf-8") # b'caf\xc3\xa9' len(b) # 5 octets b.decode("utf-8") == s # True ``` That length difference is the single most useful thing to internalise. Anything non-ASCII costs more than one octet in UTF-8, so a `str` length and the length of its encoding are different numbers. Column limits, buffer sizes, database `VARCHAR` lengths and progress counters all have to say which of the two they mean. Both `str.encode()` and `bytes.decode()` default to UTF-8 when you omit the argument, and that default does not depend on the platform. Writing the encoding out anyway is still worth it: it documents the assumption at the point where it is made, and it survives the code being moved into a helper that reads from somewhere else. ## Why anything hands you bytes at all Octets are what actually travel and what actually persists. A socket read gives you `bytes` because the network moves octets and the peer's charset is not carried in the data itself — at best it is declared out of band, in an HTTP header or a file format's own header. A file opened with mode `"rb"` gives you `bytes` for the same reason: the disk stores octets. Hashing, compression and cryptographic signing all operate on octets, so those APIs take and return `bytes`. So every real program has a boundary, and the discipline is to make it explicit and thin: 1. Receive `bytes` at the edge. 2. Decode once, with an encoding you obtained or chose deliberately. 3. Do all business logic in `str`. 4. Encode once on the way out. The alternative — letting the two types travel together through the middle of an application — is where the classic bugs live: a value that is sometimes `str` and sometimes `bytes` depending on which code path produced it, and a `TypeError` or a mangled output that only appears for non-ASCII input, which is exactly the input your test fixtures never contain. ## Constructing each type `bytes` literals use a `b` prefix and may contain only ASCII characters plus `\xNN` escapes. You can build one from a list of integers with `bytes([104, 105])`, and `bytes(3)` gives three zero octets rather than the digit. `str(b"hi")` does *not* decode — it produces the string `"b'hi'"`, which is almost never what anyone wants; `b"hi".decode("utf-8")` is the real conversion. ## What an interviewer is checking They want to hear that you know the two types are not interchangeable, that conversion is always a codec operation, and that you have a rule for where in your program the conversion happens. Candidates who came from Python 2 sometimes still describe `str` as "bytes with characters in it"; that model was true in Python 2, where `str` was octets and `unicode` was text with implicit ASCII coercion between them, and it is the source of most confusion about this topic today.

  • What encoding does str.encode() use when you omit the argument?
    UTF-8, on every platform — the default of `str.encode()` and `bytes.decode()` is fixed, not taken from the environment. Naming it explicitly is still the better habit, because it documents the assumption where it is made and keeps the call correct if the surrounding code is later reused for a source that is not UTF-8.
  • Why does "a" == b"a" evaluate to False rather than raising?
    Equality between unrelated types is defined to return `False` rather than to fail, so the comparison is legal but never true. That makes it a quiet bug: a dictionary keyed by `str` will silently miss every `bytes` lookup. Ordering the two, or concatenating them, does raise `TypeError`, which is why mixed-type bugs often surface only on the branch that concatenates.
  • Where in an application should the decode happen?
    At the edge, once. Read `bytes` from the socket or the binary file, decode with an encoding you obtained from a header or chose deliberately, then keep everything internal as `str` and encode again only when writing out. A value that is sometimes `str` and sometimes `bytes` in the middle of the code is the shape most encoding bugs take.

A str is the sentence you mean; bytes is the ink on the page. Going between them requires an alphabet — the encoding — and using the wrong alphabet gives you nonsense rather than an error.

saying these in an interview costs you the question

  • Describes bytes as just a string of ASCII characters
  • Thinks Python converts str to bytes automatically when needed
  • Assumes len() of a str equals its size in octets
  • Calls encode() on bytes or decode() on str
  • Uses str(some_bytes) expecting it to decode
  • Believes a str can compare equal to a bytes literal

context

open as a page

What is the difference between an f-string, str.format and %-formatting in Python?

level: juniorimportance: must knowfreq 80%

basics

~20 s

All three build a str. An f-string is a literal whose expressions are evaluated where it is written, so it cannot be stored as a reusable template; str.format and the % operator format a template string supplied at call time.

open as a page

Why does calling s.upper() on a Python string leave s unchanged?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Python str objects are immutable, so no method can change one in place. str.upper() builds and returns a brand-new string and leaves the original untouched, which is why you must assign the result: s = s.upper().

open as a page

What is the difference between bytes and bytearray in Python?

level: juniorimportance: must knowfreq 62%

basics

~20 s

bytes is an immutable sequence of octets; bytearray is the mutable version you can append to and edit in place. Only bytes is hashable, so only bytes works as a dict key or set member.

open as a page

Why doesn't 'scores.csv'.strip('.csv') simply remove the file extension?

level: juniorimportance: must knowfreq 62%

basics

~10 s

str.strip takes a set of characters, not a suffix. It peels any of '.', 'c', 's' or 'v' off both ends, so 'scores.csv'.strip('.csv') returns 'ore'. Use str.removesuffix('.csv'), added in Python 3.9.

open as a page

When does Python raise UnicodeDecodeError rather than UnicodeEncodeError?

level: middleimportance: must knowfreq 60%

basics

~20 s

UnicodeDecodeError comes from bytes.decode(): the octets are not valid under the codec you named. UnicodeEncodeError comes from str.encode(): the codec has no representation for a character you have. Decode is octets to text; encode is text to octets.

open as a page

Why is `s += part` in a Python loop unsafe to rely on for speed?

level: middleimportance: must knowfreq 60%

basics

~20 s

CPython sometimes resizes the left-hand string in place when the loop variable holds the only reference, so the loop looks linear. That is an unguaranteed implementation detail: add one more reference and every iteration copies everything accumulated so far.

open as a page

What does memoryview give you that slicing a bytes object does not?

level: middleimportance: must knowfreq 46%

basics

~20 s

Slicing bytes copies the sliced region into a new object; memoryview wraps the original memory, so slicing a view costs nothing but a small window object. A view of a bytearray can also write straight through to the source.

open as a page

How does str.split() with no argument differ from str.split(' ')?

level: middleimportance: must knowfreq 66%

basics

~20 s

With no argument, str.split treats any run of whitespace as one separator and discards leading and trailing whitespace, so it never yields empty strings. With ' ' it splits on each single space, so runs of spaces produce empty strings.

open as a page

What does the Python 3.14 t-string literal t"Hi {name}" evaluate to, and how does it differ from an f-string?

level: juniorimportance: should knowfreq 30%

basics

~20 s

A t-string evaluates to a string.templatelib.Template object rather than a str. It stores the literal text and the interpolated values separately instead of joining them, so a renderer can process each value before any final string exists.

open as a page

In Python's format mini-language, what does the spec in f"{total:>12,.2f}" mean?

level: middleimportance: should knowfreq 55%

basics

~20 s

Everything after the colon is the format spec: right-align in a field 12 characters wide, group thousands with commas, show two digits after the decimal point, and render as fixed-point. It formats -1234.5678 as " -1,234.57".

open as a page

How do str.find and str.index differ when the substring is absent?

level: middleimportance: should knowfreq 48%

basics

~20 s

str.find returns -1 when the substring is absent; str.index raises ValueError instead. Both return the index of the first occurrence otherwise. For a yes/no test use the in operator, never the truthiness of find's result.

open as a page

How do you walk a string.templatelib.Template to render it with per-value escaping?

level: middleimportance: should knowfreq 22%

basics

~20 s

Iterate the Template: it yields the literal string chunks and Interpolation objects in source order, skipping empty chunks. Append the literal chunks unchanged, format and escape each interpolation's value for the target syntax, then join the pieces into one str.

open as a page

Why can verifying a webhook signature after decoding the body to str fail?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A signature is computed over octets, and decoding a body to str then re-encoding it is not guaranteed to reproduce the exact octets that arrived. Verify against the raw bytes the connection delivered, then decode afterwards with an explicitly named encoding.

open as a page

Why is calling str.format on a user-supplied template string dangerous?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A hostile template walks attributes of the arguments you pass, so "{row.init.globals[SECRET]}" reads module globals and leaks secrets, and a huge width such as "{v:>100000000}" allocates hundreds of megabytes per render. Render untrusted templates with string.Template instead.

open as a page

When do you reach for io.StringIO instead of ''.join() to build a large Python string?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Use ''.join() when every fragment can be collected in a list first — it allocates the result once. Use io.StringIO when the text comes from scattered write() calls, or when you must rewind and discard a half-written section.

open as a page

A document-conversion worker keeps a memoryview slice of each payload; why does memory never drop?

level: seniorimportance: should knowfreq 33%

basics

~20 s

A memoryview keeps a reference to the object it views, so an eight-byte slice pins the entire source buffer for as long as it is retained. Copy the field out with tobytes() and let the payload be freed.

open as a page

Why does b'abc'[0] give 97 while b'abc'[0:1] gives b'a' in Python 3?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

A bytes object is a sequence of integers 0 to 255, so indexing one element yields an int, while slicing preserves the container type and yields bytes. str is the unusual sequence whose elements are themselves str.

open as a page

What do !r, !s and = do inside an f-string replacement field in Python?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

They sit before the colon. !s calls str(), !r calls repr() and !a calls ascii() on the value before formatting. A trailing = is the debug form: it prints the expression source, an equals sign, then the value's repr.

open as a page

What do int.to_bytes and int.from_bytes do, and why does the byteorder argument matter?

level: middleimportance: nice to knowfreq 24%

basics

~20 s

int.to_bytes renders an integer as a fixed-width bytes value and int.from_bytes reads one back. The byteorder argument, "big" or "little", picks which end holds the most significant byte; getting it wrong silently yields a different number rather than an error.

open as a page

When is str.translate with str.maketrans better than chained str.replace calls?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

str.translate applies a whole per-character mapping in one pass, so replacements never cascade into each other and cost stays linear in the text. Chained str.replace calls each scan the string again and can re-transform an earlier call's output.

open as a page

What does a Python 3.14 t-string actually defer, and what does an API typed to accept Template gain over str?

level: seniorimportance: nice to knowfreq 15%

basics

~20 s

Only the joining is deferred: each interpolated expression is evaluated eagerly, as in an f-string. What survives is the boundary, so a Template parameter tells the callee which text is source literal and which is a value.

open as a page