skip to content

Why does b'abc'[0] give 97 while b'abc'[0:1] gives b'a' in Python 3?

level: middleimportance: nice to knowfreq 22%

answer

  1. One of the two follows the general rule
  2. The elements are not characters
  3. Slices keep the container's type
  4. Only one sequence type is self-similar
  5. Indexing yields a number, slicing an object

basics

~20 s

A bytes object is a sequence of integers 0 to 255, so indexing one element yields an int, while slicing preserves the container type and yields bytes. str is the unusual sequence whose elements are themselves str.

solid answer

~40 s

`bytes` and `bytearray` are sequences of small integers; there is no one-octet type, so `data[0]` returns the number 97. Slicing follows the general Python rule that a slice of a sequence is a sequence of the same type, so `data[0:1]` returns `b'a'`. `str` is the exception among Python sequences: since there is no character type, indexing a `str` returns a one-character `str`, which is why the two types behave differently under the same syntax. The practical consequences are that iterating a `bytes` object yields ints, so `for x in data: if x == b'\n'` never matches, and that `data[i:i+1]` is the idiom when you want a one-octet `bytes` to compare against a literal.

code

pycon · 11 lines
pycon
>>> data = b"abc"
>>> data[0]
97
>>> data[0:1]
b'a'
>>> list(data)
[97, 98, 99]
>>> "abc"[0]
'a'
>>> bytes([104, 105])
b'hi'

go deeper

for a junior

Remember the shape: indexing a bytes object gives a number, slicing gives a bytes object, and indexing a str gives a str. If a comparison in a loop over binary data never matches, this is usually why.

for a middle

Explain that bytes is a sequence of ints and that slices preserve the container type, so the real anomaly is str being self-similar. Know the one-element-slice idiom and the difference between bytes(3), bytes([3]) and b'3'.

for a senior

Show where this bites in practice: buffer scanning for delimiters and length prefixes, where an int-versus-bytes comparison silently returns False and the parser finds nothing. Have an opinion on working in integers throughout rather than scattering one-element slices.

for a principal

Frame it as a convention question for binary-protocol code: which layer speaks in integers, which in bytes objects, and how that is enforced by typing and review so parsers do not fail silently. The cost of getting it wrong is data accepted and misinterpreted rather than rejected.

## bytes is a sequence of numbers A `bytes` object is defined as an immutable sequence of integers in the range 0 to 255. `bytearray` is the mutable counterpart with the same element type. Python has no "single octet" scalar type, so the natural result of indexing one element is the integer itself: ```python data = b"abc" data[0] # 97 list(data) # [97, 98, 99] sum(data) # 294 ``` The `b'abc'` repr is a display convenience. It prints octets that happen to fall in the printable ASCII range as characters because that is far easier to read than three numbers, and it uses `\xNN` escapes for everything else. Nothing about that repr implies the object holds characters. Slicing behaves differently for a reason that has nothing to do with bytes specifically: in Python, a slice of a sequence is a sequence of the same type. Slicing a list gives a list, slicing a tuple gives a tuple, and slicing a `bytes` gives a `bytes` — even when the slice is one element long. ```python data[0:1] # b'a' data[0] # 97 ``` So the asymmetry is not a special case for `bytes` at all. The special case is `str`. ## str is the odd one out Python has no character type. A `str` is a sequence of code points, but the only object able to hold one code point is a one-character `str`. That makes `str` self-similar: `"abc"[0]` is `"a"`, an object of the same type as its container, and every element of a `str` is itself a `str` containing elements that are `str`. Almost no other sequence in the language works this way. Because developers meet `str` first, its behaviour feels like the rule and `bytes` feels like the anomaly. Reversing that intuition — `bytes` follows the general sequence rule, `str` is the exception — makes the whole family of related surprises predictable. ## The bugs this causes **Iteration yields integers.** The most common form is a scan over a binary buffer comparing against a literal: ```python for octet in data: if octet == b"\n": # never true: int == bytes is False ... if octet == 0x0A: # correct ... ``` The comparison does not raise; it quietly returns `False`, so the loop runs and finds nothing. When you want to compare against a `bytes` literal instead, slice a single element: `data[i:i+1] == b"\n"`. **Containment accepts both.** `b"\n" in data` and `10 in data` are both legal and both do what you would expect — Python 3 allows an integer on the left of `in` for `bytes` as a convenience. That inconsistency with iteration is a fine trap in a code review. **Construction from an int is a count, not a value.** `bytes(3)` is `b'\x00\x00\x00'` — three zero octets — while `bytes([3])` is `b'\x03'` and `b"3"` is the ASCII digit, octet 51. Three superficially similar expressions produce three different objects, and mixing them up in a protocol implementation produces a message the peer rejects for no visible reason. **Construction from text needs a codec.** `bytes("abc")` raises `TypeError: string argument without an encoding`, because turning text into octets always requires a codec choice; `"abc".encode("utf-8")` or `bytes("abc", "utf-8")` are the ways to say it. ## Where it actually bites Protocol and file-format work, where you walk a buffer looking for delimiters, length prefixes and magic numbers. Code written by someone thinking in `str` reads each element and compares it to a one-character literal, and the resulting parser silently accepts everything or finds nothing. Working in integers throughout — comparing to `0x0A`, using integer constants for framing octets — is usually clearer than sprinkling one-element slices, and it makes the intent visible: at that layer you are dealing with numbers, not text. ## What an interviewer is checking This is a curiosity question rather than a gate. Nobody's offer turns on it, but a candidate who can explain it has genuinely internalised that `bytes` is not "a string of ASCII", and that is the model that prevents the harder encoding bugs. Being able to add that `str` is the anomalous self-similar sequence, rather than `bytes` being weird, is the difference between having memorised the behaviour and having understood it.

  • What does bytes(3) produce in Python 3?
    Three zero octets, `b'\x00\x00\x00'` — the integer is a length, not a value. `bytes([3])` gives the single octet `b'\x03'`, and `b"3"` gives the ASCII digit, octet 51. The three expressions look alike and produce different objects, which is a reliable source of protocol bugs.
  • How do you compare each element of a bytes object against a literal newline?
    Either work in integers and compare to `0x0A`, or take a one-element slice and compare to a bytes literal: `data[i:i+1] == b"\n"`. What you must not do is compare an iterated element to `b"\n"` directly — iteration yields ints, so the comparison is always False and the loop silently finds nothing.
  • Why does indexing a str return another str?
    Because Python has no character type, so the only object that can represent one code point is a one-character str. That makes str a self-similar sequence, unlike every other builtin sequence. Recognising str as the exception rather than bytes as the anomaly is what makes the difference predictable instead of surprising.

A bytes object is a numbered ruler: pointing at one mark gives you a number, while asking for the span between two marks gives you a shorter ruler.

saying these in an interview costs you the question

  • Expects b'abc'[0] to be b'a'
  • Compares an iterated element to a bytes literal
  • Thinks bytes(3) produces the digit b'3'
  • Describes bytes as a string of ASCII characters
  • Calls bytes(text) without giving an encoding

context