skip to content

What is the difference between bytes and bytearray in Python?

level: juniorimportance: must knowfreq 62%

answer

  1. One is frozen, one you can edit
  2. Think dict keys and hashing
  3. Indexing gives a number, slicing gives a sequence
  4. Build in one, freeze at the boundary

basics

~20 s

bytes is an immutable sequence of octets; bytearray is the mutable version you can append to and edit in place. Only bytes is hashable, so only bytes works as a dict key or set member.

solid answer

~40 s

Both hold a sequence of 8-bit integers and share almost the same methods; the difference is mutability. `bytes` is immutable and hashable, has literal syntax (`b"abc"`), and is what you return, cache or use as a `dict` key. `bytearray` supports index assignment, slice assignment, `append` and `extend`, which makes it the buffer you build binary data in before freezing it with `bytes(buf)`. Both index to `int` and slice to their own type, so `b"hello"[0]` is `104` while `b"hello"[:2]` is `b'he'`. Both export the buffer protocol, but `bytes` exports a read-only buffer and `bytearray` a writable one — which is also why a `bytearray` refuses to resize while a `memoryview` of it is alive.

code

pycon · 15 lines
pycon
>>> data = b"hello"
>>> data[0]
104
>>> data[:2]
b'he'
>>> buf = bytearray(data)
>>> buf[0] = 72
>>> buf
bytearray(b'Hello')
>>> data == bytearray(data)
True
>>> hash(buf)
Traceback (most recent call last):
  ...
TypeError: unhashable type: 'bytearray'

go deeper

for a junior

Recall the one-line difference — bytes immutable, bytearray mutable — and the indexing rule that an index gives an int while a slice gives the same type. Knowing that only bytes is hashable is usually enough to pass this screen.

for a middle

Explain the mechanics: literal syntax only for bytes, in-place and slice assignment on bytearray, construction from an int meaning zero-fill, and cross-type equality. Say which one you would return from a function and why.

for a senior

Show you pick the type on purpose in real code: bytearray while assembling or patching a frame and as a readinto destination, bytes at every API boundary and anywhere the value is cached, compared or shared across threads.

for a principal

Own the boundary rule for a codebase: where binary values are frozen, who is allowed to hand out a mutable buffer, and what the copy at the freeze point costs. An accidentally-shared bytearray is an aliasing bug that outlives the function that made it.

### Two spellings of "a sequence of octets" `bytes` and `bytearray` are Python's two built-in sequences of 8-bit integers. They model the same thing — raw binary data, the octets that come off a socket, a file opened in binary mode, or a compression routine — and they share nearly the whole method surface. The one structural difference is that `bytes` is **immutable** and `bytearray` is **mutable**. Everything else follows from that. `bytes` is hashable, so it can be a `dict` key, a `set` member, or a value you cache safely; `bytearray` is unhashable and `hash()` on one raises `TypeError`. `bytes` literals have syntax (`b"abc"`); `bytearray` has none and must be constructed by calling `bytearray(...)`. A `bytes` object can be shared between threads or stored in a long-lived structure with no fear that someone edits it underneath you; a `bytearray` cannot. ### Sequence semantics that surprise people Both types index to **integers**, not to one-character slices. `b"hello"[0]` is `104`, the ordinal of `h`, while `b"hello"[:2]` is `b'he'`. This asymmetry — integer from an index, same-type object from a slice — is the single most common stumble, because `str` behaves differently: `"hello"[0]` is `"h"`, a one-character `str`. Iterating a `bytes` object therefore yields a stream of `int`, and `for byte in data: if byte == b"a"` never matches. Comparison crosses the type boundary: `b"ab" == bytearray(b"ab")` is `True`. Equality is by content, so a function can accept either and compare against a literal. Hashing does not cross the boundary, because only one side is hashable at all. ### What mutability buys you A `bytearray` supports in-place index assignment (`buf[0] = 72`), slice assignment of a different length (`buf[:3] = b"PUT"`, which resizes the object), `append`, `extend`, `insert`, `pop`, `del`, and `+=`. That makes it the natural **accumulation buffer**: append chunks as they arrive, patch a header field once the body length is known, then call `bytes(buf)` to freeze a copy for anything that wants an immutable value. It is also the natural **destination buffer** for the family of `readinto`-style I/O calls that fill a caller-supplied region instead of allocating a fresh object per read. Construction is worth memorising. `bytearray(5)` and `bytes(5)` both produce five zero bytes, not the digit five — an integer argument means "this many zero bytes". `bytearray(b"abc")` and `bytes(bytearray(b"abc"))` copy. `bytearray([104, 105])` builds from an iterable of ints, each of which must be in `range(256)` or you get a `ValueError`. ### Both export the buffer protocol `bytes` and `bytearray` both implement the C-level buffer protocol, which is how a consumer such as `memoryview` — or a C extension — gets at their memory without going through Python-level indexing. The difference shows up in the flags: `bytes` exports a **read-only** buffer, `bytearray` exports a **writable** one. So `memoryview(b"x")[0] = 1` raises `TypeError: cannot modify read-only memory`, while the same assignment through a view of a `bytearray` succeeds. Since Python 3.12 the protocol is also expressible in pure Python and the abstract base class `collections.abc.Buffer` names the concept for type checkers. The flip side of exporting a writable buffer is that a `bytearray` **cannot be resized while a view of it is outstanding**; `buf += b"d"` under a live `memoryview` raises `BufferError`. That is a feature, not a bug: it stops a reallocation from leaving the view pointing at freed memory. ### Choosing between them Reach for `bytes` by default: as a return value, as anything that will be hashed, cached, compared or shared, and as the thing you hand to a library boundary. Reach for `bytearray` when you are *building* or *editing* a block of binary data — assembling a frame from chunks, zero-filling a fixed-size record, patching a length prefix, or giving an I/O call somewhere to write. Convert once at the boundary with `bytes(buf)`, and be aware that the conversion copies. ### A worked shape The everyday pattern looks like this: start with `buf = bytearray()`, `+=` or `extend` each chunk as it arrives, patch fixed-offset fields by index or slice once you know them, and hand the result out as `bytes(buf)`. The mutable object never escapes the function that built it, so no caller can edit your value after the fact, and the single copy at the end buys you an object that is safe to cache, hash and share. The inverse mistake is instructive. Accumulating into a `bytes` object with `+=` means every append constructs a whole new object and copies everything seen so far, because there is nothing to append *to* — immutability guarantees it. `bytearray` grows amortized in place, which is why it, along with joining a list of chunks in one call, is the standard answer whenever binary data is assembled incrementally. Neither type is text. `bytes` holds octets; producing a `str` requires an explicit decode with a named encoding, and the reverse requires an explicit encode. Treating a `bytes` object as "a string that prints funny" is the origin of a large share of Python's encoding bugs, and the interviewer asking this question is usually checking that the boundary is clear in your head.

  • Why does iterating over a bytes object give integers rather than one-byte values?
    Because `bytes` is a sequence of 8-bit integers, and Python 3 kept indexing consistent with that model: `b"hi"[0]` is `104`. Slicing returns the sequence type, so `b"hi"[0:1]` is `b'h'`. In a loop, compare against integers, or slice a single element when you need a bytes value to compare with a literal.
  • What does bytearray(4) construct, and why does that trip people up?
    Four zero bytes — `bytearray(b'\x00\x00\x00\x00')`. An integer argument means "allocate this many zero-filled bytes", not "hold the number 4". The same is true of `bytes(4)`. To get the digits, encode a string or use `int.to_bytes`. It is the standard way to preallocate a fixed-size destination buffer.
  • Is b'ab' == bytearray(b'ab') true, and can you use either as a dict key?
    The comparison is `True` — equality between the two types is by content. Only `bytes` can be a key, though: `bytearray` is unhashable because its contents can change, and hashing it would break the invariant that a key's hash never moves. Convert with `bytes(buf)` first.

bytes is a printed page and bytearray is the same page on a whiteboard: identical content, but only one of them can be filed away as a permanent reference.

saying these in an interview costs you the question

  • Says bytearray is just another name for bytes
  • Thinks b'abc'[0] evaluates to b'a'
  • Uses bytes as an accumulation buffer, rebuilding it each append
  • Claims a bytearray can be used as a dict key
  • Assumes b'ab' == bytearray(b'ab') is False
  • Treats bytes as text rather than octets

context