Why slice a memoryview of a large bytes object instead of the bytes itself?
answer
- Slicing a big blob keeps allocating
- A window onto memory, not a duplicate
- PEP 3118 exporters and consumers
- Slice cost is O(1), not O(k)
- tobytes() is where the copy returns
basics
~20 sSlicing a bytes object allocates a new bytes and copies the requested range. A memoryview exposes the same memory through the buffer protocol, so slicing it returns another view over the original block in constant time, copying nothing.
solid answer
~40 s`data[2:5]` builds a brand-new `bytes` object and memcpy's the range into it, so repeatedly slicing a multi-megabyte blob costs time and allocations proportional to everything you slice. `memoryview(data)` asks the object for its underlying memory through the PEP 3118 buffer protocol; slicing that view produces another `memoryview` that records an offset and a length over the *same* bytes, which is O(1) and a fixed-size object no matter how big the slice is. The view keeps its exporter alive through `memoryview.obj`, inherits its read-only flag (a view of `bytes` is read-only, a view of `bytearray` is writable), and reports `memoryview.nbytes` for the byte count. The copy only happens when you force it with `memoryview.tobytes()` — many stdlib calls, such as `struct.unpack_from` or a hash object's `update`, take the view directly.
code
python · 4 linesdata = bytes(range(10))
view = memoryview(data)
chunk = view[2:5]
print(chunk.tobytes(), chunk.nbytes, chunk.obj is data, chunk.readonly)go deeper
Recall the one-line contrast: slicing bytes allocates and copies, slicing a memoryview does not. Know that memoryview(x) works on bytes, bytearray and array.array, and that tobytes() is what produces a real copy.
Explain the mechanics: the buffer protocol hands over a pointer, a length and a read-only flag; a slice of a view is offset arithmetic. Be able to say why a view of bytes cannot be written to and why memoryview.obj matters.
Show the production judgement: where views actually cut allocations in a hot parse or I/O path, how a small view can pin a huge buffer, and where you deliberately convert back to bytes at an API boundary rather than leaking views into long-lived state.
Own the tradeoff at the codebase level: zero-copy buys throughput and costs readability, lifetime discipline and a class of aliasing bugs. Decide where the zero-copy region ends and copies are mandated, and make that boundary explicit in the design rather than per-call.
## The copy you did not ask for `bytes` and `bytearray` are value types, and slicing them behaves like slicing any Python sequence: `data[2:5]` calls `bytes.__getitem__` with a slice object, allocates a fresh `bytes` of the requested length, copies the bytes into it, and returns it. That is O(k) time and O(k) memory in the size of the slice. Parsing a 40 MB blob by repeatedly slicing off headers and payloads can therefore copy far more data than the blob itself contains, and every one of those copies is a separate heap object for the allocator and the garbage collector to deal with. ## What the buffer protocol is PEP 3118 defines a C-level contract between an *exporter* — an object that owns a contiguous block of memory — and a *consumer* that wants to look at that block. The exporter fills in a description: a base pointer, the total byte length, an item format code, the number of dimensions, a shape, strides, and a read-only flag. Nothing is copied; the consumer just learns where the memory is and how it is laid out. `bytes`, `bytearray`, `array.array` and most extension types that hold blocks of numbers are exporters. Since Python 3.12 (PEP 688) the same contract is reachable from Python code as the `__buffer__` and `__release_buffer__` methods plus the `collections.abc.Buffer` ABC, so a pure-Python class can now be an exporter and can be type-checked as one. ## memoryview is the built-in consumer `memoryview(obj)` requests a buffer from `obj` and wraps the result in a Python object. Indexing it with a slice does not touch memory: it constructs another `memoryview` whose base pointer is shifted and whose length is shortened. Two consequences follow: * Slicing is constant time, and the resulting object's size does not depend on the size of the slice. * Mutating a writable view mutates the exporter, and vice versa — there is exactly one copy of the data, seen through two names. The view exposes what the protocol told it: `memoryview.nbytes` is the byte count, `len()` is the *item* count (they differ once you cast to a wider format), `memoryview.itemsize` and `memoryview.format` describe one element, `memoryview.readonly` says whether writes are allowed, and `memoryview.shape` and `memoryview.strides` describe the layout. `memoryview.obj` is the exporter itself, and holding the view holds a reference to it, so the underlying memory cannot be freed while any view is alive. ## Read-only is inherited, not chosen A view of `bytes` is read-only because `bytes` is immutable; assigning through it raises `TypeError: cannot modify read-only memory`. A view of `bytearray` is writable. You can narrow a writable view with `memoryview.toreadonly()` before handing it to code you do not want mutating your buffer — that returns a read-only view over the same memory, not a copy. Relatedly, a writable view is unhashable (`hash()` raises `ValueError`) precisely because its contents can change under a dict; a read-only view of `bytes` hashes equal to that `bytes`. ## Where the copy comes back `memoryview.tobytes()`, `memoryview.tolist()` and `bytes(view)` all materialize a copy — that is their job, and it is the moment to be deliberate about. Anything that needs a real `bytes` object (a `str.encode` result compared by identity, an API annotated for `bytes`, storing the value in a long-lived dict) forces one. But a surprising amount of the standard library accepts a buffer directly and never copies: `struct.unpack_from(fmt, view, offset)`, a hash object's `update`, a binary file's `write`, `socket.socket.send`, and `int.from_bytes`. ## The price of a view Zero-copy is not free of consequences. A live view *pins* its exporter: while any view of a `bytearray` exists, that `bytearray` cannot be resized, and attempting it raises `BufferError`. A view also keeps the whole exporter alive, so keeping a 12-byte view of a 40 MB buffer keeps all 40 MB resident — the classic memoryview leak. And the ergonomics are thinner than `bytes`: no `split`, no regex, arithmetic on offsets instead of friendly slicing helpers. The rule of thumb is to use views inside the hot parsing or I/O path where the copies actually show up in a profile, and to convert to `bytes` at the boundary where the value is stored or returned.
- Is a memoryview of a bytes object writable?No. The read-only flag comes from the exporter, not from you: `bytes` is immutable, so its view raises `TypeError: cannot modify read-only memory` on assignment. A view of a `bytearray` or an `array.array` is writable. You can go one way only — `memoryview.toreadonly()` returns a read-only view over the same memory — but there is no call that makes a read-only view writable.
- How can holding a small memoryview leak a large amount of memory?A view holds a reference to its exporter through `memoryview.obj`, and the exporter owns the whole block. A 20-byte view carved out of a 40 MB `bytearray` keeps all 40 MB alive for as long as the view is reachable. If you intend to keep only the small piece — caching it, returning it, storing it in a dict — call `memoryview.tobytes()` to detach a real copy and drop the view.
- When does using memoryview actually not pay off?When the slices are small and short-lived. Copying 30 bytes is cheap, and the view object itself is an allocation too, so a per-slice memoryview can be slower than plain slicing. Views pay off on large ranges, on hot parse loops over big blobs, and wherever you need to write into a buffer someone else owns. Measure before converting readable code into offset arithmetic.
Slicing bytes is photocopying the pages you want; slicing a memoryview is laying a cardboard frame over the open book. The frame is cheap, but the book cannot be reshelved while a frame sits on it.
saying these in an interview costs you the question
- Claims memoryview copies the data lazily on creation
- Thinks a memoryview of bytes can be assigned to
- Says tobytes() is free because it is zero-copy
- Confuses len(view) with the byte count for a cast view
- Believes the exporter can be garbage collected while a view lives
- Assumes memoryview always beats slicing, even for tiny slices