Why read a binary file with readinto and a preallocated bytearray instead of read?
answer
- Who allocates the chunk each iteration?
- The call returns a count, not an object
- Zero means end of file here
- Slice the view, never the bytearray
- Short reads leave stale trailing bytes
basics
~20 sread(n) allocates a fresh bytes object on every call. readinto(buf) writes into a writable buffer you already own and returns how many bytes it actually wrote, so a streaming loop reuses one allocation instead of producing garbage per chunk.
solid answer
~40 sA binary file's `read(n)` returns a new `bytes` each call, so streaming a large file in chunks allocates and frees one object per chunk. `readinto(buf)` instead takes any writable object supporting the buffer protocol — typically a `bytearray` or a `memoryview` of one — copies bytes directly into it, and returns the count actually written, which may be less than the buffer's length and is `0` at end of file. Because it returns a count rather than a sized object, you track the fill position yourself; slice a `memoryview` of the buffer to hand `readinto` the remaining tail, since slicing the `bytearray` directly would copy. The target must be writable: passing `bytes` raises `TypeError`. `socket.socket.recv_into` is the same pattern for sockets.
code
python · 13 linesimport io
src = io.BytesIO(bytes(range(32)))
buf = bytearray(32)
view = memoryview(buf)
filled = 0
while filled < len(buf):
n = src.readinto(view[filled:])
if not n:
break
filled += n
view.release()
print(filled, buf[:6])go deeper
Recall that read(n) hands you a new bytes while readinto(buf) fills a buffer you already made and returns a count. Know that the target must be mutable, so a bytearray and not a bytes.
Explain the mechanics: one allocation reused across iterations, a returned count that can be short, 0 as end of file, and why the unfilled tail must be a memoryview slice rather than a bytearray slice.
Demonstrate the safety discipline around a reused buffer: copy or parse before the next read, release the view so the buffer is not pinned, handle short reads explicitly, and be able to say from a profile why the allocation churn mattered here.
Own where the zero-copy region starts and stops across a service. Buffer reuse trades allocation churn for aliasing hazards and harder-to-review code, so decide which layers may hand out views and which must hand out copies, and write that boundary down.
## Two shapes of the same read Binary file objects expose two ways to pull bytes out: * `read(n)` — returns a new `bytes` object of at most `n` bytes. The allocation size is decided by the call. * `readinto(b)` — copies at most `len(b)` bytes into the writable buffer `b` you supply, and returns the number of bytes actually copied. Functionally they are equivalent; operationally they differ in who allocates. Streaming a large file with `read(65536)` in a loop produces one fresh `bytes` object per iteration; the allocator and the reference-counting machinery handle every one of them, and if you then slice or concatenate the chunk you copy again. `readinto` moves the allocation out of the loop: you make one `bytearray` up front and refill it. ## What readinto accepts, and what it returns The argument must be a *writable* object supporting the buffer protocol — `bytearray`, a writable `memoryview`, an `array.array`. Passing an immutable `bytes` raises `TypeError: readinto() argument must be read-write bytes-like object, not bytes`, which is the buffer protocol's read-only flag being enforced at the boundary. The return value is the byte count, and the two edge cases matter: * A short read is legal. On a buffered reader a full buffer is typical, but on a pipe, a socket-backed file or a raw unbuffered file you can get fewer bytes than you asked for with more data still to come. Code that assumes `readinto` fills the buffer will silently process stale trailing bytes from the previous iteration. * `0` means end of file. That is the loop's terminating condition; there is no falsy empty-`bytes` sentinel as with `read`. `readinto1` exists on buffered readers and issues at most one call to the underlying raw stream, which matters when you want latency rather than a full buffer. ## Tracking the fill position with a memoryview The natural pattern is to read a record of known size into a fixed buffer, looping until it is full. The trap is expressing "the unfilled tail" as `buf[filled:]` — that slices the `bytearray`, which *copies* the tail into a new object, and `readinto` then fills the copy that you promptly throw away. The whole point of the exercise disappears, silently and without an error. The fix is to take one `memoryview` of the buffer and slice the view instead: `view[filled:]` is a window onto the tail of the same memory, so `readinto` writes into the real buffer. ```python view = memoryview(buf) filled = 0 while filled < len(buf): n = src.readinto(view[filled:]) if not n: raise EOFError("short record") filled += n ``` Release the view (`memoryview.release()` or a `with` block) when the loop ends, or the `bytearray` stays pinned and cannot be resized. ## Reusing the buffer safely Because there is only one buffer, everything derived from an iteration must be consumed or copied before the next `readinto` overwrites it. Parse the record and keep the parsed values; if you keep a view of the raw bytes instead, the next read mutates what you kept. Anything that outlives the iteration needs `memoryview.tobytes()` or an equivalent copy. This is the discipline that makes buffer reuse safe, and it is where reused-buffer code usually goes wrong. ## Where else the pattern appears `socket.socket.recv_into` is the same contract for sockets, and it is where reuse matters most, because a receive loop runs continuously. Hash objects' `update`, `struct.unpack_from` and a binary file's `write` all accept a buffer, so a whole read-parse-hash-forward path can run without materializing a single intermediate `bytes` object. ## When not to bother For a small configuration file, `read()` in one shot is clearer and the allocation is irrelevant. `readinto` earns its extra state — a buffer, a fill counter, a short-read branch — in loops that run millions of times, in receive paths where allocation churn shows up as latency variance, and where the destination memory is fixed by something outside your control. Reach for it after a profile, not before.
- Why slice the memoryview rather than the bytearray when tracking the fill offset?`buf[filled:]` slices the `bytearray`, which allocates a copy of the tail; `readinto` then fills that throwaway copy and the real buffer never changes. `memoryview(buf)[filled:]` is a window onto the same memory, so the write lands in the buffer. The bug is silent — no exception, just a buffer that stays empty past the first chunk.
- What must be true before the next readinto overwrites the buffer?Everything derived from this iteration must be consumed or copied. Parsed integers and strings are fine because they are independent objects, but a `memoryview` slice of the buffer is not — the next read mutates what it points at. Anything that outlives the iteration needs `memoryview.tobytes()` or an equivalent copy first.
- How do you know the read finished versus hit end of file?`readinto` returns the number of bytes written; `0` means end of file. A non-zero count smaller than the buffer is a legal short read, common on pipes and raw unbuffered streams, so a correct loop accumulates the count and keeps going rather than assuming one call fills the buffer.
saying these in an interview costs you the question
- Assumes readinto always fills the whole buffer
- Passes a bytes object as the readinto target
- Slices the bytearray instead of a memoryview of it
- Treats the return value as the data rather than a count
- Keeps views of the reused buffer across iterations
- Uses readinto everywhere without measuring the gain