skip to content

Why use struct.unpack_from and struct.pack_into over struct.unpack and pack?

level: middleimportance: should knowfreq 22%

answer

  1. Reading a record out of a bigger blob
  2. A bytes slice is a copy
  3. Offset in bytes, trailing data ignored
  4. pack_into needs a writable buffer
  5. bytearray, memoryview or mmap, not bytes

basics

~20 s

struct.unpack_from reads fields at an offset inside a larger buffer without slicing a copy out of it, and ignores trailing bytes. struct.pack_into writes fields in place into a pre-allocated writable buffer instead of allocating a new bytes object.

solid answer

~50 s

`struct.unpack(fmt, buf)` demands a buffer of exactly `calcsize(fmt)` bytes, so reading a header out of a big blob means slicing — and a slice of `bytes` is a **copy**. `struct.unpack_from(fmt, buf, offset)` reads the same fields directly at `offset` and simply ignores whatever follows, so you walk a record stream by stepping the offset instead of copying each record out. It accepts any buffer-protocol object: `bytes`, `bytearray`, a `memoryview`, an `mmap` region. `struct.pack_into(fmt, buf, offset, *values)` is the mirror image and needs a **writable** buffer — a `bytearray` or a writable `memoryview`; a `bytes` object fails with a `TypeError`. That is how you fill or patch one field of a fixed record in place, for example back-filling a length field once the body is written, without rebuilding the record. Both raise `struct.error` when the buffer is shorter than `offset + calcsize(fmt)`.

code

python · 9 lines
python
import struct

HEADER = "<4sHHI"
buf = bytearray(struct.calcsize(HEADER))
struct.pack_into(HEADER, buf, 0, b"IDX1", 1, 3, 0)
print(bytes(buf).hex())

struct.pack_into("<I", buf, 8, 4096)
print(struct.unpack_from(HEADER, buf, 0))

go deeper

for a junior

Know that both calls take an extra offset argument measured in bytes, and that pack_into writes into a bytearray you allocated rather than returning new bytes. Recognise them when reading existing parser code.

for a middle

Explain why a bytes slice is a copy and how the offset forms avoid it, which objects satisfy the buffer protocol, and why pack_into rejects an immutable bytes buffer. Derive offsets from calcsize, never from literals.

for a senior

Show judgement about when the exact-length check of plain unpack is the safety net you want, and use memoryview or mmap to keep large record streams copy-free while still validating declared lengths before trusting an offset.

for a principal

Frame it as a memory and allocation policy: which layers own buffers, where zero-copy views cross module boundaries, and how you keep an mmap-backed reader safe when the underlying file can be rewritten beneath it.

### The problem the `_from` / `_into` pair solves `struct.pack` allocates and returns a fresh `bytes` object. `struct.unpack` insists on a buffer of exactly the format's size. Both are perfectly fine for one record at a time, and both become awkward the moment the record lives inside something bigger — a 4 MB block read from a file, a memory-mapped index, a datagram with a header followed by a payload. Slicing works, and it costs. `blob[32:46]` builds a new 14-byte `bytes` object: an allocation, a copy, and later a deallocation, per record. Over millions of records those small copies are real, and the code that computes the two slice bounds is also the code that gets the arithmetic wrong. `struct.unpack_from(fmt, buffer, offset=0)` removes both problems. It reads the fields starting at `offset` inside a buffer of any length and ignores the rest — no exact-length requirement, no slice, no copy of the data being read. `struct.pack_into(fmt, buffer, offset, *values)` writes the fields at `offset` into a buffer you already own. ### What counts as a buffer Both functions speak the **buffer protocol**, so they accept far more than `bytes`: * `bytes` — readable only, fine for `unpack_from`. * `bytearray` — readable and writable, the usual target for `pack_into`. * `memoryview` — a zero-copy window over another buffer; slicing a `memoryview` does not copy, which is exactly why it pairs well with these calls. * `mmap.mmap` — a memory-mapped file, so `unpack_from` reads a record straight out of a file's pages and `pack_into` writes one back. * `array.array` — also a buffer. `pack_into` needs the buffer to be **writable**. Handing it an immutable `bytes` raises a `TypeError` complaining that a read-write bytes-like object was required — which is a good error, because the alternative would be silent data loss. ### Offsets, and their edges The offset is in **bytes**, always, whatever the field types are. The rules worth knowing: * The buffer must hold at least `offset + calcsize(fmt)` bytes, or you get a `struct.error` naming the size required and the offset given. * A **negative** offset is allowed for `unpack_from` and counts back from the end of the buffer — handy for a trailer record at the end of an index file. * Nothing checks that your offset lands on a record boundary. Stepping by a hard-coded number instead of `struct.calcsize(fmt)` (or `Struct.size`) is the classic way to read a stream one byte out of phase and get plausible garbage. ```python import struct rec = struct.Struct("<IH") blob = rec.pack(1, 10) + rec.pack(2, 20) + b"trailer" view = memoryview(blob) for off in range(0, 2 * rec.size, rec.size): print(rec.unpack_from(view, off)) ``` Note what happens with plain `unpack` here: the trailing bytes make the buffer the wrong length and the call fails outright. Tolerating trailing data is the second reason to reach for `unpack_from` — a header parser should not care what follows the header. ### Patching a field in place `pack_into` earns its keep when a record must be written before one of its fields is known. Write the record with a placeholder, keep the offset, finish the body, then overwrite just that field: ```python import struct HEADER = "<4sHHI" # magic, version, shards, doc count buf = bytearray(struct.calcsize(HEADER)) struct.pack_into(HEADER, buf, 0, b"IDX1", 1, 3, 0) # ... build the body, counting documents as you go ... struct.pack_into("<I", buf, 8, 4096) # back-fill the count at byte 8 ``` The alternative — rebuilding the whole header as a new `bytes` and reassembling the record — is more code and more allocation for the same result. It also generalises: the same technique fills a fixed-size output buffer record by record in a loop, reusing one `bytearray` rather than allocating per record and joining at the end. ### When plain `pack` / `unpack` is still right Do not reach for the offset forms reflexively. When you have exactly one record's worth of bytes — a datagram you just received, a row you just read — `unpack`'s exact-length requirement is a **free length check**, and losing it means a truncated read now returns a plausible tuple instead of raising. Use `unpack` when the strictness is a feature, and `unpack_from` when you are genuinely reading inside a larger buffer. And in either case derive every offset from `struct.calcsize` or a `Struct`'s `size`, never from a literal.

  • What happens if you call struct.pack_into with a bytes object as the buffer?
    It fails with a `TypeError` saying a read-write bytes-like object is required. `bytes` is immutable, so there is nowhere to write; the target must be a `bytearray`, a writable `memoryview` over one, an `mmap` region, or another writable buffer. The failure is immediate and loud, which is the right behaviour.
  • How does struct.unpack_from behave when the buffer is shorter than offset plus the format size?
    It raises `struct.error`, reporting the number of bytes the format needed and the offset it was given — it never returns a partial tuple. A negative offset is legal and counts back from the end of the buffer, which is useful for a trailer record, but it is checked the same way.
  • Why pass a memoryview rather than the bytes object itself?
    Because slicing a `memoryview` is free while slicing `bytes` copies. If your loop only ever passes an offset to `unpack_from`, either works; the `memoryview` matters when you also need to hand a sub-range of the buffer to something else — a writer, a hash, a decompressor — without duplicating megabytes to do it.

saying these in an interview costs you the question

  • Slices a large buffer just to unpack a header
  • Calls struct.pack_into on an immutable bytes object
  • Expects struct.unpack to ignore trailing bytes
  • Steps offsets by a literal instead of calcsize
  • Thinks unpack_from copies the whole buffer first
  • Uses unpack_from where unpack's length check was the safety net

context