skip to content

struct Packing and Endianness

struct.pack and unpack turn a fixed-width record into a tuple and back, once you get the format string, byte order and padding right. Interviewers use it to see whether you can read a binary protocol.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

How do struct.pack and struct.unpack convert Python values to and from bytes?

level: juniorimportance: must knowfreq 40%

answer

  1. A format string describes a fixed layout
  2. One character per field, sizes fixed
  3. unpack always hands back a tuple
  4. Buffer length must equal calcsize exactly
  5. struct.error on wrong size or range

basics

~20 s

struct.pack takes a format string plus values and returns a bytes object laid out exactly as that format describes. struct.unpack reverses it and always returns a tuple, even when the format has a single field.

solid answer

~40 s

`struct.pack(fmt, *values)` walks a format string such as `'<IHf'` — an optional byte-order prefix followed by one character per field — and writes each value into a fixed-width binary layout, returning `bytes`. `struct.unpack(fmt, buffer)` is the inverse and **always** returns a tuple, so a single-field read needs `(n,) = struct.unpack('<I', data)` or an index. The buffer given to `unpack` must be exactly `struct.calcsize(fmt)` bytes; longer or shorter raises `struct.error`, which is the usual first bug. Values must also fit their format character — packing 300 into `'B'` raises `struct.error` too. Text is never implicit: the `'s'` character consumes and produces `bytes`, so you encode and decode yourself. `struct` is the stdlib tool for fixed-width C-style records and protocol headers, not for self-describing formats.

code

python · 8 lines
python
import struct

packed = struct.pack("<IHf", 4096, 7, 0.5)
print(len(packed), packed.hex())
print(struct.unpack("<IHf", packed))

(count,) = struct.unpack("<I", packed[:4])
print(count)

go deeper

for a junior

Be ready to write a pack and an unpack from memory and to say out loud that unpack returns a tuple. Know that the buffer length must match the format exactly, and that text fields need encoding before packing.

for a middle

Explain the mechanics: format characters and their widths, the repeat count versus the length prefix on 's', pad bytes, and the range and type checks that raise struct.error. Show how you layer a fixed header over a variable-length body.

for a senior

Demonstrate defensive parsing of untrusted binary: validate declared lengths before allocating or slicing, assert calcsize against the documented record size, and log the format string with the buffer length when a struct.error escapes in production.

for a principal

Own the choice of wire format. Argue when a fixed-width struct layout is the right contract — stable, dense, cheap — and when its lack of self-description and version tolerance makes a schema-driven or self-describing format the cheaper long-run decision.

### What the module actually does `struct` is the standard library's bridge between Python objects and **fixed-width binary layouts** — the kind of record a C `struct`, a file header or a wire protocol defines. It has no schema of its own and no self-description: you tell it the layout with a *format string*, and it trusts you. That is the whole design, and it is why the format string is the only thing that matters. A format string is an optional **byte-order/size/alignment prefix** followed by one or more **format characters**, each standing for one field of a known width. The common ones: | char | C type | standard size | |---|---|---| | `b` / `B` | signed / unsigned char | 1 | | `?` | bool | 1 | | `h` / `H` | short | 2 | | `i` / `I` | int | 4 | | `q` / `Q` | long long | 8 | | `e` / `f` / `d` | half / float / double | 2 / 4 / 8 | | `s` | bytes | count-prefixed, e.g. `16s` | | `x` | pad byte | 1, consumes no value | A digit before a character is a **repeat count**, with one exception that trips everyone: for `s` the digit is a *length*, not a repetition. `'4i'` means four integers and needs four values; `'4s'` means one four-byte bytes field and needs one value. ### pack ```python import struct packed = struct.pack("<IHf", 4096, 7, 0.5) print(len(packed)) # 10 = 4 + 2 + 4 ``` `pack` returns a new `bytes` object of exactly `struct.calcsize(fmt)` length. Each value is range-checked against its character: `struct.pack('B', 300)` raises `struct.error: ubyte format requires 0 <= number <= 255`. Types are checked too — an `'i'` field will not silently accept a float, and an `'s'` field will not accept a `str`. That last one is the single most common beginner error, and it is deliberate: `struct` moves bytes, and choosing an encoding is your decision, not the module's. Encode first (`name.encode('utf-8')`), and remember that a `'16s'` field pads short input with trailing NUL bytes and truncates long input without complaint. ### unpack ```python values = struct.unpack("<IHf", packed) print(values) # (4096, 7, 0.5) ``` Two properties matter in an interview. First, the return is **always a tuple**, because a format can describe any number of fields and the API refuses to special-case one. Junior code that writes `n = struct.unpack('<I', data)` and then does arithmetic on `n` fails with a `TypeError` several lines later; the idiom is tuple-unpacking with a trailing comma, `(n,) = struct.unpack('<I', data)`, or `struct.unpack('<I', data)[0]`. Second, `unpack` demands an **exact-length** buffer. Handing it a 3-byte buffer for `'<I'` raises `struct.error: unpack requires a buffer of 4 bytes`; handing it 8 bytes raises the same class of error. That strictness is a feature — it catches a mis-sized read at the point of the read rather than as garbage three records later. When you genuinely have a bigger buffer and want a record out of the middle of it, that is what `struct.unpack_from` and an offset are for. `x` pad characters are skipped in both directions: they consume a byte of buffer and no value, which is how you reproduce a C layout that has holes in it. ### Where struct stops `struct` describes layouts that are **fixed at compile time**. There is no repetition count driven by a field's value, no length-prefixed variable field, no optional field, no version negotiation. Real formats have all of those, so real parsers use `struct` in layers: unpack a fixed header, read the count or length it declares, then unpack the body with a format built for that count — often with an f-string such as `f"<{count}I"`. That layering is the practical technique, and being able to describe it is what separates "I have used struct" from "I have parsed a format with struct". It is also not a serialisation format in its own right. A `struct` payload carries no type tags, no field names and no length, so both ends must already agree on the layout; the moment you want a self-describing payload, the module you want is `json`, `pickle` or a schema-driven format, not `struct`. ### Sanity checks that pay off `struct.calcsize(fmt)` reports the exact byte length a format describes, and asserting your reader's `calcsize` against the record size the format document promises is the cheapest possible regression test for a binary parser. Wrap a first parse in `try: ... except struct.error as exc:` and log the format and the buffer length together — that pair is nearly always enough to identify a wrong format character or a short read.

  • What does struct.calcsize give you, and why call it before a read?
    `struct.calcsize(fmt)` returns the exact number of bytes that format describes. It is how you size a read, step an offset through a record stream, and assert at import time that your format still matches the record length the format document promises — a one-line regression test for a binary parser.
  • How do you handle a text field with struct's 's' format character?
    Encode and decode yourself. `'16s'` consumes and produces `bytes`, not `str`; packing a short value pads it with trailing NUL bytes and packing a long one truncates silently. So pack `value.encode('utf-8')` and, on the way back, `raw.rstrip(b'\x00').decode('utf-8')` — and validate the decode, because a truncated multi-byte character will raise.
  • How do you parse a record whose field count is only known at runtime?
    In two steps. Unpack the fixed header with a static format, read the count it declares, then build the body format from that count — `struct.unpack(f'<{count}I', body)` — after checking the count against the bytes you actually have. `struct` formats are fixed at parse time, so anything variable-length has to be layered like this.

A format string is a stencil cut for one shape of record: pack presses values through it into bytes, and unpack lays the same stencil back over bytes to read the fields out.

saying these in an interview costs you the question

  • Thinks struct.unpack returns a bare value, not a tuple
  • Passes a str to a struct 's' field instead of bytes
  • Assumes struct.unpack tolerates extra trailing bytes
  • Reads '4s' as four separate bytes fields
  • Believes struct can describe variable-length or optional fields
  • Treats struct output as a self-describing serialisation format

context

open as a page

Why does struct.calcsize('@ci') exceed struct.calcsize('<ci')?

level: middleimportance: must knowfreq 34%

basics

~20 s

The '@' prefix means native byte order with native alignment, so struct inserts pad bytes before the int — typically eight bytes in total. The '<' prefix selects little-endian standard sizes with no alignment padding, giving exactly five.

open as a page

Why use struct.unpack_from and struct.pack_into over struct.unpack and pack?

level: middleimportance: should knowfreq 22%

basics

~20 s

struct.unpack_from reads fields at an offset inside a larger buffer without slicing a copy out of it, and ignores trailing bytes. struct.pack_into writes fields in place into a pre-allocated writable buffer instead of allocating a new bytes object.

open as a page

struct.unpack per record dominates a 45-second index cold start — how do you cut it?

level: seniorimportance: should knowfreq 20%

basics

~20 s

Compile the layout once as a module-level struct.Struct and stream the buffer through Struct.iter_unpack instead of slicing and calling struct.unpack per record. The C-level iterator removes the per-record slice, the lookup and most of the Python loop overhead.

open as a page