skip to content

Why does struct.calcsize('@ci') exceed struct.calcsize('<ci')?

level: middleimportance: must knowfreq 34%

answer

  1. The prefix sets three things, not one
  2. Native mode obeys compiler alignment rules
  3. Standard mode never inserts padding
  4. '@' is the default when no prefix
  5. '!' equals '>' — big-endian, standard sizes

basics

~20 s

The '@' prefix means native byte order with native alignment, so struct inserts pad bytes before the int — typically eight bytes in total. The '<' prefix selects little-endian standard sizes with no alignment padding, giving exactly five.

solid answer

~40 s

The prefix chooses three things at once: byte order, field sizes and alignment. `'@'` — the default when you write no prefix — means **native** order, **native** sizes and **native alignment**, so the C compiler's padding rules apply and `'@ci'` typically measures 8 bytes on a 64-bit machine: one for the char, three pad, four for the int. Every other prefix uses **standard sizes with no alignment padding**: `'<'` little-endian, `'>'` big-endian, `'!'` network order (identical to `'>'`), `'='` native order but standard sizes. So `'<ci'` is exactly 5. Anything crossing a machine boundary — a file, a socket, a shared index — must name an explicit prefix; the native default is only safe inside one process on one machine. Insert holes deliberately with the `'x'` pad character rather than relying on `'@'`.

code

python · 5 lines
python
import struct

print(struct.calcsize("@ci"), struct.calcsize("=ci"), struct.calcsize("<ci"))
print(struct.calcsize("@ic"), struct.calcsize("<l"), struct.calcsize("<q"))
print(struct.pack("<i", 1).hex(), struct.pack(">i", 1).hex())

go deeper

for a junior

Remember that a format string can start with a prefix and that leaving it off means native. Know that '<' is little-endian and '>' is big-endian, and that struct.calcsize tells you the real size.

for a middle

Explain all three effects of the prefix — order, sizes, alignment — and work an example such as '@ci' against '<ci'. Know the standard widths, that 'l' is four bytes there, and that 'n', 'N' and 'P' are native-only.

for a senior

Show the operational rule: every layout leaving the process names its prefix, holes are explicit 'x' bytes, and calcsize is asserted against the documented record size so a format drift fails at import rather than corrupting a parse.

for a principal

Own the interoperability contract across a fleet: pin one byte order for all stored and transmitted records, document the layout with its padding, and plan how a layout version is signalled so records written before a change stay readable.

### The prefix does three jobs, not one Newcomers read the leading character of a `struct` format as "endianness". It is actually a triple: **byte order**, **field size convention**, and **alignment**. Getting that wrong is how a binary file written on one machine is read as garbage on another. | prefix | byte order | sizes | alignment | |---|---|---|---| | `@` (default) | native | native | native — pad bytes inserted | | `=` | native | standard | none | | `<` | little-endian | standard | none | | `>` | big-endian | standard | none | | `!` | big-endian (network) | standard | none | **Native mode** (`'@'`, and therefore any format with no prefix at all) asks the C compiler what an `int` looks like on this build and where it must sit. On a typical 64-bit platform an `int` wants 4-byte alignment, so in `'@ci'` the single `c` byte is followed by **three pad bytes** before the `i` — `struct.calcsize('@ci')` reports 8. Reverse the fields and the padding disappears: `'@ic'` is 5, because nothing needs aligning after the char. `struct` does *not* append trailing padding automatically; if you need the record rounded up the way a C compiler would round a real `struct`, end the format with a zero-repeat field of the widest type, as in `'@ci0i'`. **Standard mode** is everything else. Sizes are fixed by the module regardless of platform: `b`/`B` 1, `h`/`H` 2, `i`/`I`/`l`/`L` 4, `q`/`Q` 8, `e` 2, `f` 4, `d` 8. Note `l` — in standard mode it is **4 bytes**, not the 8 a 64-bit C `long` gives you; if you mean 64 bits, write `q`. No pad byte is ever inserted, so the record length is simply the sum of the field widths and `'<ci'` is 5. Two format characters exist only in native mode: `n` and `N` (signed and unsigned `size_t`) and `P` (a pointer) are rejected with a `struct.error` under `'<'`, `'>'`, `'!'` or `'='`, precisely because their width is a property of the machine and standard mode refuses to guess. ### Why this shows up as a bug, not a question The failure is asymmetric and quiet. A writer that omits the prefix produces native-order, natively-padded records. On the same architecture, the reader agrees and everything works — including all your tests. The day a record is written by one host and read by another with a different word size, a different alignment rule, or the opposite byte order, the reader gets plausible-looking nonsense: a field shifted by three bytes, or an integer with its bytes reversed. Byte-swapped values are recognisable once you have seen them (`1` reads as `16777216`); misalignment is subtler, because every field after the hole is wrong. The discipline is one line: **any layout that leaves the process gets an explicit prefix.** Little-endian `'<'` is the usual choice for files, because every mainstream CPU is little-endian and the bytes then match a memory dump. `'!'` is the usual choice for anything on a socket, because network byte order is big-endian by convention and the `!` documents that intent to the next reader; it is otherwise identical to `'>'`. ### Reproducing a C layout on purpose If you are matching an existing C `struct`, do not lean on `'@'` and hope the alignment matches — the C side may be packed with a pragma, and the padding depends on the compiler and the target. Write the padding out explicitly instead with `'x'`: ```python import struct native = struct.calcsize("@ci") # 8 on a typical 64-bit build explicit = struct.calcsize("<cxxxi") # 8 everywhere, and it says why ``` Now the layout is stated in the format string, is identical on every platform, and reviews cleanly. Pair it with an assertion that `struct.calcsize(FMT)` equals the record size the specification promises, and a wrong format character becomes an import-time failure instead of a corrupt parse. ### The two-value check in an interview Being able to say *why* `'@ci'` and `'<ci'` differ, and then say *which one you would ship*, is the whole question. The answer is that native mode is for talking to this process's own C code — a `ctypes` structure, a memory-mapped region defined by a library on this machine — and standard mode with an explicit byte order is for everything that has to be read by someone else, later, somewhere else.

  • What does the '!' prefix mean, and when do you reach for it?
    `'!'` is network byte order — big-endian with standard sizes and no padding, byte-for-byte identical to `'>'`. Reach for it when the layout goes on a socket and the specification talks about network order, because the character documents that intent; use `'<'` for on-disk formats, where little-endian matches every mainstream CPU and a hex dump reads naturally.
  • A file written with no prefix reads as garbage on another host. What happened?
    No prefix means `'@'`: native order, native sizes, native alignment. The writer's compiler inserted pad bytes and used its own word size and endianness, and the reading host disagrees on at least one of the three. Fix it by pinning the format — `'<'` or `'!'` — and writing any holes as explicit `'x'` pad bytes, then re-generating the file; there is no way to read the old one without knowing the writer's platform.
  • How wide is the 'l' format character?
    Four bytes in standard mode — `'<l'` is 4, not 8 — because the module fixes standard sizes independently of the platform. Only in native mode does `'@l'` follow the C `long`, which is 8 bytes on 64-bit Unix and 4 on Windows. If you mean a 64-bit integer, always write `'q'`.

Native mode lets the local compiler arrange the furniture to suit this room; standard mode ships flat-packed to fixed dimensions so it assembles identically in someone else's house.

saying these in an interview costs you the question

  • Thinks '@' and '=' mean the same thing
  • Assumes 'l' is eight bytes in standard mode
  • Writes a file format with no byte-order prefix
  • Believes struct always packs fields back to back
  • Calls '!' a special network encoding rather than big-endian
  • Expects '@' to append trailing padding like a C compiler

context