How do you define a ctypes.Structure and pass it to C by reference?
answer
- Fields are declared in memory order
- Each field reads bytes in the instance
- argtypes decides value versus address
- byref is call-only, pointer is storable
- Immutable bytes is never an output buffer
basics
~10 sSubclass ctypes.Structure and set fields to an ordered list of (name, type) pairs, which fixes the memory layout. Declare the parameter as ctypes.POINTER(YourStruct) in argtypes and pass ctypes.byref(instance) at the call.
solid answer
~40 sA struct is a `ctypes.Structure` subclass whose `_fields_` list gives the fields in memory order; each field is a descriptor reading and writing bytes inside the instance's own block, so mutations are visible to C and back. Layout follows the platform compiler's alignment rules by default — check `ctypes.sizeof` against the C side's `sizeof` before trusting a transcription. Whether the struct goes by value or by reference is decided by `argtypes`: `[Digest]` copies it into the call, `[ctypes.POINTER(Digest)]` takes an address, and you pass `ctypes.byref(d)` for a cheap call-only reference or `ctypes.pointer(d)` when you need a storable pointer object. For a buffer C writes into, allocate a mutable one with `ctypes.create_string_buffer(n)` or `(ctypes.c_uint32 * n)()` and pass its length too — never a Python `bytes`, which is immutable.
code
python · 10 linesimport ctypes
class Digest(ctypes.Structure):
_fields_ = [("recipients", ctypes.c_uint32), ("subject", ctypes.c_char * 32)]
d = Digest(1200, b"Nightly digest")
print(ctypes.sizeof(Digest), d.recipients, d.subject)
print(type(ctypes.byref(d)).__name__, type(ctypes.pointer(d)).__name__)go deeper
Know the shape: subclass ctypes.Structure, list the fields in order with their C types, create an instance, and read or set fields as ordinary attributes.
Explain the mechanics: fields are descriptors over one memory block, argtypes decides by-value versus by-pointer, byref differs from pointer, and output buffers must be mutable and length-bounded.
Demonstrate verification and lifetime judgment — asserting sizeof against the C side, round-tripping known bytes through from_buffer_copy, and keeping anything the library retains a pointer to referenced from Python.
Own the interface style: fewer, coarser calls with structs describing the payload rather than chatty per-field calls, and a policy that any struct crossing the boundary has a size assertion and a versioned transcription.
### Declaring the layout A C struct is described by subclassing `ctypes.Structure` and setting the `_fields_` class attribute to an ordered list of `(name, type)` pairs. The order matters, because it is the memory order: ```python class Digest(ctypes.Structure): _fields_ = [("recipients", ctypes.c_uint32), ("subject", ctypes.c_char * 32)] ``` The class is now a real C type. `ctypes.sizeof(Digest)` reports its size in bytes and `ctypes.alignment(Digest)` its alignment. An instance owns a block of memory, and each declared field is a descriptor that reads and writes bytes inside that block — so `d.recipients = 5` mutates the actual bytes a C function will see, and a C function that writes into the block changes what Python reads back afterwards. Fixed-size arrays are spelled `type * n`, nested structs by naming another `Structure` subclass as a field type, and bitfields by adding a third element giving the width in bits. `ctypes.Union` is the same machinery with every field at offset zero. By default the layout follows the platform C compiler's alignment and padding rules, which is exactly what you want: the padding is what makes the Python-side struct match the C-side one. `_pack_` overrides the maximum alignment, for packed on-the-wire structs. Since 3.13 that is coupled with `_layout_`, which names the layout algorithm explicitly — `"ms"` for the MSVC rules, `"gcc-sysv"` for the System V ones. On 3.14, using `_pack_` without `_layout_` emits a `DeprecationWarning` saying the implicit choice is slated to become an error in 3.19, so write both. The first thing to do after transcribing any struct is compare `ctypes.sizeof` against the `sizeof` the C side reports. A mismatch means the layouts differ, and every field after the first divergence is being read from the wrong offset — which produces wrong values rather than an error. ### By value versus by reference Which one happens is decided by `argtypes`, not by the call syntax. **By value:** declare `argtypes = [Digest]` and pass the instance. `ctypes` copies the whole struct into the call the way C would. Struct-by-value is the most ABI-sensitive thing `ctypes` does — how a struct is split across registers is a detailed platform rule — so it is where a wrong declaration bites hardest. **By reference:** declare `argtypes = [ctypes.POINTER(Digest)]` and pass `ctypes.byref(d)`. `ctypes.byref` builds a lightweight `CArgObject` that exists only to be an argument: cheap, and the right default. `ctypes.pointer(d)` instead constructs a real `LP_Digest` object — more expensive, but a first-class value you can store in a variable, put in an array, or assign to a pointer field of another struct. `ctypes.addressof(d)` gives the raw integer address, useful for logging or for passing as a `c_void_p`, but it carries no reference to the object that owns the memory, so nothing keeps that memory alive. ### Output buffers C's usual idiom for returning bytes is caller-allocates. The allocation must be a **mutable** `ctypes` object: - `ctypes.create_string_buffer(n)` makes a zero-filled mutable array of `n` `char`. Read it back with `.raw` for all `n` bytes, or `.value` for the bytes up to the first NUL. `ctypes.create_unicode_buffer(n)` is the `wchar_t` equivalent. - `(ctypes.c_uint32 * n)()` makes a zero-filled array of any other element type, indexable and sliceable from Python after the call. - For a single scalar out-parameter, construct the scalar and pass a reference: `n = ctypes.c_int()`, then `lib.count(ctypes.byref(n))`, then read `n.value`. Never pass a Python `bytes` object as a buffer for C to write into. `bytes` is immutable, `ctypes` converts it to a bare `char *`, and a C function writing through that pointer corrupts a shared — possibly interned — object, silently, and usually far from where the crash finally appears. Sizing is your job as well: pass the buffer's length alongside it, `len(buf)` or `ctypes.sizeof(buf)`, so the library knows the bound, and treat any function that takes a destination pointer with no length parameter as a buffer overflow waiting to happen. ### Lifetime The struct or buffer must stay referenced from Python for as long as the C side can reach it. That is automatic for a call-scoped out-parameter, because `ctypes` keeps the arguments alive for the duration of the call. It is emphatically not automatic when the library stores the pointer for later, nor when you assign a freshly built temporary to a pointer field of another struct: the temporary's last reference disappears at the end of the statement, and the field is left pointing at freed memory. Bind such objects to a name that outlives the C side's use of them. ### Overlaying existing bytes Going the other way — you have bytes and want the struct — the `ctypes.Structure.from_buffer` classmethod overlays the struct on an existing writable buffer with no copy, keeping the two views in sync and keeping the exporter alive, while `ctypes.Structure.from_buffer_copy` copies, and is what you need for an immutable `bytes` payload. `bytes(d)` serialises an instance back out. These are also the cheapest way to test a transcription: round-trip a known byte pattern and check every field.
- When would you use ctypes.pointer instead of ctypes.byref?`ctypes.byref(x)` produces a lightweight object that is only valid as a call argument — it is the cheaper choice and covers most calls. `ctypes.pointer(x)` builds a real pointer object of type `LP_X`, which you need when the pointer must outlive the call: stored in a variable, placed in an array of pointers, or assigned to a pointer field of another struct. The pointer object also keeps a reference to its target, which byref's argument wrapper only does for the duration of the call.
- Why does ctypes.sizeof of your struct sometimes exceed the sum of its field sizes?Alignment padding. The platform's C compiler inserts gaps so each field starts at an address that is a multiple of its alignment, and pads the whole struct to its own alignment so arrays of it stay aligned. A `uint8` followed by a `uint32` occupies 8 bytes, not 5. ctypes reproduces those rules deliberately, because the C side has them too; `_pack_` (with `_layout_`) removes them when you are describing a packed wire format.
- What goes wrong if you pass a Python bytes object where C writes into the buffer?`bytes` is immutable and possibly shared or interned, but ctypes converts it to a plain `char *` with no way to signal read-only. The C function writes through that pointer and corrupts an object other Python code is still using, with no error at the point of damage. Allocate with `ctypes.create_string_buffer(n)` instead, and read `.raw` or `.value` after the call.
saying these in an interview costs you the question
- Thinks _fields_ order does not affect memory layout
- Passes an immutable bytes object as an output buffer
- Expects sizeof to equal the sum of field sizes
- Believes byref copies the struct into the call
- Passes a buffer without telling C its length
- Assigns a temporary array to a struct's pointer field