skip to content

With int.to_bytes, what must a binary-index writer pin so another machine decodes the same values?

level: seniorimportance: should knowfreq 38%

answer

  1. The int has no width until you pick one
  2. Three decisions, and both sides must agree
  3. Only the too-big case raises
  4. Length, byteorder, signed
  5. Defaults are 1 byte and big-endian since 3.11

basics

~20 s

Pin three things and record them in the format: the byte length, the byteorder, and whether the value is signed. int.from_bytes must use the identical three. A value too large for the length raises OverflowError; a mismatched byteorder corrupts silently.

solid answer

~50 s

A Python `int` has no width, so encoding one is where you choose a width - and every part of that choice has to be agreed with the reader. `int.to_bytes(length, byteorder, signed=False)` fixes the number of bytes, the ordering, and whether negatives are representable; `int.from_bytes` must be called with the same three. Get the length wrong on write and you get a loud `OverflowError`; get the byteorder or signedness wrong on read and you get a wrong number with no exception at all, which is far worse. Since Python 3.11 both calls default to `length=1` and `byteorder='big'`, so a format that relies on defaults silently encodes something nobody intended. In a persisted index, write the width and byteorder into a header, pick a width with headroom, validate that the byte count read back is a multiple of the record size, and round-trip the encoder in a test.

code

python · 13 lines
python
value = 1000
raw = value.to_bytes(4, 'big')
assert raw.hex() == '000003e8'
assert int.from_bytes(raw, 'big') == value
assert (-1).to_bytes(2, 'big', signed=True).hex() == 'ffff'
assert int.from_bytes(bytes.fromhex('ffff'), 'big', signed=True) == -1

try:
    (300).to_bytes(1, 'big')
except OverflowError as exc:
    print('overflow:', exc)

print('read with the wrong order:', int.from_bytes(raw, 'little'))

go deeper

for a junior

Know that int.to_bytes needs a byte length and a byteorder, that int.from_bytes reverses it with the same arguments, and that a value too large for the length raises OverflowError rather than truncating.

for a middle

Explain all three parameters including signed, why two's complement makes the signed range half as wide, and why relying on the 3.11 defaults of one byte and big-endian is a bug waiting in any persisted format.

for a senior

Show the operational judgment: overflow is loud but a byteorder or length mismatch is silent, so the format needs a versioned header, a record-size validation on read, generous widths, and closing writers through a context manager so nothing is lost unflushed.

for a principal

Own the format contract itself: who may change the width, how a version bump is rolled out across writers and readers, and whether a self-describing serialization framework should replace hand-rolled fixed-width records at your scale.

## Where the width comes from A Python `int` is unbounded, so it carries no width of its own. The moment it has to leave the process - a binary index file, a cache key, a length prefix on a socket, a checksum field - you must choose a width, and `int.to_bytes` is where you choose it. The API takes three decisions: - **length**: how many bytes the value occupies. - **byteorder**: `'big'` or `'little'`. `sys.byteorder` tells you the platform's native order, which is exactly what you should not use for a format two machines share. - **signed**: whether negatives are representable, using two's complement within the chosen length. `int.from_bytes` is the inverse and takes the same three. Nothing about the encoded bytes records those choices; a byte string is just bytes. The agreement lives in your format, or it does not exist. ```python value = 1000 raw = value.to_bytes(4, 'big') assert raw.hex() == '000003e8' assert int.from_bytes(raw, 'big') == value assert (-1).to_bytes(2, 'big', signed=True).hex() == 'ffff' ``` ## The two failure modes, and only one of them is loud Overflow is loud. `(300).to_bytes(1, 'big')` raises `OverflowError`, and so does `(-1).to_bytes(2, 'big')` without `signed=True`, because an unsigned encoding cannot hold a negative value. Those are the good failures: they happen at the write, at the moment the value first exceeds your assumption. Everything else is silent. Read a big-endian record as little-endian and you get a plausible, entirely wrong integer. Read four bytes as though they were eight and you get the wrong number with no complaint. Read an empty slice and `int.from_bytes` returns `0` - a valid-looking offset. Decoding cannot detect any of this, because it has no expectations to violate. ## A worked failure A translation-memory updater appended segment offsets to a side index, one fixed-width record per segment, writing each offset with `to_bytes`. Two things went wrong at once. The writer wrote through a file object that was never closed - the run finished, the process exited on a path that skipped the close, and the last buffered block never reached disk. The reader then decoded a truncated tail: a short final slice produced a small offset rather than an error, so lookups for the newest segments returned text from the beginning of the memory instead of failing. The other half was the width: the offsets had been encoded in 4 bytes, and once the memory grew past 4 GiB the writer started raising `OverflowError` in production for exactly the segments nobody had test data for. Because the integration suite took 27 minutes, both problems were discovered a full run after the change that caused them. The fixes are structural, not clever. Write through a context manager so the file is closed and flushed on every exit path. Choose 8 bytes for an offset, not 4 - the width costs nothing at index scale and buys decades of headroom. Put a magic marker and a format version in a header so an old reader refuses a new file instead of misreading it. Validate on read that the payload length is an exact multiple of the record width and raise if it is not. And round-trip the encoder in a unit test with the boundary values: zero, the maximum for the width, and the first value that must overflow. ## The defaults are a trap in a format Since Python 3.11, `length` defaults to 1 and `byteorder` defaults to `'big'` in both `to_bytes` and `from_bytes`; before that, `byteorder` was required. The defaults are pleasant in a REPL and dangerous in a file format, because `value.to_bytes()` silently encodes one byte and raises for anything above 255. In persistence code, pass all three arguments explicitly - it is also the only way a reviewer can check the format from the call site. ## Choosing an ordering Big-endian is the conventional choice for anything crossing a machine boundary, and it has a concrete property worth knowing: for unsigned values of equal width, byte-wise lexicographic comparison of big-endian encodings matches numeric comparison. That makes big-endian fixed-width integers directly usable as sorted keys in an ordered store, which little-endian encodings are not. Signed values break the property unless you bias them first. ## When to reach past to_bytes `int.to_bytes` handles one integer. A record with several fields, mixed widths or floats is better served by the `struct` module, whose format strings pin both the byte order and the field widths in one place and whose prefix characters let you demand a standard size instead of the platform's native layout. `from_bytes` accepts any bytes-like object - `bytes`, `bytearray`, a `memoryview` slice - so you can decode straight out of a buffer without copying. What does not change is the discipline: the width, the order and the signedness are part of the contract, and the contract belongs in the format, not in the reader's memory.

  • What happens if the reader uses the wrong byteorder, and how would you catch it?
    Nothing raises - the bytes are simply reinterpreted, so you get a wrong number that looks legitimate. Nothing in the encoding records the order, so detection has to come from the format: a magic marker and version in a header, a checksum over the payload, or a sanity range check on decoded values. A round-trip test alone will not catch it, because a symmetric bug in one process is self-consistent.
  • How do you encode a negative offset, and what does length have to be?
    Pass `signed=True` on both sides; the value is encoded in two's complement inside the chosen length, so `(-1).to_bytes(2, 'big', signed=True)` is the two bytes `ffff`. The length must fit the signed range, which is half the unsigned range for the same width - `-129` does not fit in one signed byte and raises `OverflowError`. Forgetting `signed=True` on read decodes the same bytes as a large positive number.
  • When would you use the struct module instead of int.to_bytes?
    When the unit is a record rather than a single integer: several fields, mixed widths, or floats. A `struct` format string pins the byte order and every field width in one reviewable place and encodes or decodes the whole record in one call, and its prefix characters let you demand a standard size rather than the platform's native layout and alignment. For one integer with an explicit length and order, `int.to_bytes` is clearer.

Encoding an int is like writing a number in a fixed-size box on a paper form: how many boxes, which end you start from, and whether a minus sign is allowed all have to be agreed with whoever reads the form.

saying these in an interview costs you the question

  • Relies on the default length and byteorder in a file format
  • Thinks the encoded bytes describe their own byteorder
  • Omits signed=True and is surprised by OverflowError
  • Trusts a short slice decoded by from_bytes
  • Uses the platform native order for a shared format
  • Assumes from_bytes validates the length it is given

context