A payment reconciliation job writes array.array('q') with tofile; why can another host read back wrong values?
answer
- Raw bytes carry no self-description
- Two machines must agree on layout
- Native byte order, native item size
- byteswap, or an explicit struct format
- Plausible wrong numbers, not an exception
basics
~20 sarray.array.tofile writes the raw native machine representation — no typecode, no length, no byte-order marker. A reader with different endianness, or one assuming a different typecode, decodes the same bytes into plausible but wrong numbers instead of failing.
solid answer
~50 s`tobytes()` and `tofile()` emit exactly the packed items as the current machine stores them — native byte order, native item size, and nothing else. The file is not self-describing, so `fromfile(f, n)` cannot detect a mismatch: it happily reinterprets the bytes under whatever typecode you gave the receiving array. Read `'q'` data as `'i'` and every record silently splits in two. The only errors you *do* get are structural: `frombytes` raises `ValueError` when the byte count is not a multiple of `itemsize`, and `fromfile` raises `EOFError` when the file runs short — while keeping the items it managed to read. The fixes are to record `sys.byteorder` and the typecode in a header and `byteswap()` on mismatch, or to stop using a raw dump for interchange and write an explicit `struct` format such as `'<q'` where width and byte order are stated in the format string.
code
python · 16 linesimport array, pathlib, sys, tempfile
amounts = array.array('q', [1250, -400, 99]) # cents
path = pathlib.Path(tempfile.mkdtemp()) / 'amounts.bin'
with path.open('wb') as f:
amounts.tofile(f)
print(path.stat().st_size, amounts.itemsize, sys.byteorder) # 24 8 little
back = array.array('q')
with path.open('rb') as f:
back.fromfile(f, 3)
print(back.tolist()) # [1250, -400, 99] - same host, same layout
foreign = array.array('q', back)
foreign.byteswap() # what the other endianness would decode
print(foreign.tolist())go deeper
Know that array.array.tofile and tobytes write only the raw packed bytes — no typecode, no count, no header — and that reading them back means telling frombytes or fromfile which typecode to use in advance.
Explain that the bytes are in native byte order and native item size, so a dump is portable only between machines that agree on both. Name sys.byteorder for detection and byteswap() for repairing a mismatch in place.
Diagnose the quiet version of this failure: wrong endianness or a wrong typecode yields plausible numbers, not an exception, so the job reconciles against garbage. Validate a known record and an independent total before trusting a batch, and remember fromfile leaves partial state on EOFError.
Decide whether a raw machine dump belongs in an interchange path at all. Weigh a bulk tofile against a self-describing format carrying width, byte order and schema version, and make the resulting contract explicit so the next reader is not guessing from code.
### What actually goes into the file `array.array.tofile(f)` writes the array's internal buffer to a binary file object verbatim, and `tobytes()` returns the same buffer as a `bytes`. That buffer is the machine's own representation: each item occupies `itemsize` bytes in the platform's native byte order. There is no header, no typecode byte, no element count, no version field. Three items of `'q'` are 24 bytes and nothing more. That is a feature for speed — dumping ten million integers is one write of a contiguous block, with no per-element Python work — and a trap for interchange, because every piece of context needed to interpret the bytes lives outside the file, in the code of whoever reads it. ### The three ways a reader gets it wrong **Byte order.** `sys.byteorder` is `'little'` on x86-64 and on Apple silicon, but big-endian hosts exist (and network payloads are conventionally big-endian). Read little-endian `'q'` bytes on a big-endian host and each 8-byte group is reversed, so 1250 becomes 8863084066665136128-ish nonsense. Nothing raises: the bit pattern is a perfectly valid integer. **Item width.** `'q'` is at least 8 bytes everywhere, but `'l'` is 8 on 64-bit Unix and 4 on Windows. A file written from an `'l'` array on one and read into an `'l'` array on the other silently reframes every record. This is why a shared format should never use a typecode whose width is platform-defined. **Typecode confusion.** The reader constructs `array.array('i')` where the writer used `'q'`, and 3 records become 6 numbers — the low half and the high half of each original value, most of which are small and look believable. All three failures share a shape a reconciliation job cannot tolerate: the output is *plausible*. A totals check against an independent source is what catches it, not an exception. ### The errors you do get `frombytes(data)` raises `ValueError: bytes length not a multiple of item size` when the buffer does not divide evenly — a partial-write or truncated-transfer detector, and the one alignment mistake the module will tell you about. `fromfile(f, n)` raises `EOFError: read() didn't return enough bytes` when the file holds fewer than `n` items, **and keeps the items it did read**, so the array is left partially populated exactly as a failed `extend` leaves it. Code that catches `EOFError` and retries has to rebuild the array, not top it up. ### Making it safe There are three honest options, in increasing order of discipline. 1. **Keep the raw dump, add a header.** Write a small fixed prefix recording the typecode character, `itemsize`, `sys.byteorder` and a count, then on read compare it to the local values and call `byteswap()` when the byte order differs. `byteswap()` reverses the bytes of every item in place and only works for item sizes 1, 2, 4 and 8. Cheap, and it turns a silent misread into a checked conversion. 2. **Use an explicit format instead.** `struct.pack('<q', value)` and `struct.unpack` state width and byte order in the format string itself: `'<'` little-endian, `'>'` big-endian, `'!'` network order, and standard sizes independent of the platform. `struct.calcsize('<q')` is 8 on every machine, whereas `struct.calcsize('@l')` follows the local C `long`. Slower than a bulk `tofile`, and correct by construction. 3. **Do not use a machine dump for interchange at all.** If the file crosses a team boundary or outlives a release, a self-describing format that carries schema, version and byte order pays for itself the first time someone adds a column. ### The judgement call A useful rule: `array.array.tofile` is an excellent *cache* and a poor *contract*. Inside one process, or one host writing a scratch file it will re-read itself with the same code, the raw dump is the fastest thing available and the absent header costs nothing. The moment a second reader exists — a different service, a different architecture, or a colleague in a four-person team who will read the file next quarter with code they wrote fresh — the missing metadata becomes an unwritten contract, and unwritten contracts on financial data fail quietly. Write the header, or write the `struct` format, and validate a known record before trusting a batch.
- What state is the array in after fromfile raises EOFError?Partially populated. `fromfile(f, n)` appends everything it managed to read before the file ran out, then raises `EOFError`, so an array asked for five items from a two-item file ends up holding those two. Treat it like a failed `extend`: discard the array and rebuild, or record the length you had before the call, otherwise a retry duplicates the prefix.
- How would you make an array.array dump safe to read on another architecture?Write a small header recording the typecode, `itemsize`, `sys.byteorder` and the item count, then on read compare it against the local values and call `byteswap()` when the byte order differs — it reverses each item in place for widths 1, 2, 4 and 8. If the file is a real contract rather than a cache, prefer `struct` with an explicit format such as `'<q'`, where width and byte order are stated rather than inherited.
- Which errors will array.array raise on a corrupt dump, and which will it not?It raises `ValueError` from `frombytes` when the byte count is not a multiple of `itemsize`, and `EOFError` from `fromfile` when the file is short. It raises nothing at all for wrong byte order, a wrong typecode of the same width, or a wrong item width that still divides evenly — those decode into valid-looking numbers, which is why a batch needs an independent totals check.
saying these in an interview costs you the question
- Assumes tofile writes a typecode or length header
- Expects an exception when byte order does not match
- Thinks tofile output is portable across architectures
- Believes fromfile leaves the array untouched on EOFError
- Uses a platform-width typecode such as 'l' in a shared file
- Trusts a decoded batch without an independent totals check