Why does appending 2**31 to an array.array('i') raise OverflowError?
answer
- One character decides the whole array
- Bytes per item, and what fits
- Loud on integers, quiet on floats
- OverflowError, and extend stops mid-batch
basics
~20 sThe typecode 'i' means a signed C int, four bytes wide on CPython, so the largest value that fits is 2**31 - 1. array.array range-checks every value against the typecode and refuses the write instead of truncating it.
solid answer
~50 sAn `array.array`'s one-character typecode fixes the C type of every slot, and `a.itemsize` reports how many bytes that is on the current machine: `'b'` 1, `'h'` 2, `'i'` 4, `'q'` 8, `'f'` 4, `'d'` 8. Because `'i'` is a signed 4-byte int, `2**31` is out of range and CPython raises `OverflowError: signed integer is greater than maximum` — it never wraps or truncates. The same check rejects a negative on an unsigned code such as `'I'`, while a wrong Python type (a `str`, or a `float` into an integer array) raises `TypeError` instead. Two edges bite in production: `extend` applies items one at a time, so a rejected item leaves everything before it already appended; and the float typecodes do **not** raise — `'f'` silently rounds to 4-byte precision and turns an over-large value into `inf`.
code
python · 13 linesimport array
ints = array.array('i', [9])
print(ints.typecode, ints.itemsize) # i 4 -> fits -2**31 .. 2**31-1
try:
ints.extend([1, 2, 2**31])
except OverflowError as exc:
print(exc) # signed integer is greater than maximum
print(ints.tolist()) # [9, 1, 2] - the prefix was applied
try:
ints.append(1.5)
except TypeError as exc:
print(exc) # wrong kind, not wrong magnitudego deeper
Recall that an array.array is typed: the one-character typecode fixes both which values fit and how many bytes each takes, and a.itemsize reports that width. Knowing that an out-of-range integer raises rather than wraps is the core point.
Explain the range check itself: integer typecodes raise OverflowError instead of wrapping, an unsigned code rejects negatives with the same error, and a wrong type raises TypeError. Know that 'l' is 8 bytes on 64-bit Unix but 4 on Windows.
Demonstrate the failure you actually meet in production: extend applies items until one is rejected, so a failed batch leaves a partly-filled array and a naive retry duplicates the prefix. Add that the float typecodes never raise — 'f' rounds and saturates to inf silently.
Take a position on typecode policy for data the team shares: prefer widths that are guaranteed ('q'/'Q'), assert itemsize at startup, or move to an explicit struct format so no one's platform quietly redefines the record. Decide it once and write it down.
### The typecode is a C type `array.array` takes a one-character typecode at construction and never changes it. That character names a C type, and the C type decides two things at once: which values are representable, and how many bytes each element occupies. `array.typecodes` is the full string of accepted codes; the numeric ones are `'b'`/`'B'` (signed/unsigned char, 1 byte), `'h'`/`'H'` (short, at least 2), `'i'`/`'I'` (int, at least 2 by the language spec but 4 on every mainstream CPython build), `'l'`/`'L'` (long, at least 4), `'q'`/`'Q'` (long long, at least 8), `'f'` (float, 4) and `'d'` (double, 8). The documented sizes are *minimums*, not promises. `'l'` is 8 bytes on 64-bit Linux and macOS but 4 bytes on Windows, because it maps to the platform's C `long`. That is why the authoritative answer to "how wide is this array?" is never a table you memorised but `a.itemsize` on the machine you are running on — and why, when a width has to be stable across machines, you pick `'q'`/`'Q'` or verify `itemsize` at startup. ### The range check, and why it is loud When you assign or append an integer, CPython converts the Python `int` to the C type and checks that it fits. `array.array('i').append(2**31)` therefore raises `OverflowError: signed integer is greater than maximum`, and `-2**31 - 1` raises the mirror-image `... is less than minimum`. An unsigned typecode rejects negatives the same way: `array.array('I').append(-1)` raises `OverflowError: can't convert negative value to unsigned int`. This is the opposite of C, where assigning an out-of-range value to an `int` wraps or is undefined. Python refuses. That is the design you want: a reconciliation job that stores amounts in cents as `'i'` and meets a value above 2.1 billion stops with an exception instead of writing a wrapped negative that reconciles to nothing. A *type* mismatch is a different error. `array.array('i').append(1.5)` raises `TypeError: 'float' object cannot be interpreted as an integer`, as does appending a `str`. Distinguish the two in an interview: `OverflowError` means "right kind, wrong magnitude"; `TypeError` means "wrong kind". ### Two edges that cause real bugs **`extend` is not atomic.** It consumes the iterable item by item and appends as it goes, so when item *k* is rejected, items 0..*k*-1 are already in the array: ```python import array a = array.array('i', [9]) try: a.extend([1, 2, 2**31]) except OverflowError: print(a.tolist()) # [9, 1, 2] - partially applied ``` If that array is a batch accumulator, catching the `OverflowError` and retrying the batch double-appends the accepted prefix. Either rebuild the array from scratch on failure, or validate the batch before extending. **The float typecodes are quiet.** There is no range check that will save you on `'f'`. `array.array('f', [0.1])[0]` is `0.10000000149011612`, because 0.1 was rounded to the nearest 4-byte float; appending `1e40` stores `inf` without a word. So the silent-truncation failure people fear on the integer side is real — it just lives on the float side. Money never belongs in `'f'`, and rarely belongs in a binary float at all; integer cents in `'q'` is the honest representation. ### Choosing a typecode Work backwards from the value range and the lifetime of the data. For values that stay inside one process and comfortably inside 32 bits, `'i'` halves the memory of `'q'`. For anything that crosses a machine, a file or a year of growth, `'q'` costs 8 bytes and removes a whole class of incident. For unsigned counters `'Q'`; for byte-oriented data `'B'`. When the array is a fixed-width record shared with another system, do not infer the width — assert it, or describe the record with an explicit `struct` format string where the width is written down rather than inherited from the platform. ### The Unicode typecodes Two non-numeric codes exist and are worth a sentence: `'w'`, added in 3.13, stores 4-byte Unicode character codes, and the older `'u'` has been deprecated since 3.3 — on 3.14 it still constructs but emits a `DeprecationWarning`, and it is scheduled for removal in 3.16. New code that wants an array of characters uses `'w'`.
- How wide is the 'l' typecode, and why is that the wrong question?It depends on the platform: 8 bytes on 64-bit Linux and macOS, 4 on Windows, because it maps to the C `long`. The documented figure is a minimum, not a guarantee, so the answer at runtime is `array.array('l').itemsize`. When a width must be fixed — a file format, a shared record — use `'q'`/`'Q'`, which are at least 8 everywhere, or assert `itemsize` at startup.
- Which error tells you the value was the wrong type rather than out of range?`TypeError`. Appending `1.5` or a `str` to an integer array raises `TypeError: 'float' object cannot be interpreted as an integer` — wrong kind of value. `OverflowError` means the kind was right but the magnitude did not fit the typecode, including the unsigned case where a negative raises `can't convert negative value to unsigned int`.
- Would you store monetary amounts in an array.array('f')?No. `'f'` is a 4-byte binary float, so 0.1 comes back as 0.10000000149011612 and large values saturate to `inf` with no exception at all — the failure is silent and cumulative. Store integer minor units (cents) in `'q'`, where the range check is loud and the arithmetic is exact, and convert to a decimal type only for display.
saying these in an interview costs you the question
- Says the value is silently truncated to 32 bits
- Believes typecode widths are identical on every platform
- Assumes extend is atomic and rolls back on failure
- Confuses OverflowError with TypeError for a wrong-type value
- Thinks the 'f' typecode also raises on an over-large value
- Claims array.array widens its items automatically as needed