skip to content

Compact Binary Sequences

Holding many numbers or bytes without one object per element: array.array's typed storage, bytearray, and memoryview slices that copy nothing. Interviewers ask when a list of ints costs too much.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

7

How does array.array store integers differently from a Python list?

level: juniorimportance: must knowfreq 40%

answer

  1. What one list slot really holds
  2. Pointers to objects versus packed values
  3. One typecode, fixed itemsize, contiguous block
  4. Memory win, but every read boxes

basics

~20 s

A Python list stores pointers to full int objects scattered on the heap. array.array stores raw machine values of one typecode packed back to back, so it holds only numbers of that type and costs far less memory.

solid answer

~40 s

A `list` is a dynamic array of `PyObject*` pointers: for `list(range(1000))` you pay one machine word per slot **plus** a separate 28-byte `int` object for every value that misses CPython's small-integer cache. `array.array('i', range(1000))` allocates one contiguous block of 1000 four-byte C ints — `a.itemsize` is 4 and `sys.getsizeof(a)` is about 4200 bytes, against 8056 bytes for the list's slots alone before you count the int objects they point at. The price is homogeneity and boxing: an `array.array` accepts only values matching its typecode, and every read builds a fresh `int` object out of the packed value, so element access is not faster than a list. It is a storage optimisation, not a compute one — `a * 2` repeats the array exactly like a list rather than doubling each element.

code

python · 8 lines
python
import array, sys

a = array.array('i', range(1000))
lst = list(range(1000))
print(a.typecode, a.itemsize, sys.getsizeof(a))   # i 4 4200
print(sys.getsizeof(lst))                         # 8056 - slots only
print(sys.getsizeof(lst[999]))                    # 28 - one of 743 int objects
print(a[0], a[-1], len(a))

go deeper

for a junior

Be ready to say what one list slot actually holds — a pointer to a separate int object — and that an array.array packs raw machine values instead. Naming typecode and itemsize and knowing the array is memory-cheaper is enough here.

for a middle

Explain the boxing cost: every read out of an array.array builds a fresh int, so you buy memory and not speed. Be able to size both options roughly and to say why sys.getsizeof on a list under-reports its real footprint.

for a senior

Show where the win is real: hundreds of thousands of homogeneous values, or data you hand straight to a file, a socket or C via tobytes. Say plainly that per-element Python loops over an array.array gain nothing at all.

for a principal

Own the decision rule. array.array is a stdlib-only footprint optimisation with no arithmetic; a dedicated numeric array library buys vectorised math at the cost of a dependency and a build story. Frame the choice as memory versus dependency, and set the threshold where it flips.

### What a list slot actually holds A CPython `list` is a dynamic array of `PyObject*` pointers. Each slot is one machine word — eight bytes on a 64-bit build — and it does **not** contain a number; it contains the address of a separate `int` object living elsewhere on the heap. That object carries a type pointer, a reference count and a digit array, and `sys.getsizeof(1000)` reports 28 bytes for it. So `list(range(1000))` is roughly 8 KB of pointers *plus* about 28 bytes for each value large enough to miss the small-integer cache that CPython pre-allocates for −5 through 256 — close to 29 KB all in. `sys.getsizeof` on the list reports only the slot array (8056 bytes), which is exactly why the real footprint surprises people. ### What array.array holds instead `array.array(typecode, initializer)` allocates one contiguous C buffer and packs machine values into it end to end, with no per-element object at all. The one-character typecode, fixed at construction, decides both the C type and the width: `'b'`/`'B'` one byte, `'h'`/`'H'` two, `'i'`/`'I'` four, `'q'`/`'Q'` eight, `'f'` a 4-byte C float, `'d'` a C double. `array.typecodes` lists every code the module accepts, `a.typecode` reports the one this array uses, and `a.itemsize` reports its byte width on the machine you are actually running on. For the same thousand values, `sys.getsizeof(array.array('i', range(1000)))` is about 4200 bytes — 4000 bytes of packed items plus a small header and some growth room — roughly a seventh of the list's true cost. ### What you give up **Homogeneity.** Every element must match the typecode. Push a `str` or a `float` into an integer array and you get `TypeError`; push an integer too large for the width and you get `OverflowError`. There is no room for `None`, no room for a mixed record, no room for a Python object of any kind. **Boxing on every read.** The buffer holds machine values, but Python code can only see Python objects, so indexing an `array.array` builds a brand-new `int` (or `float`) each time, where a list index merely returns the pointer it already had and bumps a refcount. Element access is therefore slightly *slower* than a list, not faster. If you write a Python `for` loop over an `array.array` doing arithmetic, you have paid the memory saving and gained nothing on speed. **No elementwise arithmetic.** `array.array` is a plain sequence type. `a * 2` repeats the contents like a list; `a + b` concatenates and raises `TypeError` if the typecodes differ; there is no `a * 2` that doubles values and no vectorised math anywhere in the module. Elementwise numeric work is the job of a dedicated numeric array library, not of `array`. ### Where the win is real The payoff shows up in two places. First, **footprint at scale**: hundreds of thousands or millions of homogeneous machine numbers held in memory at once — sample buffers, id columns, offset tables, counters — where a seven-fold reduction is the difference between fitting in RAM and not. A few hundred integers is a few kilobytes either way and not worth a typecode. Second, **operations that stay in C**. Because the storage is one contiguous block in native layout, whole-array operations skip Python entirely: `a.tobytes()` hands you the packed bytes, `a.frombytes(data)` fills the array from them, `a.tofile(f)` and `a.fromfile(f, n)` stream to and from a binary file, and the object exposes the buffer protocol so C code can read the same memory without copying. A list of ints can do none of that without a per-element conversion loop. ### The everyday API Apart from being typed, an `array.array` behaves like a mutable sequence: `append`, `extend`, `insert`, `pop`, `remove`, `index`, `count`, `reverse`, slicing, `len`, iteration, and `tolist()`/`fromlist()` to cross back and forth with a list. `append` amortises the same way a list's does — the buffer is grown with headroom rather than reallocated per item — so building an array element by element is not quadratic. ### The decision rule Use a list by default. Reach for `array.array` when the data is (a) large, (b) uniformly one machine numeric type, and (c) either memory-bound or destined for bytes, a file, a socket or C. If what you actually need is arithmetic across the elements, `array.array` is the wrong tool and the answer is a numeric array library — which is a dependency decision, not a language one.

  • Is indexing an array.array cheaper than indexing a list?
    No — usually a shade more expensive. A list index returns the `PyObject*` that is already there and bumps its refcount; `array.array.__getitem__` has to construct a fresh `int` or `float` object from the packed machine value on every read. The saving is memory, plus bulk operations that stay in C such as `tobytes`, `tofile` or handing the buffer to a C library — never per-element Python loops.
  • When would you keep the list even though the array.array is smaller?
    When the values are not one uniform machine numeric type, when you need `None` or arbitrary objects in the sequence, when you rely on elementwise arithmetic, or when the collection is small. The footprint win only matters at hundreds of thousands of elements; below that a list is simpler, faster to read from Python, and costs a few kilobytes either way.
  • Why does sys.getsizeof under-report a list of integers?
    `sys.getsizeof` measures one object, not the graph it references. For a list it returns the object header plus the pointer slots and stops there; the `int` objects the slots point at are separate allocations of about 28 bytes each. For an `array.array` there is nothing to miss — the values live inside the object's own buffer — which is why the two numbers are only comparable if you say so explicitly.

A list is a coat-check board of numbered tickets, each pointing at a coat somewhere in a big room; an array.array is a shelf where identical boxes are stacked edge to edge, and the contents are the shelf.

saying these in an interview costs you the question

  • Claims array.array makes element access faster than a list
  • Thinks a * 2 doubles each element of an array.array
  • Says a Python list stores the integers themselves inline
  • Believes an array.array can hold mixed types or None
  • Quotes sys.getsizeof(list) as the list's full memory cost
  • Reaches for array.array to speed up a numeric Python loop

context

open as a page

Why does slicing a memoryview of a bytearray avoid the copy that slicing the bytearray itself makes?

level: juniorimportance: must knowfreq 35%

basics

~20 s

Slicing a bytearray allocates a second bytearray and copies the bytes. A memoryview slice only records an offset and length into the same storage, so it costs the same tiny amount whatever the slice size.

open as a page

Why does appending 2**31 to an array.array('i') raise OverflowError?

level: middleimportance: should knowfreq 30%

basics

~20 s

The typecode 'i' means a signed C int, four bytes wide on CPython, so the largest value that fits is 2**31 - 1. array.array range-checks every value against the typecode and refuses the write instead of truncating it.

open as a page

Why does assigning into memoryview(b'abc') raise TypeError while memoryview(bytearray(b'abc')) allows it?

level: middleimportance: should knowfreq 40%

basics

~20 s

A bytes object is immutable, so the buffer it exports is flagged read-only and the view refuses writes with TypeError: cannot modify read-only memory. A bytearray exports writable memory, so its view can be written through into the original.

open as a page

Why does a live memoryview make bytearray.append raise BufferError, and how do you scope the view in a long-running ingest worker?

level: seniorimportance: should knowfreq 30%

basics

~10 s

A memoryview holds an export of the bytearray's storage, and a bytearray refuses to reallocate while any export is outstanding. So append, extend and clear raise BufferError; release the view before resizing.

open as a page

What does memoryview.cast() change about a view, and why can a view's len differ from its nbytes?

level: middleimportance: nice to knowfreq 20%

basics

~10 s

memoryview.cast reinterprets the same memory under a different element format without copying, so a 4-element unsigned-int view becomes a 16-element byte view. len counts elements, nbytes counts bytes, so they differ for wider elements.

open as a page

A payment reconciliation job writes array.array('q') with tofile; why can another host read back wrong values?

level: seniorimportance: nice to knowfreq 18%

basics

~20 s

array.array.tofile writes the raw native machine representation — no typecode, no length, no byte-order marker. A reader with different endianness, or one assuming a different typecode, decodes the same bytes into plausible but wrong numbers instead of failing.

open as a page