skip to content

A sensor-telemetry collector buffers millions of float readings in a list - how do you cut the memory per reading?

level: seniorimportance: should knowfreq 38%

answer

  1. One object per number is the problem
  2. Pointer plus boxed float per element
  3. Contiguous typed storage instead of pointers
  4. Fixed-width records in one byte buffer
  5. Slice a view, do not copy bytes

basics

~20 s

Stop making one Python object per reading. A list of floats costs about 32 bytes per value on 3.14 - a pointer plus a boxed float; array.array("d") stores the same doubles contiguously at about 8 bytes each.

solid answer

~40 s

A `list` of a million floats stores a million eight-byte pointers plus a million separate 24-byte `float` objects - roughly 32 bytes per reading for eight bytes of data. Replacing the buffer with `array.array("d", ...)` keeps raw doubles in one contiguous block at about 8 bytes each, a 4x cut with no dependency; if a reading is a record rather than a number, pack fixed-width records into a `bytearray` with `struct.pack_into` and read them back with `struct.unpack_from`. Hand out windows onto that buffer with `memoryview`, whose slices share memory instead of copying. The secondary win matters in a collector: a million individually allocated objects are also a million objects the cycle collector walks, which is often what shows up as an intermittent ingest timeout. Measure with `tracemalloc` before and after rather than assuming.

code

python · 16 lines
python
import array
import tracemalloc

n = 500_000

tracemalloc.start()
as_list = [float(i) for i in range(n)]
list_bytes, _ = tracemalloc.get_traced_memory()
tracemalloc.stop()

tracemalloc.start()
as_array = array.array("d", range(n))
array_bytes, _ = tracemalloc.get_traced_memory()
tracemalloc.stop()

print(list_bytes / n, array_bytes / n)

go deeper

for a junior

Know that a list of numbers is a list of pointers to number objects, and that the standard library has contiguous typed storage - array.array - for when you have a lot of one numeric type.

for a middle

Explain the arithmetic: pointer plus boxed float versus a raw double, why that is roughly 32 bytes against 8, and how struct plus bytearray covers fixed-width records that array.array cannot express.

for a senior

Show the operating judgement: confirm with tracemalloc that the buffer is the memory, convert one hot path rather than the codebase, keep conversions at explicit boundaries, and choose the item width deliberately.

for a principal

Decide the representation at the design level. Row-of-objects versus column-of-buffers sets the service's memory ceiling and its collection pauses, and it is far cheaper to choose once than to migrate a running ingest path later.

### Where the memory actually goes A `list` of a million Python floats does not hold a million doubles. It holds a million **pointers**, eight bytes each, in one contiguous array, and each pointer targets a separate `float` object of 24 bytes (two-word header plus the double). Measured on 3.14 that is about 32 bytes per reading for eight bytes of information — a 4x tax — plus the list's own header and its over-allocation slack. There is a second, quieter cost. Those million objects are individually allocated, individually reference counted, and individually visited by the cycle collector when it runs. In a telemetry collector that is what turns a memory problem into a latency problem: the intermittent ingest timeout that nobody can reproduce is often a collection pass walking a container graph that should never have been a container graph. ### The fix: store the numbers, not objects that hold numbers **`array.array`** is the direct replacement. `array.array("d", values)` keeps raw doubles in one contiguous block: about 8 bytes per element, a 4x reduction with no third-party dependency. It supports `append`, indexing, slicing and iteration; reading an element boxes a `float` on the way out, so a hot per-element Python loop is not faster — it is *smaller*, and it stops being a million objects. **`bytes` / `bytearray` with `struct`** is the move when a reading is a record rather than a number — a timestamp, a sensor id and a value. Pack fixed-width records into one `bytearray` with `struct.pack_into` and read them back with `struct.unpack_from` at an offset. You get one object for the whole buffer. **`memoryview`** is what keeps that cheap at the edges. Slicing a `bytes` object copies; slicing a `memoryview` does not — it hands out a window onto the same buffer. So a parser, a checksum routine or a socket write can work on a region of the buffer without duplicating it, which matters exactly when the buffer is the thing you were trying to shrink. **Columns instead of rows.** If a reading has three fields, three typed arrays of a million entries beat a million three-field objects on both memory and locality. This is the same reason an external array library is shaped the way it is; you can get most of the memory win from the standard library alone. ### The list's own slack Appending grows a list in steps, and it deliberately over-allocates so that repeated `append` stays amortized O(1). Building a 17-element list with `append` reports 248 bytes on 3.14; building the same list at a known size reports 200. On a million-element buffer that slack is a few percent, not the main problem — but if you are pre-sizing anyway, building from a sized iterable avoids it, and `array.array` grows the same way. ### Doing this on a real service The order matters more than the trick: 1. **Measure first.** Confirm the readings are the memory, with `tracemalloc` snapshots diffed across an ingest window, and confirm the timeout correlates with memory pressure rather than with the network. 2. **Change the hot buffer only.** Converting one ingest path from a list of floats to `array.array` is a contained diff a 4-person team can review; converting the whole codebase to typed buffers is not. 3. **Keep the boundary honest.** Something eventually has to hand these numbers to a serializer or a query layer. Convert at that boundary explicitly and measure the conversion cost, rather than letting each consumer box a million floats back into a list. 4. **Pick the width deliberately.** `"d"` is 8 bytes, `"f"` is 4 with less precision, and an integer typecode with a fixed scale is often enough for sensor data. Halving the item size halves the buffer, and it is the cheapest remaining win once the boxing is gone. ### What not to say `__slots__` is not the answer here. It shrinks the *instance* that holds attributes; it does nothing about a million boxed floats, because those floats are not instances of your class. If the shape is "lots of numbers", stop creating one Python object per number.

  • When is `array.array` the wrong answer for this buffer?
    When the readings are heterogeneous records, when downstream code needs a real Python object per reading anyway, or when the work is vectorised arithmetic rather than storage. Then a `bytearray` with `struct` for fixed-width records, or an external array library with its own typed buffers, fits better. `array.array` wins where the data is one numeric type and the code mostly appends, indexes and iterates.
  • What does `memoryview` buy over slicing a `bytes` object?
    Slicing `bytes` or `bytearray` copies the region; slicing a `memoryview` does not - it produces another view onto the same buffer. That lets a parser, a checksum or a socket write operate on a window of a large buffer without duplicating it, which matters precisely when the buffer is the memory you were trying to save.
  • Would `__slots__` help here?
    No. `__slots__` shrinks an instance that holds attributes; the cost in this buffer is a million boxed floats, which are not instances of your class. If each reading were a small object with several fields and you truly needed one object per reading, `__slots__` would trim it - but with plain numbers the answer is typed storage, not leaner instances.

saying these in an interview costs you the question

  • Assumes a Python list of floats stores doubles contiguously
  • Reaches for __slots__ when the data is plain numbers
  • Blames list over-allocation for a 4x memory difference
  • Slices bytes repeatedly and copies the buffer each time
  • Changes the representation without measuring before and after
  • Claims array.array also makes element-by-element loops faster

context