skip to content

How does a Python range object differ from a list of the same numbers?

level: juniorimportance: must knowfreq 72%

answer

  1. Ask what the object actually stores
  2. Three integers, not n integers
  3. Memory is constant in the length
  4. start, stop, step recomputed on demand
  5. sys.getsizeof(range(10**9)) is 48

basics

~20 s

A range object stores only start, stop and step, so it occupies the same few dozen bytes whether it spans ten values or a billion. A list of the same numbers materializes every int in memory.

solid answer

~40 s

`range` is an immutable **sequence** type that holds three integers — `start`, `stop` and `step` — and computes everything else on demand. `sys.getsizeof(range(10**9))` is 48 bytes, the same as `sys.getsizeof(range(10))`; the equivalent list allocates a pointer array plus an `int` object per element. It is lazy but it is not an iterator: it supports `len()`, indexing, slicing, `in`, `.index()` and `reversed()`, and you can iterate the same range object repeatedly because each loop asks it for a fresh iterator. The tradeoff is that a range only ever describes evenly spaced integers and cannot be mutated, so you convert to a list when you need to sort, shuffle, append or hold arbitrary values.

code

pycon · 7 lines
pycon
>>> import sys
>>> sys.getsizeof(range(10))
48
>>> sys.getsizeof(range(10**9))
48
>>> len(range(10**9))
1000000000

go deeper

for a junior

Recall the headline: a range holds start, stop and step, not the values. Be able to say that range(10**9) is instant and tiny while list(range(10**9)) is not, and that you convert to a list only when you need to mutate.

for a middle

Explain the mechanics: indexing is start + i * step with a bounds check, len() is a ceiling division, and the object is a Sequence rather than an iterator, so it is reusable and supports len, slicing and in.

for a senior

Show the production judgment. Point out where list(range(...)) quietly becomes a memory problem or a stale cached value, and be able to reason about resident-set impact for large n rather than just quoting a rule.

for a principal

Own the tradeoff framing: computed sequences buy constant memory at the cost of immutability and arithmetic-only content. Be ready to say when a team should prefer a lazy computed view over a materialized collection, and where that discipline stops paying.

### What a `range` object actually is In Python 3, `range` is a built-in **immutable sequence type** — not a function that returns a list, and not a generator. Calling `range(0, 6800, 500)` builds one small object that stores three integers, exposed as the read-only attributes `start`, `stop` and `step`. Everything else the object can tell you is *computed* from that triple on demand. Because the triple is all it holds, the object's footprint does not grow with its length: ```pycon >>> import sys >>> sys.getsizeof(range(10)) 48 >>> sys.getsizeof(range(10**9)) 48 ``` A list of the same numbers is a different animal. It allocates a resizable array of pointers plus a separate `int` object for every value outside CPython's small-integer cache. `list(range(1000))` costs roughly 8 KB for the pointer array alone, and a list of ten million numbers runs into hundreds of megabytes once the `int` objects are counted. ### Lazy, but emphatically not an iterator The most common mistake is calling a range "a generator". A generator object is *one-shot*: iterate it once and it is exhausted, it has no `len()`, and you cannot index it. A range is a full sequence. It registers as a `collections.abc.Sequence`, and it supports `len()`, indexing, slicing, `in`, `.index()`, `.count()` and `reversed()`. You may iterate the same range object any number of times, because each `for` statement asks it for a fresh iterator. It is lazy in the sense that no element exists until you ask for one; it is not lazy in the sense of being consumed. Indexing is arithmetic rather than a lookup: `r[i]` is `start + i * step`, bounds-checked against a computed length, so `range(3, 20, 4)[2]` is `11` without ever producing the two values before it. `len()` is arithmetic too — a ceiling division of `stop - start` by `step`. ### What the cheapness costs you Two consequences follow from "three integers, computed on demand". *Immutability.* You cannot rebind `r.start` (it raises `AttributeError: readonly attribute`), and item assignment raises `TypeError: 'range' object does not support item assignment`. In exchange, a range is hashable and makes a perfectly good dict key. *Only arithmetic progressions.* A range holds evenly spaced integers and nothing else. The moment you need to shuffle, filter in place, sort, append, or hold values that are not evenly spaced, you need a real list — and that is when you pay for the memory. One more edge worth carrying: the computed length must fit a C `Py_ssize_t`, so an enormous range still indexes fine while `len()` overflows. ```pycon >>> r = range(0, 10**100) >>> r[10**50] 100000000000000000000000000000000000000000000000000 >>> len(r) Traceback (most recent call last): OverflowError: Python int too large to convert to C ssize_t ``` ### Where the difference bites in real code Consider an ETL export that ships rows to a warehouse in fixed chunks. A chunk plan written as `range(0, total_rows, 500)` is rebuilt in constant time and constant memory on every run, so it always reflects the row count the job just measured. Materialize it once as `list(range(0, total_rows, 500))` and stash it at module import, and you have manufactured a stale cached value: when the batch grows to 6,800 rows, the cached offsets still describe yesterday's smaller export and the tail rows silently never ship. The range costs nothing to recompute, which is precisely why nobody is tempted to cache it. The second place it bites is scale intuition. Candidates who "know range is lazy" still reach for `list(range(n))` because "Python is fast". That is true at n = 1,000 and a very different conversation at n = 100,000,000, where the list is the difference between a resident set of a few megabytes and a killed process. ### The Python 2 history that still leaks into advice In Python 2, `range()` returned a real list and xrange() was the lazy object. Python 3 deleted xrange and made `range` itself the lazy sequence. Any guidance that says "use xrange for big loops" is Python 2 guidance; on 3.14 there is nothing to switch to, and wrapping a range in `list()` is nearly always a step backwards. ### The rule of thumb Use `range` for loop counters, chunk offsets, index arithmetic and membership tests over an arithmetic progression. Convert to a list only when you need the values as a mutable container — and when you do write `list(range(...))`, be able to say out loud why you needed the list.

  • Is a range object an iterator?
    No. An iterator is one-shot, has no `len()` and cannot be indexed. A range is a reusable immutable sequence: each `for` statement calls `iter()` on it to get a fresh range-iterator, so looping twice yields the same values twice. It also registers as a `collections.abc.Sequence`.
  • When would you still build list(range(n))?
    When you need the values as a mutable container — to shuffle, sort, append, assign by index, or pass to an API that requires a real list. Also when you will index the same values many times and want them materialized. Outside those cases the list only adds allocation and memory.
  • Does range(0, 10**100) work?
    Constructing it works, and indexing works with arbitrary-precision integers: `range(0, 10**100)[10**50]` returns a 51-digit number. But `len()` must fit a C `Py_ssize_t`, so `len(range(0, 10**100))` raises `OverflowError`. The lesson is that a range is arithmetic, and only the length is bounded by the C layer.

A list of numbers is a printed lookup table; a range is the formula printed on one line. The formula answers any row you ask for without the paper.

saying these in an interview costs you the question

  • Says range() returns a list in Python 3
  • Calls a range object a generator
  • Thinks range(10**9) allocates a billion int objects
  • Believes a range can only be iterated once
  • Claims you can reassign r.start to reuse a range
  • Reaches for xrange, which does not exist in Python 3

context