skip to content

How do you bound itertools.count(), cycle() and repeat() so a loop terminates?

level: juniorimportance: must knowfreq 60%

answer

  1. They never signal the end of iteration
  2. The consumer must do the stopping
  3. Anything that drains it runs forever
  4. islice, takewhile, or zip against something finite
  5. repeat takes an optional times argument

basics

~10 s

itertools.count, cycle and repeat never raise StopIteration, so draining one runs forever. The consumer has to stop them: wrap with itertools.islice or itertools.takewhile, zip against a finite iterable, or break out of the loop.

solid answer

~40 s

The three infinite iterators in itertools are `count(start, step)`, `cycle(iterable)` and `repeat(obj)`. None of them ever raises `StopIteration`, so anything that drains an iterator to exhaustion - a bare `for` loop, `list()`, `sum()`, `sorted()`, a comprehension - runs until you kill the process or memory is gone. Termination is the consumer's job. The four normal ways to supply it: `itertools.islice(it, n)` to take a prefix lazily; `itertools.takewhile(pred, it)` to stop at the first element that fails a predicate; the built-in `zip`, which stops at its shortest argument, so `zip(itertools.count(), rows)` numbers a finite sequence safely; or an explicit `break`. `repeat` is the one exception with a built-in end - `repeat(obj, times)` stops after `times` items.

code

python · 14 lines
python
import itertools

# islice: take a lazy prefix
print(list(itertools.islice(itertools.count(10, 5), 5)))
# [10, 15, 20, 25, 30]

# explicit next(): pull exactly as many as you need
colours = itertools.cycle(["red", "green", "blue"])
print([next(colours) for _ in range(4)])
# ['red', 'green', 'blue', 'red']

# zip stops at the shortest argument
print(list(zip("abc", itertools.repeat(0))))
# [('a', 0), ('b', 0), ('c', 0)]

go deeper

for a junior

Recall the three names - count, cycle, repeat - and that none of them ends by itself. Be ready to write, on the spot, the one-liner that takes the first five values of a counter using itertools.islice.

for a middle

Explain the mechanics: iteration ends when next raises StopIteration, these three never do, so list(), sum() and a bare for loop all run forever. Name all four bounding tools and say which suits a count, a predicate and a parallel sequence.

for a senior

Show the operational judgement: cycle's hidden copy of its input, count's float drift, and the fact that a takewhile predicate which never goes false is still an infinite loop. Be ready to say how you would spot such a loop in a service that is pinned at 100% CPU.

for a principal

Own the API-design angle: exposing a lazy endless iterator pushes the termination decision onto every caller, which is powerful but easy to misuse. Discuss when a library should hand out an unbounded iterator versus a bounded one with an explicit limit parameter.

### Three iterators that never stop `itertools` ships three iterators with no end. `itertools.count(start=0, step=1)` yields `start`, `start + step`, `start + 2 * step` and so on forever. `itertools.cycle(iterable)` yields the elements of its argument in order, then starts over from the beginning, forever. `itertools.repeat(obj)` yields the *same object* again and again, forever - unless you pass its optional second argument, `times`, which is the only built-in stopping condition any of the three has. "Forever" has an exact mechanical meaning. An iterator signals that it is finished by raising `StopIteration` from its `__next__` method; a `for` loop catches that and exits normally. These three never raise it. So every construct that drains an iterator to exhaustion will run until you kill the process: `for x in itertools.count():` without a `break`, `list(...)`, `tuple(...)`, `sum(...)`, `max(...)`, `sorted(...)`, a list comprehension, `str.join`. With `list()` the failure is worse than a hang, because the list grows until the machine starts swapping and then the process dies. None of this is a bug in the interpreter; the loop is doing exactly what it was asked to do. ### Termination is the consumer's job The mental model that makes this easy: an infinite iterator is a *source*, and a source has no opinion about how much you want. Something downstream must decide. There are four normal ways to be that something. **`itertools.islice`** takes a lazy prefix: `itertools.islice(itertools.count(10, 5), 5)` yields `10, 15, 20, 25, 30` and then stops. This is the workhorse, and it is the reason this pairing is taught as a unit - `count` alone is a footgun, `count` plus `islice` is a lazy range with a float or `Decimal` step, or with no end decided in advance. **`itertools.takewhile`** stops on a predicate rather than a count: `itertools.takewhile(lambda n: n * n < 500, itertools.count())` ends as soon as the predicate is false. This is only safe when the predicate is *guaranteed* to go false - a predicate that stays true over an endless source is the same infinite loop with extra steps. **The built-in `zip`** stops at its shortest argument, which makes it a bounding device: `zip(itertools.count(), rows)` attaches an index to each row and ends when `rows` does; `zip(names, itertools.repeat(0))` pairs every name with a default. Note the interaction with `zip`'s `strict=True` keyword (added in 3.10): strict mode demands that all arguments end together, so pairing it with an endless argument raises `ValueError` the moment the finite one runs out. Strict and infinite are mutually exclusive by design. **An explicit `break`** inside the loop is perfectly idiomatic when the stopping condition depends on work done in the body, not on the values alone. ### What each of the three is actually for `count` is the counter you reach for when the built-in `enumerate` does not fit: a step other than 1, a non-integer start, or a counter fed to something other than a `for` loop. It accepts floats, but repeated addition accumulates error - `itertools.islice(itertools.count(0.1, 0.1), 3)` gives `[0.1, 0.2, 0.30000000000000004]` - so for exact float positions compute `start + step * i` from an integer counter, or use `Decimal`/`Fraction`. `cycle` has a memory cost that catches people out: on its first pass it saves a copy of every element it sees so it can replay them later. Cycling a large or expensive iterable buys an unbounded copy of it. Cycling three colour names is fine; cycling a million-row source is a leak. `repeat` yields one object repeatedly - the same object, not copies, which matters if it is mutable. Its two idiomatic uses are supplying a constant argument to `zip` or `map`, and `itertools.repeat(None, n)` as the fastest way to run a loop body exactly `n` times without building a `range` object. ### A word on identity and laziness None of the three builds a container. `itertools.count` holds only its current value and its step, `repeat` holds one object reference, and `cycle` holds its saved copy plus a position. Their `repr` reflects that: after two `next()` calls a counter prints as `count(2)`, which is a handy debugging cue that an iterator is stateful and has already moved. That statefulness is the other half of the model - passing the same infinite iterator to two consumers does not give each of them the full stream, it splits the stream between them. ### The thing to say out loud in an interview That these iterators are *lazy*, that laziness is the point - you can express "all the even numbers" without materialising them - and that the price of laziness is that the boundary moves to the consumer. Show the boundary explicitly and the answer is complete.

  • What does itertools.cycle keep in memory, and when does that become a problem?
    On its first pass `cycle` stores a copy of every element it pulls from the source, because the source may be a one-shot iterator that cannot be replayed. Memory therefore grows to the full size of the input. Cycling a handful of labels is free; cycling a large or lazily-produced sequence buys an unbounded copy of it, and you are better off re-creating the source each pass.
  • How does the strict=True keyword of the built-in zip interact with an infinite iterator?
    Badly, and deliberately so. Plain `zip` stops at its shortest argument, which is what makes `zip(itertools.count(), rows)` a safe bound. `strict=True` inverts that contract: it requires every argument to end at the same point and raises `ValueError` as soon as one runs out while another still has items. Against an endless argument that error is guaranteed, so never combine the two.
  • When would you reach for itertools.count() instead of the built-in enumerate?
    `enumerate` already pairs an index with each item of one iterable, so prefer it for the common case. Use `count` when you need a step other than 1, a non-integer or non-zero start, a counter shared across several `zip` arguments, or a counter consumed by `next()` outside any loop - for example handing out sequential ids.

Think of count, cycle and repeat as a tap with no float valve: the tap will not stop on its own, so the size of the bucket you hold under it is the only thing that decides how much you get.

saying these in an interview costs you the question

  • Says a for loop over itertools.count() eventually stops on its own
  • Calls list() on itertools.cycle() to inspect what it holds
  • Thinks itertools.repeat can never be given an end
  • Tries to bound an infinite iterator with len()
  • Assumes itertools.cycle replays without storing the elements
  • Believes the producer rather than the consumer decides the length

context