skip to content

questions

4

What do Python's built-in map() and filter() return, and how do you get a list?

level: juniorimportance: must knowfreq 70%

answer

  1. Not a list in Python 3
  2. Lazy, computed on demand
  3. Read it once, then it is spent
  4. len() and indexing both fail
  5. list(), tuple(), or a for loop

basics

~20 s

Both return a lazy, one-shot iterator object, not a list. Nothing is computed until you iterate. Wrap the result in list() or tuple(), or feed it straight to a for loop, sum(), any() or all().

solid answer

~40 s

In Python 3, `map(func, iterable)` returns a `map` object and `filter(pred, iterable)` returns a `filter` object. Both are **iterators**: they hold the source and the callable, compute one result per `__next__` call, and are exhausted after a single full pass. That makes them cheap over large inputs and composable without building intermediate lists, but it also means `len()` and indexing raise `TypeError`, a second `list()` on the same object returns `[]`, and printing one shows a repr rather than the values. Materialize with `list()`, `tuple()` or `set()`, or consume directly in a `for` loop or an aggregate such as `sum()`. In Python 2 both returned real lists; Python 3.0 made them lazy, and that is still the behaviour on 3.14.

code

pycon · 7 lines
pycon
>>> m = map(str.upper, ["a", "b"])
>>> type(m).__name__
'map'
>>> list(m)
['A', 'B']
>>> list(m)
[]

go deeper

for a junior

Recall the one-line fact: map() and filter() give back lazy iterators, so wrap them in list() when you want to print, index, count or reuse the results.

for a middle

Explain the iterator protocol behind it - iter/next, StopIteration, and why one full pass exhausts the object - and show the two-consumption bug rather than just naming it.

for a senior

Show the judgment about where to materialise: keep pipelines lazy over large or streaming inputs, but convert to a list at the boundary where the result becomes shared data, and note that exceptions surface at the consumer.

for a principal

Own the API-boundary call: returning an iterator from a public function makes callers responsible for the one-shot semantics, so decide deliberately whether your library hands back a stream or a materialised collection, and document it.

### Two different things: an iterable and an iterator An **iterable** is anything you can start a loop over — a list, a tuple, a string, a dict. An **iterator** is the cursor that walks it: an object with `__next__()` that yields the next value and finally raises `StopIteration`. Every iterator is also an iterable (its `__iter__` returns itself), which is why a `map` object works in a `for` loop and confuses people into thinking it is a collection. `map(func, iterable)` and `filter(predicate, iterable)` both return iterators. `map` stores the callable and a cursor over the source; each time you ask for the next item it pulls one source element, calls the function on it, and hands back the result. `filter` does the same but pulls until the predicate is truthy, skipping the rest. **No function call happens at the moment you write `map(...)`** — it happens later, once per item, as you consume. ### What laziness buys you Memory is the obvious win. Reading a million-line log and transforming each line with a list comprehension builds a million-element list; `map` over the same file holds one item at a time, so peak memory does not grow with the input. Composition is the second win: `filter(pred, map(func, source))` is a two-stage pipeline that still touches each element once and still holds one element at a time. The third is early exit — `next(filter(pred, items))` finds the first match and stops, without transforming the tail. ### What laziness costs you The cost is that a `map` or `filter` object behaves nothing like the list a beginner expects: * `len(m)` raises `TypeError` — an iterator does not know how many items it will produce, and asking would require consuming it. * `m[0]` raises `TypeError` — there is no indexing protocol, only forward stepping. * `print(m)` prints something like `<map object at 0x...>`, not the values. Wrap it in `list()` to look at it. * **It is one-shot.** After a full pass the underlying cursor is at the end. A second `list(m)` returns `[]`, a second `for` loop body never runs, and a second `sum(m)` is `0`. This is the single most common real bug: code logs `len(list(results))`, then loops over `results` again and silently processes nothing. * An empty result and an exhausted iterator are indistinguishable from the outside, which makes that bug quiet rather than loud. If you need the values more than once, materialize once (`items = list(map(func, source))`) and reuse the list. If you need laziness *and* repetition, keep the source and rebuild the pipeline, or store the source in a list and map over it each time. ### Errors surface late Because the function is not called until consumption, a callable that raises does so from inside `list()`, `sum()` or the `for` statement — not from the line that constructed the `map`. Tracebacks point at the consumer, and a wrapped `try` around only the construction line catches nothing. This is worth knowing before you spend twenty minutes reading the wrong line. ### `filter` and its `None` shortcut `filter(None, iterable)` is a documented special case: passing `None` as the predicate keeps the items that are **truthy**, dropping `0`, `""`, `[]`, `{}`, `None` and `False`. It is not "filter by nothing" and it is not a bug — it is the idiom for dropping blanks. Beyond that special case, `filter` takes exactly one predicate and one iterable. ### The version story In Python 2, `map` and `filter` returned lists, and `map` even padded shorter inputs with `None`. Python 3.0 turned both into lazy iterators as part of the same sweep that made `range`, `dict.keys()` and `zip` lazy. That is still the behaviour on Python 3.14 — no release since has changed it. Interviewers ask this partly because it is the cleanest way to check whether a candidate has internalised the iterator protocol or is still thinking in list-shaped Python 2 habits. ### Practical rule Treat a `map` or `filter` result as a *stream you may read once*. Either consume it in the same expression that created it, or convert it to a list at the boundary where it stops being a pipeline and starts being data.

  • Why does len() fail on a map object when len() works on the list you mapped over?
    `len()` calls `__len__`, which an iterator does not implement, so it raises `TypeError`. An iterator cannot know its length without consuming itself, and consuming it would destroy the values it is meant to hand you. If you need a count, materialise first with `list()` and take the length of that, or count as you consume.
  • What does filter(None, items) do, and how is it different from filter(bool, items)?
    Both keep the truthy items and drop `0`, `""`, `[]`, `{}`, `None` and `False`. Passing `None` is a documented special case handled in C, so it skips a Python-level call per element and is marginally faster; `filter(bool, items)` calls `bool` on every item. Behaviourally they are the same, and `filter(None, ...)` is the conventional spelling.
  • If a function passed to map() raises, which line does the traceback point at?
    The line that **consumes** the iterator — the `list()`, `sum()` or `for` statement — not the line that built the `map` object, because the callable is not invoked until consumption. Wrapping only the construction line in `try` catches nothing. Put the guard around the consumption, or materialise eagerly if you want failures to surface where the pipeline is defined.

A list is a printed page you can reread; a map object is a ticker tape being typed as you read it. Look away and let it run to the end, and the tape is spent — the values were never stored.

saying these in an interview costs you the question

  • Saying map() returns a list in Python 3
  • Calling len() or indexing a map or filter object
  • Assuming the function runs when map() is written
  • Iterating the same map object twice and expecting values
  • Thinking filter(None, xs) filters nothing out
  • Printing a map object and reporting it as empty

context

open as a page

How does Python's map() behave when given several iterables of different lengths?

level: middleimportance: should knowfreq 35%

basics

~20 s

map() takes one item from each iterable per step and calls the function with that many arguments. It stops silently as soon as the shortest iterable runs out, so extra items in the longer ones are never seen.

open as a page

What does the initializer argument to functools.reduce() change about the result?

level: middleimportance: should knowfreq 45%

basics

~20 s

It is the starting accumulator, combined before the first item. It seeds the result type, makes an empty iterable return it rather than raise TypeError, and forces the function to run even for a one-element input.

open as a page

When is map() or filter() clearer than the equivalent comprehension in Python?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Use map() or filter() when the callable is already named: map(str.strip, lines) reads well. Once you write a lambda inside them, a comprehension reads better. Where an aggregate such as sum or any fits, it beats both.

open as a page