skip to content

How does re.findall differ from re.finditer, and when do you need each?

level: middleimportance: should knowfreq 55%

answer

  1. One returns a list, one an iterator
  2. One throws the positions away
  3. Adding parentheses changes one return type
  4. Match objects carry span and group
  5. Large input, early exit, offsets

basics

~20 s

re.findall builds a list of strings (or tuples, once the pattern has capture groups) for every non-overlapping match. re.finditer returns a lazy iterator of Match objects, so you keep offsets and spend no memory on the full list.

solid answer

~40 s

Both walk the string once and find the same non-overlapping matches; they differ in what they hand back. `re.findall` returns a fully-materialised `list` of **strings** when the pattern has no capture group, a list of the captured text when it has exactly one, and a list of tuples when it has several — so adding a group to a working pattern silently changes the return shape, which is its notorious trap. `re.finditer` returns a lazy iterator of `re.Match` objects, so you keep everything the match knows: `Match.group()`, `Match.start()`, `Match.end()`, `Match.span()`. Use `findall` for a quick list of matched text, `finditer` when you need positions, want to stop early, or are scanning input large enough that materialising every match matters.

code

python · 6 lines
python
import re

text = "clip_720p.mp4 clip_1080p.mp4"
print(re.findall(r"\d+p", text))      # ['720p', '1080p']
print(re.findall(r"(\d+)p", text))    # ['720', '1080']
print(re.findall(r"(\d+)(p)", text))  # [('720', 'p'), ('1080', 'p')]

go deeper

for a junior

Recall that one of the two gives you plain strings in a list and the other gives you Match objects you loop over. Be able to say which one still knows where each match was found.

for a middle

Explain the group-dependent return shape of findall — no groups, one group, several groups — and the laziness of finditer. This is the layer beneath everyday use that an interviewer is probing for.

for a senior

Show the operational judgment: choose based on offsets needed, early exit and input size, and treat findall on unbounded input as a memory decision. Mention counting via a generator rather than materialising a list.

for a principal

Frame it as an API-contract risk: a return type that silently changes when someone adds parentheses is the kind of thing you want caught by types and tests, or avoided by standardising on finditer in shared parsing code.

## Same scan, different product `re.findall` and `re.finditer` both scan a string left to right for every **non-overlapping** occurrence of a pattern, advancing past each match before looking for the next. Given the same pattern and string they agree on *which* matches exist. What differs is the object each one gives you. `re.finditer(pattern, string)` returns an **iterator of `re.Match` objects**, produced lazily as you consume it. Each Match carries the matched text and its position: `Match.group()` for the text, `Match.start()`, `Match.end()` and `Match.span()` for the offsets. `re.findall(pattern, string)` returns a **list of strings** — the whole list, built before it returns. The positions are gone, and the Match objects were never handed to you. ## findall's shape-shifting return value This is the detail interviewers are usually fishing for. What `findall` puts in the list depends on how many capture groups the pattern contains: - **no groups** → a list of the full matched strings; - **exactly one group** → a list of what that group captured, *not* the full match; - **two or more groups** → a list of tuples, one element per group, still with the full match absent. ```python import re text = "clip_720p.mp4 clip_1080p.mp4" re.findall(r"\d+p", text) # ['720p', '1080p'] re.findall(r"(\d+)p", text) # ['720', '1080'] - group only re.findall(r"(\d+)(p)", text) # [('720', 'p'), ('1080', 'p')] ``` Nothing about the pattern's *meaning* changed between the first two lines — parentheses were added for grouping — yet the result type changed. This is a real production hazard: a colleague wraps part of a pattern in parentheses to apply an alternation, and downstream code that expected whole matches starts receiving fragments, with no exception anywhere. `re.finditer` has no such behaviour, because a Match object always exposes both the full match and the groups. ## When each one is right **Reach for `findall`** when you want a quick list of the matched text, the input is small, and positions are irrelevant — counting occurrences, pulling every token out of a short line, feeding a `set`. **Reach for `finditer`** when any of these hold: 1. **You need offsets.** Highlighting, slicing around a match, or reporting "line 12, column 34" all need `Match.span()`, and `findall` has thrown it away. 2. **You may stop early.** `next(re.finditer(...), None)` finds the first match and abandons the scan; `findall` always scans to the end of the string. For a first-hit search `re.search` is clearer still. 3. **The input is large.** `findall` holds every matched substring simultaneously. Consider a metadata extractor sweeping a 340-case regression pack of caption tracks: `findall` over each file builds a full list per file, while `finditer` lets you process and discard one match at a time. The scan cost is the same; the peak memory is not. 4. **You want to enrich each hit.** With a Match in hand you can carry the offsets forward, keep the source object beside it, or build a small record per occurrence. ## Gotchas worth naming **Non-overlapping.** Neither function reports overlapping occurrences — after a match, scanning resumes at its end. Overlapping hits require a lookahead construct or manual iteration with a moving start offset. **Zero-length matches.** A pattern that can match nothing (a `*` quantifier over an optional class, for example) yields zero-length matches, and both functions will produce one at each position rather than looping forever — but the resulting list of empty strings is rarely what the author wanted. If your `findall` result is full of `''`, the pattern can match emptiness. **`finditer` is an iterator, not a sequence.** It has no `len()`, cannot be indexed, and is exhausted after one pass. `list(re.finditer(...))` materialises it, at which point you have chosen `findall`'s memory profile deliberately rather than by accident — which is fine, and is exactly how you get a list of Match objects, something `findall` cannot give you at all. **It is a *lazy* iterator over a fixed string.** The string is not re-read; mutating the source name afterwards does not affect an in-flight iteration, because the Match objects reference the original string object. A reasonable house rule: default to `finditer`, and use `findall` only when you genuinely want a plain list of text and have looked at how many groups the pattern has.

  • Your re.findall result is a list of tuples instead of strings. What changed?
    The pattern gained a second capture group. With no groups findall returns whole matches, with one group it returns that group's text, and with two or more it returns a tuple per match. The usual fix is to make the extra parentheses non-capturing, or to switch to `re.finditer`, whose Match objects expose the full match regardless of grouping.
  • How do you count matches without building the list of them?
    Feed `re.finditer` to a counting construct — `sum(1 for _ in re.finditer(pattern, text))` — which consumes the iterator one Match at a time and holds nothing. `len(re.findall(pattern, text))` gives the same number but materialises every matched substring first, which is wasted work on large input.
  • Do these functions report overlapping matches?
    No. Both find non-overlapping occurrences: after each match the scan resumes at that match's end offset. To find overlapping occurrences you use a zero-width lookahead so the engine consumes nothing, or you iterate manually with a compiled pattern's search method and a start offset advanced by one each time.

saying these in an interview costs you the question

  • Thinking re.findall returns Match objects
  • Not knowing capture groups change findall's return shape
  • Calling len() on the re.finditer result
  • Believing finditer finds overlapping matches
  • Using re.findall then re-searching just to get positions
  • Assuming findall is lazy because it scans once

context