skip to content

Quantifiers, Anchors, and Flags

Greedy quantifiers grab as much as possible and lazy ones as little — one HTML tag against the whole document. Flags like IGNORECASE, MULTILINE and DOTALL change what a pattern matches globally.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

What is the difference between `.*` and `.*?` in a Python `re` pattern?

level: juniorimportance: must knowfreq 70%

answer

  1. Two appetites for the same repetition
  2. One takes everything, one takes nothing
  3. A trailing `?` flips the preference
  4. `<.*>` runs to the last `>`
  5. `[^>]*` removes the choice entirely

basics

~20 s

.* is greedy: it grabs as much text as it can, then gives characters back only when the rest of the pattern fails. .*? is lazy: it starts empty and grows one character at a time, only when forced.

solid answer

~50 s

Every repetition operator in Python's `re` has a greedy and a lazy form: `*` / `*?`, `+` / `+?`, `?` / `??`, `{m,n}` / `{m,n}?`. The greedy form consumes the longest run it can and then backtracks, surrendering one character at a time until the remainder of the pattern matches; the lazy form consumes the fewest and expands only when the remainder fails. The classic demonstration is `<.*>` against `<a> text <b>`: it matches the whole string, because `.` happily matches `>` as well. `<.*?>` matches just `<a>`. Laziness changes the **length** of the match, never where it starts — the engine still takes the leftmost starting position it can. In practice a negated character class such as `<[^>]*>` beats both: it cannot cross the delimiter at all, so it stops for structural reasons rather than by backtracking.

code

pycon · 7 lines
pycon
>>> import re
>>> re.search(r"<.*>", "<a> text <b>").group()
'<a> text <b>'
>>> re.search(r"<.*?>", "<a> text <b>").group()
'<a>'
>>> re.search(r"<[^>]*>", "<a> text <b>").group()
'<a>'

go deeper

for a junior

Be ready to state the rule in one sentence and demonstrate it on <.*> versus <.*?> against a string with two tags. Knowing that *?, +? and {m,n}? are all lazy forms is enough at this level.

for a middle

Explain the mechanics: greedy takes the maximum then backtracks one character at a time, lazy takes the minimum then grows. Say explicitly that both matches begin at the same leftmost position, and show the negated character class as the better tool.

for a senior

An interviewer expects you to reach for [^delim]* rather than either quantifier in code review, and to notice that a greedy .* silently changes reach the moment someone enables the flag that lets . cross newlines. Say what you would commit and why.

for a principal

Own the standard: where in your codebase regular expressions are acceptable at all, versus where a real parser or a split-and-strip belongs. Greedy-versus-lazy arguments in review are usually a signal the pattern is doing structural work it should not be doing.

### What a quantifier actually decides A quantifier in a Python regular expression says how many times the preceding element may repeat. `*` means zero or more, `+` means one or more, `?` means zero or one, and `{m,n}` means between `m` and `n` times inclusive (`{m,}` is at least `m`, `{,n}` is at most `n`, `{m}` is exactly `m`). By default every one of these is **greedy**. Appending a `?` to the quantifier itself — `*?`, `+?`, `??`, `{m,n}?` — makes it **lazy** (the docs call it non-greedy). Note the pun that trips people up: the `?` in `a?` is a quantifier meaning "optional", while the `?` in `a*?` is a modifier on the `*` that already sits there. ### How the engine resolves the choice CPython's `re` is a backtracking engine. It scans for the leftmost position at which the whole pattern can succeed, and at that position it walks the pattern left to right. When it reaches a greedy quantifier it takes as many repetitions as it possibly can — for `.*` that is every remaining character on the line — and only then tries to match what follows. If that fails, it *backtracks*: it hands one character back and retries the rest of the pattern, over and over, until either something matches or it runs out of characters to return. A lazy quantifier inverts the order of attempts: it first tries zero repetitions, tries the rest of the pattern, and only if that fails does it consume one more character and try again. So greedy and lazy are not two different answers to "what is the match" — they are two different **orders in which candidate matches are tried**. Both start at the same leftmost position; they differ in which candidate they reach first. ```python import re text = "<a> text <b>" re.search(r"<.*>", text).group() # '<a> text <b>' re.search(r"<.*?>", text).group() # '<a>' ``` The greedy version does not stop at the first `>` because nothing in `<.*>` forbids `.` from matching `>`. It runs to the end of the string, backtracks to the *last* `>`, and succeeds there. The lazy version stops at the first `>` it can. ### The leftmost rule beats laziness A very common misconception is that `.*?` yields "the shortest match in the string". It does not. The engine's first commitment is to the earliest possible starting position; laziness only chooses among the candidates that start there. If a shorter match exists further right, greedy and lazy alike will ignore it, because the leftmost start already succeeded. ### Bounded repeats `{m,n}` participates in the same rule: `\d{2,4}` prefers four digits and backs down to three, then two; `\d{2,4}?` prefers two and grows. `re.findall(r"\d{2,4}", "1 12 123 12345")` returns `['12', '123', '1234']` — note the last one is truncated at four digits, which is exactly the upper bound doing its job. Write `{m,n}` with no spaces; `{2, 4}` is not a quantifier at all, it is four literal characters. ### The alternative that is usually right When you are extracting a field delimited by some character, neither greedy nor lazy is the best tool — a negated character class is. `<[^>]*>` cannot cross a `>` under any circumstances, so it stops for a structural reason rather than by trial and error. It expresses the intent ("everything up to but not including the delimiter") directly, it is immune to the reader misreading the appetite of the quantifier, and it does far less work at match time because each character is decided once. The same shape applies to quoted strings (`"[^"]*"`) and to key=value fragments (`=([^;]*)`). ### Where the dot itself gets in the way One detail that changes the answer entirely: `.` does not match a newline unless the `re.DOTALL` flag is set. On multi-line input a greedy `.*` is therefore bounded by the end of its line, which is why the same pattern behaves very differently once someone adds that flag. If your subject text can contain newlines, decide deliberately whether `.` should cross them. ### What to say in an interview Define greedy and lazy as *orders of preference under backtracking*, show the `<.*>` versus `<.*?>` contrast, state clearly that both matches start at the same place, and finish by pointing at the negated character class as the version you would actually commit. That sequence — mechanism, demonstration, and the better alternative — is what separates someone who has memorised a rule from someone who writes patterns other people can maintain.

  • Does `.*?` guarantee you get the shortest match anywhere in the string?
    No. The engine commits to the leftmost starting position first, and laziness only picks the shortest candidate that begins *there*. A genuinely shorter match further right is never considered, because the leftmost start already succeeded. If you need the shortest overall, iterate the candidates yourself and choose.
  • How would you extract the text between two delimiters without relying on greedy or lazy behaviour at all?
    Use a negated character class: `<[^>]*>` or `"([^"]*)"`. The class physically cannot match the delimiter, so the repetition stops for a structural reason instead of by backtracking. It reads more clearly, behaves identically whether or not the reader knows the greedy rules, and does less work per character.
  • Is `{2,4}?` legal in Python's `re`, and what does it mean?
    Yes. It is a lazy bounded repeat: it still requires between two and four repetitions, but it tries two first and only grows toward four if the rest of the pattern fails. Contrast `{2,4}`, which tries four first and backs down. Beware `{2, 4}` with a space — that is not a quantifier, it is four literal characters.

Greedy is filling your plate at the buffet and putting food back until you can fit dessert; lazy is taking one spoonful at a time until someone says that is enough.

saying these in an interview costs you the question

  • Says `.*?` finds the shortest match anywhere in the string
  • Thinks `.` matches a newline by default
  • Believes greedy and lazy matches can start at different positions
  • Uses `<.*>` to pull one tag out and calls it correct
  • Reads the `?` in `*?` as meaning the repetition is optional

context

open as a page

How do `re.MULTILINE` and `re.DOTALL` change what `^`, `$` and `.` match?

level: middleimportance: should knowfreq 55%

basics

~20 s

re.MULTILINE makes ^ and $ match at every line boundary instead of only at the start and end of the whole string. re.DOTALL makes . match a newline as well. The two flags are independent.

open as a page

What does `\b` match in a Python `re` pattern, and what does `re.ASCII` change about it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

\b is a zero-width assertion: it matches at a position where exactly one side is a word character. ‘Word character’ means \w, which for str patterns is Unicode-aware by default; re.ASCII narrows it to [A-Za-z0-9_].

open as a page

Where may an inline `(?i)` flag appear in a Python `re` pattern, and what does `(?i:...)` do?

level: middleimportance: nice to knowfreq 20%

basics

~20 s

A bare (?i) is a global flag: it applies to the whole pattern, and since Python 3.11 it must sit at the very start or compilation raises re.error. (?i:...) is the scoped form and applies only inside its own group.

open as a page