skip to content

How do `re.MULTILINE` and `re.DOTALL` change what `^`, `$` and `.` match?

level: middleimportance: should knowfreq 55%

answer

  1. Two flags, two completely separate effects
  2. One is about lines, one about newlines
  3. `$` already forgives one trailing newline
  4. `\A` and `\Z` ignore the line flag
  5. Turning one on multiplies the match count

basics

~20 s

re.MULTILINE makes ^ and $ match at every line boundary instead of only at the start and end of the whole string. re.DOTALL makes . match a newline as well. The two flags are independent.

solid answer

~40 s

By default `^` matches only at the start of the subject string, and `$` matches at the end — plus, importantly, just before a single trailing newline, so `re.search(r"\w+$", "abc\n")` still matches `abc`. `re.MULTILINE` (`re.M`) makes both anchors fire at every line boundary, so `re.findall(r"^\w+", text, re.MULTILINE)` yields one match per line rather than one for the whole string. `re.DOTALL` (`re.S`) changes only `.`, which otherwise matches every character except `\n`. The two are orthogonal: MULTILINE never touches `.`, DOTALL never touches the anchors. `\A` and `\Z` always mean absolute start and absolute end of the string and ignore `re.MULTILINE` entirely — `\Z` also refuses the trailing newline that `$` forgives — which is what you want for whole-input validation, alongside `re.fullmatch`.

code

pycon · 10 lines
pycon
>>> import re
>>> text = "one\ntwo\nthree"
>>> re.findall(r"^\w+", text)
['one']
>>> re.findall(r"^\w+", text, re.MULTILINE)
['one', 'two', 'three']
>>> re.search(r"\w+$", "abc\n").group()
'abc'
>>> print(re.search(r"abc\Z", "abc\n"))
None

go deeper

for a junior

Remember the one-line summary for each flag: MULTILINE is about ^ and $ per line, DOTALL is about . crossing newlines. Being able to name which flag you need for a given failure is what is being checked here.

for a middle

Explain the defaults precisely, including that $ also matches just before a trailing newline, and that the two flags are orthogonal. Be able to show re.findall returning one result without MULTILINE and one per line with it.

for a senior

Show that you treat MULTILINE as a change in result cardinality, not just in reach: name the callers you would re-check, and reach for re.fullmatch or \A/\Z when the pattern is validating an entire untrusted value rather than searching within it.

for a principal

Own the convention: whether line-oriented work in your codebase goes through a flag or through an explicit splitlines() loop, and where regular-expression validation is allowed to be the gate at all. Hidden semantics in a flag are a maintenance cost you are choosing to pay.

### The default, stated precisely With no flags, `^` matches at exactly one position: offset zero of the subject string. `$` matches at the end of the string **and** at the position just before a final newline. That second case is not a bug; it is deliberate, so that a pattern written for a line still works when the line arrives with its terminator attached. It is also the single most common surprise in this area: ```python import re re.search(r"\w+$", "abc\n").group() # 'abc' — matched, despite the newline re.search(r"abc\Z", "abc\n") # None — \Z means the true end ``` `\A` is the absolute start and `\Z` the absolute end. Unlike `^` and `$` they are unaffected by flags, and `\Z` does not forgive a trailing newline. Python has no `\z`; `\Z` already means what other flavours spell `\z`. ### re.MULTILINE `re.MULTILINE` (short alias `re.M`, inline `(?m)`) redefines the two anchors in terms of lines: `^` now matches at the start of the string and immediately after every `\n`; `$` matches at the end of the string and immediately before every `\n`. ```python text = "one\ntwo\nthree" re.findall(r"^\w+", text) # ['one'] re.findall(r"^\w+", text, re.MULTILINE) # ['one', 'two', 'three'] ``` The flag says nothing about `.`. A pattern like `^a.*b$` under MULTILINE is anchored per line *and* still cannot cross a newline, because `.` remains newline-hostile — which is usually exactly the line-oriented behaviour you wanted. ### re.DOTALL `re.DOTALL` (alias `re.S`, inline `(?s)`) does one thing: it lets `.` match `\n` too. It does not touch anchors, and it does not touch character classes — a negated class such as `[^a]` already matches a newline with or without the flag, which is why `[\s\S]` is a common flag-free way to say "any character at all". ```python re.search(r"a.b", "a\nb") # None re.search(r"a.b", "a\nb", re.DOTALL) # matches 'a\nb' ``` Combine flags with the bitwise OR: `re.MULTILINE | re.DOTALL`, or write both inline at the very start of the pattern as `(?ms)`. ### The failure this causes in production A fraud-scoring service matches rules against a free-text memo field that customers can paste multiple lines into. A rule was written as `^id=(\w+)$` and, because it never matched anything on multi-line memos, someone added `re.MULTILINE`. It matched — and now it matched *per line*. The handler had been written around `re.findall`, iterating results and enqueuing a manual-review task for each; a memo containing the same identifier on two lines produced a duplicated side effect, two review cases for one payment. At a 1,200-request-per-minute peak that turned into a visible backlog before anyone connected it to a one-word flag change. The lesson is that MULTILINE does not merely make a pattern "work on multi-line input": it changes the **cardinality** of the result. Whenever you add it, re-check every caller that assumed at most one match, and decide deliberately between three different intents: * validate the entire input → `re.fullmatch`, or anchor with `\A` and `\Z`, and do not use MULTILINE at all; * find one occurrence anywhere → `re.search` with no anchors; * process line by line → either MULTILINE with a loop that expects many matches, or, often clearer, split the text into lines yourself and match each one. That third option is worth pausing on. `for line in text.splitlines():` with an unflagged pattern is frequently more readable than a MULTILINE regular expression, because the iteration is visible in the code instead of hidden in a flag, and each match is unambiguously one line. ### Security-adjacent consequence Anchoring a validation pattern with `^...$` rather than `\A...\Z` means an attacker who can inject a newline may get a value accepted whose *first* line looks benign. `re.fullmatch(r"[A-Z]{2}-\d{4,6}", value)` is the blunt, safe formulation: it requires the whole string to match, no anchors needed, and no flag can loosen it by accident. ### What to say in an interview State the defaults first, including the trailing-newline concession that `$` makes — that detail is the tell that you have actually read the semantics. Then give one line each for the two flags, stress that they are orthogonal, and close on `\A` / `\Z` / `re.fullmatch` as the correct tools for validating a whole input. ### Picking the flag from the symptom Three symptoms map cleanly onto three fixes. A pattern that matches only the first line of a multi-line subject wants `re.MULTILINE`, or an explicit `splitlines()` loop. A pattern that stops at a newline in the middle of the region you meant to capture wants `re.DOTALL`, or a negated class such as `[\s\S]` if you would rather not change the meaning of `.` across the whole pattern. And a validation pattern that accepts a value it should have rejected usually wants neither flag — it wants `re.fullmatch`, because the real defect was `^...$` anchoring on input that could contain a newline. Naming which of the three you are looking at, before reaching for a flag, is most of the work.

  • Why does `re.search(r"\w+$", "abc\n")` match even with no flags set?
    Because `$` matches both at the end of the string and immediately before a single trailing newline, so a pattern written for a line still works on a line read with its terminator. If you need to forbid that newline, anchor with `\Z`, which means the true end of the string, or use `re.fullmatch`.
  • Does `re.DOTALL` change how a negated character class such as `[^a]` behaves?
    No. `re.DOTALL` affects only the `.` metacharacter. A negated class already matches a newline unless you exclude it explicitly, which is why `[\s\S]` or `[^\x00]`-style constructions are the flag-free way to say ‘any character including a newline’.
  • You need a pattern to validate that an entire user-supplied value is a product code. Which anchoring do you choose?
    `re.fullmatch(pattern, value)` — it requires the whole string to match with no anchors at all. If you must anchor inside the pattern, use `\A` and `\Z`, never `^` and `$`: an injected newline can make a `^...$` pattern accept a value whose first line alone looks valid, and `re.MULTILINE` makes that worse.

saying these in an interview costs you the question

  • Thinks `re.MULTILINE` is what makes `.` match newlines
  • Believes `$` matches only at the very end of the string
  • Treats `\Z` and `$` as interchangeable under `re.MULTILINE`
  • Assumes a negated class like `[^a]` stops at a newline
  • Adds `re.MULTILINE` without rechecking how many matches callers expect

context