How do `re.MULTILINE` and `re.DOTALL` change what `^`, `$` and `.` match?
answer
- Two flags, two completely separate effects
- One is about lines, one about newlines
- `$` already forgives one trailing newline
- `\A` and `\Z` ignore the line flag
- Turning one on multiplies the match count
basics
~20 sre.MULTILINE makes ^ and $ match at every line boundary instead of only at the start and end of the whole string. re.DOTALL makes . match a newline as well. The two flags are independent.
solid answer
~40 sBy default `^` matches only at the start of the subject string, and `$` matches at the end — plus, importantly, just before a single trailing newline, so `re.search(r"\w+$", "abc\n")` still matches `abc`. `re.MULTILINE` (`re.M`) makes both anchors fire at every line boundary, so `re.findall(r"^\w+", text, re.MULTILINE)` yields one match per line rather than one for the whole string. `re.DOTALL` (`re.S`) changes only `.`, which otherwise matches every character except `\n`. The two are orthogonal: MULTILINE never touches `.`, DOTALL never touches the anchors. `\A` and `\Z` always mean absolute start and absolute end of the string and ignore `re.MULTILINE` entirely — `\Z` also refuses the trailing newline that `$` forgives — which is what you want for whole-input validation, alongside `re.fullmatch`.
code
pycon · 10 lines>>> import re
>>> text = "one\ntwo\nthree"
>>> re.findall(r"^\w+", text)
['one']
>>> re.findall(r"^\w+", text, re.MULTILINE)
['one', 'two', 'three']
>>> re.search(r"\w+$", "abc\n").group()
'abc'
>>> print(re.search(r"abc\Z", "abc\n"))
Nonego deeper
Remember the one-line summary for each flag: MULTILINE is about ^ and $ per line, DOTALL is about . crossing newlines. Being able to name which flag you need for a given failure is what is being checked here.
Explain the defaults precisely, including that $ also matches just before a trailing newline, and that the two flags are orthogonal. Be able to show re.findall returning one result without MULTILINE and one per line with it.
Show that you treat MULTILINE as a change in result cardinality, not just in reach: name the callers you would re-check, and reach for re.fullmatch or \A/\Z when the pattern is validating an entire untrusted value rather than searching within it.
Own the convention: whether line-oriented work in your codebase goes through a flag or through an explicit splitlines() loop, and where regular-expression validation is allowed to be the gate at all. Hidden semantics in a flag are a maintenance cost you are choosing to pay.
### The default, stated precisely With no flags, `^` matches at exactly one position: offset zero of the subject string. `$` matches at the end of the string **and** at the position just before a final newline. That second case is not a bug; it is deliberate, so that a pattern written for a line still works when the line arrives with its terminator attached. It is also the single most common surprise in this area: ```python import re re.search(r"\w+$", "abc\n").group() # 'abc' — matched, despite the newline re.search(r"abc\Z", "abc\n") # None — \Z means the true end ``` `\A` is the absolute start and `\Z` the absolute end. Unlike `^` and `$` they are unaffected by flags, and `\Z` does not forgive a trailing newline. Python has no `\z`; `\Z` already means what other flavours spell `\z`. ### re.MULTILINE `re.MULTILINE` (short alias `re.M`, inline `(?m)`) redefines the two anchors in terms of lines: `^` now matches at the start of the string and immediately after every `\n`; `$` matches at the end of the string and immediately before every `\n`. ```python text = "one\ntwo\nthree" re.findall(r"^\w+", text) # ['one'] re.findall(r"^\w+", text, re.MULTILINE) # ['one', 'two', 'three'] ``` The flag says nothing about `.`. A pattern like `^a.*b$` under MULTILINE is anchored per line *and* still cannot cross a newline, because `.` remains newline-hostile — which is usually exactly the line-oriented behaviour you wanted. ### re.DOTALL `re.DOTALL` (alias `re.S`, inline `(?s)`) does one thing: it lets `.` match `\n` too. It does not touch anchors, and it does not touch character classes — a negated class such as `[^a]` already matches a newline with or without the flag, which is why `[\s\S]` is a common flag-free way to say "any character at all". ```python re.search(r"a.b", "a\nb") # None re.search(r"a.b", "a\nb", re.DOTALL) # matches 'a\nb' ``` Combine flags with the bitwise OR: `re.MULTILINE | re.DOTALL`, or write both inline at the very start of the pattern as `(?ms)`. ### The failure this causes in production A fraud-scoring service matches rules against a free-text memo field that customers can paste multiple lines into. A rule was written as `^id=(\w+)$` and, because it never matched anything on multi-line memos, someone added `re.MULTILINE`. It matched — and now it matched *per line*. The handler had been written around `re.findall`, iterating results and enqueuing a manual-review task for each; a memo containing the same identifier on two lines produced a duplicated side effect, two review cases for one payment. At a 1,200-request-per-minute peak that turned into a visible backlog before anyone connected it to a one-word flag change. The lesson is that MULTILINE does not merely make a pattern "work on multi-line input": it changes the **cardinality** of the result. Whenever you add it, re-check every caller that assumed at most one match, and decide deliberately between three different intents: * validate the entire input → `re.fullmatch`, or anchor with `\A` and `\Z`, and do not use MULTILINE at all; * find one occurrence anywhere → `re.search` with no anchors; * process line by line → either MULTILINE with a loop that expects many matches, or, often clearer, split the text into lines yourself and match each one. That third option is worth pausing on. `for line in text.splitlines():` with an unflagged pattern is frequently more readable than a MULTILINE regular expression, because the iteration is visible in the code instead of hidden in a flag, and each match is unambiguously one line. ### Security-adjacent consequence Anchoring a validation pattern with `^...$` rather than `\A...\Z` means an attacker who can inject a newline may get a value accepted whose *first* line looks benign. `re.fullmatch(r"[A-Z]{2}-\d{4,6}", value)` is the blunt, safe formulation: it requires the whole string to match, no anchors needed, and no flag can loosen it by accident. ### What to say in an interview State the defaults first, including the trailing-newline concession that `$` makes — that detail is the tell that you have actually read the semantics. Then give one line each for the two flags, stress that they are orthogonal, and close on `\A` / `\Z` / `re.fullmatch` as the correct tools for validating a whole input. ### Picking the flag from the symptom Three symptoms map cleanly onto three fixes. A pattern that matches only the first line of a multi-line subject wants `re.MULTILINE`, or an explicit `splitlines()` loop. A pattern that stops at a newline in the middle of the region you meant to capture wants `re.DOTALL`, or a negated class such as `[\s\S]` if you would rather not change the meaning of `.` across the whole pattern. And a validation pattern that accepts a value it should have rejected usually wants neither flag — it wants `re.fullmatch`, because the real defect was `^...$` anchoring on input that could contain a newline. Naming which of the three you are looking at, before reaching for a flag, is most of the work.
- Why does `re.search(r"\w+$", "abc\n")` match even with no flags set?Because `$` matches both at the end of the string and immediately before a single trailing newline, so a pattern written for a line still works on a line read with its terminator. If you need to forbid that newline, anchor with `\Z`, which means the true end of the string, or use `re.fullmatch`.
- Does `re.DOTALL` change how a negated character class such as `[^a]` behaves?No. `re.DOTALL` affects only the `.` metacharacter. A negated class already matches a newline unless you exclude it explicitly, which is why `[\s\S]` or `[^\x00]`-style constructions are the flag-free way to say ‘any character including a newline’.
- You need a pattern to validate that an entire user-supplied value is a product code. Which anchoring do you choose?`re.fullmatch(pattern, value)` — it requires the whole string to match with no anchors at all. If you must anchor inside the pattern, use `\A` and `\Z`, never `^` and `$`: an injected newline can make a `^...$` pattern accept a value whose first line alone looks valid, and `re.MULTILINE` makes that worse.
saying these in an interview costs you the question
- Thinks `re.MULTILINE` is what makes `.` match newlines
- Believes `$` matches only at the very end of the string
- Treats `\Z` and `$` as interchangeable under `re.MULTILINE`
- Assumes a negated class like `[^a]` stops at a newline
- Adds `re.MULTILINE` without rechecking how many matches callers expect