skip to content

Text Processing and Regular Expressions

Pattern work with the re module: the matching entry points, capture groups and substitution, quantifiers and flags, plus normalizing Unicode before you compare. Easy to use and easy to misuse.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

20

How do you name a capture group in a Python re pattern, and what does .groupdict() return?

level: juniorimportance: must knowfreq 62%

answer

  1. Parentheses record what they matched
  2. Numbering follows the opening parenthesis
  3. One accessor gives a tuple, another a dict
  4. Python spells it with a P
  5. Absent groups are not empty strings

basics

~10 s

Write (?P<name>...) to name a group. The match object then exposes it as m.group('name'), and m.groupdict() returns a dict mapping every named group to the text it captured; unnamed numbered groups are not included.

solid answer

~40 s

Capturing groups are numbered by their **opening** parenthesis, left to right, from 1; group 0 is the whole match. Python's own spelling `(?P<name>...)` adds a name to a group without taking its number away, so `m.group(2)` and `m.group("start")` can be the same group. `m.groups()` returns a tuple of groups 1..N — group 0 is excluded — while `m.groupdict()` returns only the *named* ones as a dict. A group that did not participate in the match yields `None`, not an empty string, unless you pass a default: `m.groups("")`. Names must be valid identifiers and unique in the pattern; a compiled pattern exposes the mapping as `.groupindex`. Prefer names in any pattern you expect to edit, because inserting a group in the middle silently renumbers everything after it.

code

pycon · 8 lines
pycon
>>> import re
>>> m = re.search(r"(?P<chrom>chr\d+):(?P<start>\d+)-(?P<end>\d+)", "chr7:117120017-117308719")
>>> m.group("chrom"), m.group(2)
('chr7', '117120017')
>>> m.groups()
('chr7', '117120017', '117308719')
>>> m.groupdict()
{'chrom': 'chr7', 'start': '117120017', 'end': '117308719'}

go deeper

for a junior

Recall the three accessors and what each returns: group(0) is the whole match, groups() is a tuple of captures from 1, groupdict() is a dict of named groups. Be able to write (?P<name>...) from memory.

for a middle

Explain the numbering rule by opening parenthesis, why a non-participating group is None rather than empty, and how the groups()/groupdict() default argument changes that. Mention (?:...) when grouping without capturing.

for a senior

Show the maintenance judgement: numeric indexes shift when someone inserts a group, so named groups plus groupdict() are what survive a pattern edit in code that parses real files. Know the compile-time errors around duplicate names.

for a principal

Own the boundary question — when a regex with a dozen named groups has stopped being the right parser at all, and a real grammar, a line splitter, or a purpose-built format reader should take over. Name the readability and testing cost you are trading.

### What creates a group Every pair of parentheses in a Python regular expression that is not one of the `(?...)` extension forms creates a **capturing group** — a slot in which the engine records the substring that part of the pattern consumed. Groups are numbered by the position of their *opening* parenthesis, scanning left to right, starting at 1. Group 0 is reserved for the entire match. Nesting does not change the rule: in `((a)(b))` the outer group is 1, `(a)` is 2 and `(b)` is 3. Counting parentheses by hand is exactly what goes wrong on any pattern longer than one line, which is the practical argument for naming. ### Naming: `(?P<name>...)` Python spells a named group `(?P<name>...)`. The `P` is Python's own marker on the shared `(?...)` extension syntax — the .NET/PCRE spelling `(?<name>...)` is a syntax error here, and getting that wrong is one of the most common small mistakes in an interview. The name must be a valid Python identifier and must be unique within the pattern; reusing one raises `re.error` at compile time ("redefinition of group name"). A named group is *still* a numbered group: naming is additive, so `\1` and `m.group(1)` keep working alongside `m.group("name")`. ### The accessors Given a match object `m`: * `m.group()` and `m.group(0)` return the whole matched substring. * `m.group(n)` returns group `n`; `m.group("name")` returns a named group. Passing several indexes, `m.group(1, 3)`, returns a tuple. * `m.groups()` returns a tuple of **all capturing groups, 1 through N** — group 0 is not in it, and neither are non-capturing `(?:...)` groups, which record nothing. * `m.groupdict()` returns `{name: text}` for the **named groups only**. A pattern with no named groups gives an empty dict. * `m.lastindex` and `m.lastgroup` name the last group that matched, which is occasionally useful for alternations. * On the compiled pattern, `.groupindex` is a mapping of name to group number. ```python import re m = re.search(r"(?P<chrom>chr\d+):(?P<start>\d+)-(?P<end>\d+)", "chr7:117120017-117308719") m.group("chrom") # 'chr7' m.group(2) # '117120017' — same group as 'start' m.groups() # ('chr7', '117120017', '117308719') m.groupdict() # {'chrom': 'chr7', 'start': ..., 'end': ...} ``` ### Groups that did not participate An optional group, or a group inside an alternation branch the engine did not take, captured nothing. Python reports that as `None`, not `""`. This is the single most common source of `TypeError: int() argument must be...` in parsing code: `int(m.group("version"))` blows up on the record that omitted the suffix. Both `groups()` and `groupdict()` accept a default to substitute instead — `m.groups("")`, `m.groupdict("")` — and a bare `m.group("version")` still returns `None`, so a guard or a default is your choice per call site. ```python m = re.fullmatch(r"(\w+?)(?:\.(\d+))?", "BRCA1") m.groups() # ('BRCA1', None) m.groups("-") # ('BRCA1', '-') ``` Note the `(?:...)` around the optional suffix in that pattern: it groups `\.(\d+)` so the `?` applies to the whole thing, without adding a capture slot of its own. ### Referring to a group from inside the pattern Groups can be referenced later in the same pattern — a **backreference** — numerically as `\1` or by name as `(?P=name)`. That is how you assert that two parts of the text are identical, such as a repeated word or a closing tag matching its opener: ```python re.findall(r"<(?P<t>\w+)>.*?</(?P=t)>", "<a>1</a><b>2</b>") # ['a', 'b'] ``` ### Why names are worth the extra characters Three practical reasons. First, edits: add one group near the front of a pattern and every numeric index after it shifts, silently breaking `m.group(3)` and any `\2` downstream — names are immune. Second, readability at the call site: `m.group("start")` says what it means where `m.group(2)` needs the reader to go count. Third, `m.groupdict()` drops straight into a record — in a text-extraction pipeline over annotation lines, `groupdict()` is often the row you were going to build by hand anyway. The cost is that a named group is a capturing group, so where you only need grouping for a quantifier or an alternation, use `(?:...)` and keep the numbering clean. ### In an interview Say the three things out loud: numbering starts at 1 and 0 is the whole match; `groups()` is a tuple of all captures while `groupdict()` is a dict of named ones; non-participating groups are `None`. That is the whole answer, and offering the `None` detail unprompted is what makes it sound like you have parsed real text.

  • How do you refer back to an earlier group from inside the same regex pattern?
    With a backreference: `\1` by number, or `(?P=name)` by name. It asserts that the text here is identical to what that group already captured — `r"\b(\w+) \1\b"` finds a doubled word, and `r"<(?P<t>\w+)>.*?</(?P=t)>"` requires the closing tag to repeat the opening one. Note the numeric form `\1` means a backreference inside a pattern, but a *group reference* inside a re.sub replacement string; same spelling, two different grammars.
  • Can two groups in one Python pattern share a name?
    No. `re.compile` raises `re.error` with "redefinition of group name" at compile time, unlike some other engines that allow duplicate names across alternation branches. If two alternatives need to yield the same field, either give them distinct names and coalesce in Python, or restructure so the shared part sits in one group outside the alternation.
  • What is the difference between m.groups() and m.groupdict() on a pattern that mixes named and unnamed groups?
    `groups()` returns every capturing group in numeric order, named or not, as a tuple starting at group 1. `groupdict()` returns only the named ones, keyed by name; the unnamed captures are simply absent. So a pattern with three groups of which one is named gives a 3-tuple and a 1-entry dict.

saying these in an interview costs you the question

  • Says the first capture group is number 0
  • Claims .groups() includes the whole match
  • Expects .groupdict() to contain unnamed numbered groups
  • Assumes a non-participating group captures an empty string
  • Writes the .NET spelling (?<name>...) in a Python pattern
  • Thinks naming a group removes its number

context

open as a page

Why are regex patterns in Python written as raw strings like r"\d+"?

level: juniorimportance: must knowfreq 60%

basics

~10 s

A backslash means something to both the string-literal parser and the regex engine. A raw literal stops Python consuming it first, so the engine sees what you typed. Otherwise you must double every backslash.

open as a page

What is the difference between re.match, re.search and re.fullmatch?

level: juniorimportance: must knowfreq 75%

basics

~20 s

re.match only tries the pattern at the start of the string, re.search scans forward and returns the first match anywhere, and re.fullmatch requires the pattern to consume the whole string. All three return a Match object or None.

open as a page

What is the difference between `.*` and `.*?` in a Python `re` pattern?

level: juniorimportance: must knowfreq 70%

basics

~20 s

.* is greedy: it grabs as much text as it can, then gives characters back only when the rest of the pattern fails. .*? is lazy: it starts empty and grows one character at a time, only when forced.

open as a page

Why does `re.search(r'(a+)+$', 'a'*30 + '!')` take exponential time?

level: middleimportance: must knowfreq 50%

basics

~20 s

The inner a+ and the outer + can carve the run of as up in about 2^n ways. The trailing ! makes $ fail, so the backtracking engine must try every one of those splits before reporting no match.

open as a page

When should the repl argument to re.sub be a callable instead of a replacement string?

level: middleimportance: must knowfreq 52%

basics

~20 s

Pass a callable whenever the replacement depends on what matched — arithmetic, a lookup, a conditional skip. re.sub calls it once per match with the Match object and inserts whatever str it returns, with no template escape processing.

open as a page

How do you strip accents from a Python str to build a search or slug key?

level: middleimportance: must knowfreq 52%

basics

~20 s

Normalize to NFD first so each accent becomes its own combining code point, drop every character whose unicodedata.combining value is non-zero, then normalize back to NFC. Without the NFD step the accent is fused inside one precomposed code point.

open as a page

Which pattern shapes in Python's `re` module risk catastrophic backtracking?

level: juniorimportance: should knowfreq 30%

basics

~20 s

Ambiguous repetition: a quantifier nested inside a repeated group such as (a+)+, or a repeated alternation whose branches overlap such as (a|a)*. Both are cheap until the input almost matches, and then the engine tries every possible split.

open as a page

What does Python's str.casefold() do that str.lower() does not?

level: juniorimportance: should knowfreq 38%

basics

~10 s

str.lower() maps characters to their lowercase form for display. str.casefold() applies the stronger Unicode case-folding mappings meant for caseless matching, so "Stra\u00dfe".casefold() gives "strasse" while .lower() leaves the sharp s alone.

open as a page

Why does re.split return the separators when the split pattern contains a capture group?

level: middleimportance: should knowfreq 34%

basics

~20 s

re.split interleaves the text of every capturing group into the result list, by design, so you can reassemble the string. Make the delimiter group non-capturing with (?:...), or drop the parentheses, to get only the fields.

open as a page

How does re.findall differ from re.finditer, and when do you need each?

level: middleimportance: should knowfreq 55%

basics

~20 s

re.findall builds a list of strings (or tuples, once the pattern has capture groups) for every non-overlapping match. re.finditer returns a lazy iterator of Match objects, so you keep offsets and spend no memory on the full list.

open as a page

How do `re.MULTILINE` and `re.DOTALL` change what `^`, `$` and `.` match?

level: middleimportance: should knowfreq 55%

basics

~20 s

re.MULTILINE makes ^ and $ match at every line boundary instead of only at the start and end of the whole string. re.DOTALL makes . match a newline as well. The two flags are independent.

open as a page

How do NFC and NFD differ from NFKC and NFKD in unicodedata.normalize?

level: middleimportance: should knowfreq 40%

basics

~20 s

NFC and NFD are canonical: they only compose or decompose characters that are already equivalent, and they round-trip. NFKC and NFKD add compatibility folding, which rewrites formatting variants such as ligatures and Roman numerals and is lossy.

open as a page

A worker is stuck in `re.search()` on one email digest for 27 minutes and ignores Ctrl-C. How do you confirm the cause and bound it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Dump the live process stack with faulthandler to confirm it is inside the match. re has no timeout and the C matching loop defers signals, so the only hard bound is a separate process you can kill.

open as a page

Why is passing data-derived text as re.sub's replacement string unsafe, and what is the fix?

level: seniorimportance: should knowfreq 30%

basics

~20 s

The replacement string is a template that re.sub scans for group references and escapes, so a backslash in the data is interpreted rather than inserted. It corrupts output or raises re.error. Pass a callable returning the literal text instead.

open as a page

Is re.compile worth it given that the re module already caches compiled patterns?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The re module caches compiled patterns, so re.compile is no large speed win. It still helps: no per-call cache lookup, no eviction when a program uses many distinct patterns, and only its methods take pos and endpos.

open as a page

What does `\b` match in a Python `re` pattern, and what does `re.ASCII` change about it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

\b is a zero-width assertion: it matches at a position where exactly one side is a word character. ‘Word character’ means \w, which for str patterns is Unicode-aware by default; re.ASCII narrows it to [A-Za-z0-9_].

open as a page

A ticket-triage bot slices each Python str title to a fixed character budget and accents vanish - why, and how do you cut safely?

level: seniorimportance: should knowfreq 30%

basics

~20 s

len() and slicing count code points, not what a reader sees. If a title is decomposed, an accent is a separate combining code point after its base letter, so a cut can land between them and drop the accent. Normalize to NFC first and back off before a combining mark.

open as a page

What do atomic groups `(?>...)` and possessive quantifiers do in Python's `re`?

level: middleimportance: nice to knowfreq 20%

basics

~10 s

Both tell the engine to throw away the backtracking alternatives once the enclosed part has matched. Added in Python 3.11, they collapse an exponential search and can turn a would-be match into a non-match.

open as a page

Where may an inline `(?i)` flag appear in a Python `re` pattern, and what does `(?i:...)` do?

level: middleimportance: nice to knowfreq 20%

basics

~20 s

A bare (?i) is a global flag: it applies to the whole pattern, and since Python 3.11 it must sit at the very start or compilation raises re.error. (?i:...) is the scoped form and applies only inside its own group.

open as a page