skip to content

Capture Groups and Substitution

Pull pieces out of a match with numbered and named groups, then rewrite text with re.sub taking a string or a callable. Interviewers use it to see you transform text without string surgery.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

How do you name a capture group in a Python re pattern, and what does .groupdict() return?

level: juniorimportance: must knowfreq 62%

answer

  1. Parentheses record what they matched
  2. Numbering follows the opening parenthesis
  3. One accessor gives a tuple, another a dict
  4. Python spells it with a P
  5. Absent groups are not empty strings

basics

~10 s

Write (?P<name>...) to name a group. The match object then exposes it as m.group('name'), and m.groupdict() returns a dict mapping every named group to the text it captured; unnamed numbered groups are not included.

solid answer

~40 s

Capturing groups are numbered by their **opening** parenthesis, left to right, from 1; group 0 is the whole match. Python's own spelling `(?P<name>...)` adds a name to a group without taking its number away, so `m.group(2)` and `m.group("start")` can be the same group. `m.groups()` returns a tuple of groups 1..N — group 0 is excluded — while `m.groupdict()` returns only the *named* ones as a dict. A group that did not participate in the match yields `None`, not an empty string, unless you pass a default: `m.groups("")`. Names must be valid identifiers and unique in the pattern; a compiled pattern exposes the mapping as `.groupindex`. Prefer names in any pattern you expect to edit, because inserting a group in the middle silently renumbers everything after it.

code

pycon · 8 lines
pycon
>>> import re
>>> m = re.search(r"(?P<chrom>chr\d+):(?P<start>\d+)-(?P<end>\d+)", "chr7:117120017-117308719")
>>> m.group("chrom"), m.group(2)
('chr7', '117120017')
>>> m.groups()
('chr7', '117120017', '117308719')
>>> m.groupdict()
{'chrom': 'chr7', 'start': '117120017', 'end': '117308719'}

go deeper

for a junior

Recall the three accessors and what each returns: group(0) is the whole match, groups() is a tuple of captures from 1, groupdict() is a dict of named groups. Be able to write (?P<name>...) from memory.

for a middle

Explain the numbering rule by opening parenthesis, why a non-participating group is None rather than empty, and how the groups()/groupdict() default argument changes that. Mention (?:...) when grouping without capturing.

for a senior

Show the maintenance judgement: numeric indexes shift when someone inserts a group, so named groups plus groupdict() are what survive a pattern edit in code that parses real files. Know the compile-time errors around duplicate names.

for a principal

Own the boundary question — when a regex with a dozen named groups has stopped being the right parser at all, and a real grammar, a line splitter, or a purpose-built format reader should take over. Name the readability and testing cost you are trading.

### What creates a group Every pair of parentheses in a Python regular expression that is not one of the `(?...)` extension forms creates a **capturing group** — a slot in which the engine records the substring that part of the pattern consumed. Groups are numbered by the position of their *opening* parenthesis, scanning left to right, starting at 1. Group 0 is reserved for the entire match. Nesting does not change the rule: in `((a)(b))` the outer group is 1, `(a)` is 2 and `(b)` is 3. Counting parentheses by hand is exactly what goes wrong on any pattern longer than one line, which is the practical argument for naming. ### Naming: `(?P<name>...)` Python spells a named group `(?P<name>...)`. The `P` is Python's own marker on the shared `(?...)` extension syntax — the .NET/PCRE spelling `(?<name>...)` is a syntax error here, and getting that wrong is one of the most common small mistakes in an interview. The name must be a valid Python identifier and must be unique within the pattern; reusing one raises `re.error` at compile time ("redefinition of group name"). A named group is *still* a numbered group: naming is additive, so `\1` and `m.group(1)` keep working alongside `m.group("name")`. ### The accessors Given a match object `m`: * `m.group()` and `m.group(0)` return the whole matched substring. * `m.group(n)` returns group `n`; `m.group("name")` returns a named group. Passing several indexes, `m.group(1, 3)`, returns a tuple. * `m.groups()` returns a tuple of **all capturing groups, 1 through N** — group 0 is not in it, and neither are non-capturing `(?:...)` groups, which record nothing. * `m.groupdict()` returns `{name: text}` for the **named groups only**. A pattern with no named groups gives an empty dict. * `m.lastindex` and `m.lastgroup` name the last group that matched, which is occasionally useful for alternations. * On the compiled pattern, `.groupindex` is a mapping of name to group number. ```python import re m = re.search(r"(?P<chrom>chr\d+):(?P<start>\d+)-(?P<end>\d+)", "chr7:117120017-117308719") m.group("chrom") # 'chr7' m.group(2) # '117120017' — same group as 'start' m.groups() # ('chr7', '117120017', '117308719') m.groupdict() # {'chrom': 'chr7', 'start': ..., 'end': ...} ``` ### Groups that did not participate An optional group, or a group inside an alternation branch the engine did not take, captured nothing. Python reports that as `None`, not `""`. This is the single most common source of `TypeError: int() argument must be...` in parsing code: `int(m.group("version"))` blows up on the record that omitted the suffix. Both `groups()` and `groupdict()` accept a default to substitute instead — `m.groups("")`, `m.groupdict("")` — and a bare `m.group("version")` still returns `None`, so a guard or a default is your choice per call site. ```python m = re.fullmatch(r"(\w+?)(?:\.(\d+))?", "BRCA1") m.groups() # ('BRCA1', None) m.groups("-") # ('BRCA1', '-') ``` Note the `(?:...)` around the optional suffix in that pattern: it groups `\.(\d+)` so the `?` applies to the whole thing, without adding a capture slot of its own. ### Referring to a group from inside the pattern Groups can be referenced later in the same pattern — a **backreference** — numerically as `\1` or by name as `(?P=name)`. That is how you assert that two parts of the text are identical, such as a repeated word or a closing tag matching its opener: ```python re.findall(r"<(?P<t>\w+)>.*?</(?P=t)>", "<a>1</a><b>2</b>") # ['a', 'b'] ``` ### Why names are worth the extra characters Three practical reasons. First, edits: add one group near the front of a pattern and every numeric index after it shifts, silently breaking `m.group(3)` and any `\2` downstream — names are immune. Second, readability at the call site: `m.group("start")` says what it means where `m.group(2)` needs the reader to go count. Third, `m.groupdict()` drops straight into a record — in a text-extraction pipeline over annotation lines, `groupdict()` is often the row you were going to build by hand anyway. The cost is that a named group is a capturing group, so where you only need grouping for a quantifier or an alternation, use `(?:...)` and keep the numbering clean. ### In an interview Say the three things out loud: numbering starts at 1 and 0 is the whole match; `groups()` is a tuple of all captures while `groupdict()` is a dict of named ones; non-participating groups are `None`. That is the whole answer, and offering the `None` detail unprompted is what makes it sound like you have parsed real text.

  • How do you refer back to an earlier group from inside the same regex pattern?
    With a backreference: `\1` by number, or `(?P=name)` by name. It asserts that the text here is identical to what that group already captured — `r"\b(\w+) \1\b"` finds a doubled word, and `r"<(?P<t>\w+)>.*?</(?P=t)>"` requires the closing tag to repeat the opening one. Note the numeric form `\1` means a backreference inside a pattern, but a *group reference* inside a re.sub replacement string; same spelling, two different grammars.
  • Can two groups in one Python pattern share a name?
    No. `re.compile` raises `re.error` with "redefinition of group name" at compile time, unlike some other engines that allow duplicate names across alternation branches. If two alternatives need to yield the same field, either give them distinct names and coalesce in Python, or restructure so the shared part sits in one group outside the alternation.
  • What is the difference between m.groups() and m.groupdict() on a pattern that mixes named and unnamed groups?
    `groups()` returns every capturing group in numeric order, named or not, as a tuple starting at group 1. `groupdict()` returns only the named ones, keyed by name; the unnamed captures are simply absent. So a pattern with three groups of which one is named gives a 3-tuple and a 1-entry dict.

saying these in an interview costs you the question

  • Says the first capture group is number 0
  • Claims .groups() includes the whole match
  • Expects .groupdict() to contain unnamed numbered groups
  • Assumes a non-participating group captures an empty string
  • Writes the .NET spelling (?<name>...) in a Python pattern
  • Thinks naming a group removes its number

context

open as a page

When should the repl argument to re.sub be a callable instead of a replacement string?

level: middleimportance: must knowfreq 52%

basics

~20 s

Pass a callable whenever the replacement depends on what matched — arithmetic, a lookup, a conditional skip. re.sub calls it once per match with the Match object and inserts whatever str it returns, with no template escape processing.

open as a page

Why does re.split return the separators when the split pattern contains a capture group?

level: middleimportance: should knowfreq 34%

basics

~20 s

re.split interleaves the text of every capturing group into the result list, by design, so you can reassemble the string. Make the delimiter group non-capturing with (?:...), or drop the parentheses, to get only the fields.

open as a page

Why is passing data-derived text as re.sub's replacement string unsafe, and what is the fix?

level: seniorimportance: should knowfreq 30%

basics

~20 s

The replacement string is a template that re.sub scans for group references and escapes, so a backslash in the data is interpreted rather than inserted. It corrupts output or raises re.error. Pass a callable returning the literal text instead.

open as a page