skip to content

In Go's regexp package, how do you read a named capture group out of a match?

level: middleimportance: should knowfreq 42%

answer

  1. the name lives on the pattern, not on the match
  2. two slices of the same length, side by side
  3. slot zero belongs to no group
  4. look the number up, then index as usual
  5. a missing name resolves to a negative index

basics

~20 s

Name a group in the pattern with (?P<name>...), then call SubexpIndex("name") to get its slot number and index the match slice with it. SubexpNames returns the names as a slice parallel to the match, with an empty string at slot 0.

solid answer

~40 s

You write the group as `(?P<key>\w+)` in the pattern, and then still index the ordinary `[]string` that `FindStringSubmatch` returns — the name is a way of resolving the number, not a different return type. `re.SubexpIndex("key")` gives that number, or `-1` if the pattern defines no such name, so guard it before using it as an index. `re.SubexpNames()` gives the whole picture: a `[]string` exactly as long as a match slice, where slot 0 is always `""` because it stands for the whole match, unnamed groups are `""`, and named groups carry their name at their own index. Ranging over `SubexpNames` next to a match is how you build a `map[string]string` of everything the pattern captured, and it is the quickest way to check that names and indices line up as you expect.

code

go · 8 lines
go
re := regexp.MustCompile(`(?P<key>\w+)=(?P<val>\d+)`)
m := re.FindStringSubmatch("retries=5")

fmt.Println(re.SubexpNames()) // prints: [ key val]  (slot 0 is always "")
fmt.Println(len(m), len(re.SubexpNames())) // prints: 3 3

fmt.Println(m[re.SubexpIndex("val")]) // prints: 5
fmt.Println(re.SubexpIndex("missing")) // prints: -1

go deeper

for a junior

Know the (?P<name>...) spelling and that you still get back a numbered slice of strings, not a map. Being able to fetch one named group correctly is enough at this level.

for a middle

Explain the parallel layout: SubexpNames is NumSubexp()+1 long, slot 0 is empty because it stands for the whole match, and unnamed groups hold empty strings. Show the loop that builds a map from it.

for a senior

Argue for names over numbers on patterns other people will edit, and show where you resolve a name — once at startup, with the -1 check — so a renamed group fails on the first run instead of on one unlucky input.

for a principal

Own the boundary call: past a certain complexity a named-group extraction is a parser wearing a costume, and you decide when the team moves to a real parser or a typed decoder rather than growing the expression.

## Naming a group A capture group is named in the pattern itself with the syntax `(?P<name>re)` — for example `(?P<key>\w+)=(?P<val>\d+)`. Recent Go also accepts the shorter `(?<name>re)` spelling for the same thing. Naming changes nothing about matching: the group still captures, still occupies a numbered slot, and still counts toward `NumSubexp()`. What the name buys you is a stable way to find that slot without hard-coding a number that shifts the moment somebody adds a group earlier in the pattern. ## The match slice is still numbered This is the point people miss. `FindStringSubmatch` returns a plain `[]string` indexed by number, exactly as it does for an unnamed pattern. There is no `m["key"]`. The names are metadata on the *compiled pattern*, so you ask the pattern for the number and then index the match: ```go re := regexp.MustCompile(`(?P<key>\w+)=(?P<val>\d+)`) m := re.FindStringSubmatch("retries=5") if m != nil { fmt.Println(m[re.SubexpIndex("val")]) // 5 } ``` ## SubexpIndex `func (re *Regexp) SubexpIndex(name string) int` returns the index of the first subexpression with that name, or **`-1`** if the pattern has no group by that name. `-1` is a perfectly ordinary `int`, so `m[re.SubexpIndex("typo")]` panics with an out-of-range index rather than failing politely. Because the pattern is a constant of the program, the honest place to catch a typo is once at startup: ```go keyIdx := re.SubexpIndex("key") if keyIdx < 0 { panic("pattern lost its key group") } ``` That converts a mid-run panic on some unlucky input line into a failure the first test run catches. ## SubexpNames `func (re *Regexp) SubexpNames() []string` returns the names of all the parenthesized subexpressions, and its layout is deliberately parallel to a match slice: - its length is `NumSubexp() + 1`, the same as any slice `FindStringSubmatch` returns; - **slot 0 is always the empty string**, because slot 0 of a match is the whole matched text, which is not a group and cannot be named; - an unnamed group carries `""` at its index; - a named group carries its name at its index. For `(?P<key>\w+)=(?P<val>\d+)` the result is `[]string{"", "key", "val"}`. The returned slice must not be modified — it belongs to the compiled pattern, which may be shared. That parallel layout makes one loop the canonical way to turn a match into a map: ```go fields := map[string]string{} for i, name := range re.SubexpNames() { if i > 0 && name != "" { fields[name] = m[i] } } ``` The `i > 0` test skips the whole-match slot and the `name != ""` test skips unnamed groups, so the map holds exactly the named captures. Note that a named group which took part in no match contributes an empty string here, indistinguishable from one that matched empty text — the same limitation the numbered API has. ## Why names are worth the noise A pattern like `^(\S+)\s+(\S+)\s+(\S+)$` is read by counting parentheses, and every reader counts again. Worse, wrapping an existing group or inserting an alternative renumbers everything after it, and the code that indexes `m[2]` keeps compiling while silently reading the wrong field. Naming both documents the pattern inline and makes the extraction survive edits, at the cost of a slightly denser expression. Names also carry over to replacement templates: `ReplaceAllString` expands `${key}` using exactly the names `SubexpNames` reports, so one naming decision serves both reading and rewriting. ## Debugging alignment When a match does not contain what you expected, the fastest diagnostic is to print `SubexpNames()` and the match slice together — two slices of the same length, side by side — for one line that matches and one that does not. Misalignment, an unnamed group you forgot about, or a `-1` from `SubexpIndex` becomes obvious in one glance, and it takes far less time than re-reading the regular expression.

  • Why is the first element of SubexpNames always the empty string?
    Because slot 0 of a match is the whole matched text rather than a capture group, and the whole match cannot be named. Keeping the empty placeholder there makes `SubexpNames()` exactly as long as any match slice and index-for-index parallel with it, which is what lets you range over the two together.
  • How do you build a map of named captures from one match?
    Range over `re.SubexpNames()` with its index, skip index 0 and skip entries whose name is empty, and use the remaining indices to read the match slice: `fields[name] = m[i]`. The two slices are the same length by construction, so the indices always line up.
  • What happens if you index a match with the result of SubexpIndex for a name the pattern does not define?
    `SubexpIndex` returns `-1`, and indexing a slice with `-1` panics with an out-of-range error at runtime. Resolve names once at startup and fail loudly there, rather than resolving them per line and discovering the typo on whichever input first reaches that branch.

saying these in an interview costs you the question

  • Expects the match to be a map keyed by group name
  • Thinks SubexpNames starts with the first group's name
  • Uses SubexpIndex's result without checking for -1
  • Believes naming a group changes what the pattern matches
  • Assumes an unnamed group is left out of SubexpNames