skip to content

Submatches and Replacement

Pulling capture groups out with FindStringSubmatch and writing them back with $1 or ${name}, where an unescaped dollar in replacement text quietly eats the following characters.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

questions

4

In Go's regexp package, what does FindStringSubmatch return for a match and for a non-match?

level: juniorimportance: must knowfreq 55%

answer

  1. one flat slice per match
  2. count the opening parentheses
  3. its length comes from the pattern, not the input
  4. a search that finds nothing hands back nil, not an empty slice

basics

~20 s

FindStringSubmatch returns a []string in which element 0 is the entire matched text and elements 1..n are the capture groups, numbered by opening parenthesis. If the pattern matches nowhere in the input it returns nil, so check for nil before indexing.

solid answer

~40 s

`FindStringSubmatch` looks for the leftmost match of the pattern anywhere in the string and returns one `[]string`. Slot 0 holds the whole matched text; slots 1 through `NumSubexp()` hold the capture groups, numbered by the order of their opening parentheses. So for `(\w+)=(\d+)` against `"port=8080"` you get `["port=8080", "port", "8080"]` — length 3, not 2. When the pattern does not match at all, the return value is `nil`, not an empty slice and not a slice of empty strings, which is why `m[1]` panics on a non-matching line. The length is fixed by the pattern rather than by the input: a capture group that took part in no match still occupies its slot, holding `""`. For every match instead of the first, use `FindAllStringSubmatch(s, n)` with a negative `n`.

code

go · 9 lines
go
re := regexp.MustCompile(`(\w+)=(\d+)`)

m := re.FindStringSubmatch("port=8080")
fmt.Println(len(m)) // prints: 3
fmt.Println(m[0])   // prints: port=8080
fmt.Println(m[1])   // prints: port
fmt.Println(m[2])   // prints: 8080

fmt.Println(re.FindStringSubmatch("nothing here") == nil) // prints: true

go deeper

for a junior

Be ready to state the slice layout from memory: whole match at 0, groups from 1 in opening-parenthesis order, and nil when the pattern matches nowhere. Say the nil part out loud before you are asked.

for a middle

Explain why the length is NumSubexp()+1 regardless of input, what a non-participating group contains, and how the Find/All/String/Index name grammar composes into the variant you actually want.

for a senior

Show the guard you put in front of every such call in real code, and say when you reach for the Index variant to distinguish an empty capture from an absent one rather than trusting an empty string.

for a principal

Frame it as an API-shape choice: number-indexed submatches are brittle across pattern edits, so decide as a team whether parsing goes through names, a small typed helper, or a real parser instead of a regular expression.

## What the call does `func (re *Regexp) FindStringSubmatch(s string) []string` searches `s` for the leftmost match of the compiled pattern and returns the text of that match together with the text of each capture group. It is not anchored: the pattern may match anywhere inside `s` unless you wrote `^` and `$` yourself. ## The shape of the returned slice The result is a single flat `[]string`: - **Index 0** is the *entire* matched substring — not the first capture group. This is the single most common mistake with the API. - **Index 1..n** are the capture groups, numbered by the position of their **opening** parenthesis, left to right, including nested groups. - The length is always `re.NumSubexp() + 1`, decided by the pattern, never by the input. For `(\w+)=(\d+)` matched against `"port=8080"`: | index | value | |---|---| | 0 | `port=8080` | | 1 | `port` | | 2 | `8080` | A group that is present in the pattern but took part in no match — for example the second alternative of `(a)|(b)` when `a` matched — still has a slot, and that slot holds the empty string `""`. Nothing is skipped or shifted; the numbering stays stable, which is what makes indexing by number safe once you know the pattern. Because a non-participating group and a group that legitimately matched empty text both look like `""`, the string API cannot tell them apart. When that distinction matters, use `FindStringSubmatchIndex`, which returns a `[]int` of byte offset pairs and uses `-1, -1` for a group that did not participate. ## The non-match case If the pattern matches nowhere in `s`, the function returns **`nil`**. Not an empty non-nil slice, not a slice of empty strings, and not an error. `len(m)` is then 0 and `m == nil` is true, so any `m[0]` or `m[1]` panics with `index out of range [1] with length 0`. Every use of this API therefore has a guard in front of it: ```go m := re.FindStringSubmatch(line) if m == nil { continue // or handle the unmatched line } use(m[1]) ``` A `len(m) < 2` check works equally well and is what you want if a helper may be handed patterns with different group counts. ## Getting every match `FindStringSubmatch` returns the **first** match only. Its plural sibling is: ```go func (re *Regexp) FindAllStringSubmatch(s string, n int) [][]string ``` It returns a slice of the same per-match slices described above. The `n` argument caps how many matches you want; **a negative `n` (conventionally `-1`) means all of them**. `n` has nothing to do with capture groups. Like the singular form, it returns `nil` — here a nil `[][]string` — when there is no match at all, so ranging over the result of a failed search is safe while indexing it is not. ## The family of names The package names these methods systematically, and knowing the grammar saves memorising a dozen signatures: - `Find…` returns the whole match only; `Find…Submatch` also returns the groups. - No prefix means `[]byte` in and out (`FindSubmatch`); `String` means `string` in and out (`FindStringSubmatch`). - `All` means every match rather than the first, and takes the `n` limit. - `Index` means byte offsets (`[]int`) instead of text, and is the variant that can distinguish an empty capture from an absent one. So `FindAllStringSubmatchIndex(s, -1)` is "every match, as a string search, with groups, reported as offsets" — the name is fully compositional. ## Practical notes Indexing by number is fine for a small, stable pattern, but the numbers drift the moment somebody adds a group in the middle of the expression. Naming the groups and resolving their index by name removes that fragility. And when all you need is a yes/no answer, `MatchString` avoids allocating the result slice entirely.

  • How do you get every match in the string instead of just the first one?
    Use `FindAllStringSubmatch(s, n)`, which returns `[][]string` — one per-match slice of the same shape. A negative `n`, conventionally `-1`, means "no limit, return them all"; a positive `n` caps the count. It returns `nil` when nothing matches, so ranging over the result is safe even on a failed search.
  • What sits in the slot of a capture group that took part in no match?
    The empty string. The slice length is fixed at `NumSubexp()+1`, so no slot is ever skipped or shifted. The catch is that a group which genuinely matched empty text looks identical. If you must tell them apart, call `FindStringSubmatchIndex`, whose `[]int` result uses the pair `-1, -1` for a group that did not participate.
  • How do you decide the length of the returned slice before you run the search?
    It is always `re.NumSubexp() + 1` — one slot for the whole match plus one per capture group in the pattern. That is a property of the compiled pattern, not of the input, so you can assert it once at startup if a helper is parameterised by pattern.

saying these in an interview costs you the question

  • Says index 0 is the first capture group
  • Expects an empty slice rather than nil when nothing matches
  • Indexes m[1] without checking the result first
  • Thinks FindStringSubmatch returns every match in the input
  • Assumes a non-participating group is omitted, shifting the later indices
open as a page

In Go's regexp package, how do you read a named capture group out of a match?

level: middleimportance: should knowfreq 42%

basics

~20 s

Name a group in the pattern with (?P<name>...), then call SubexpIndex("name") to get its slot number and index the match slice with it. SubexpNames returns the names as a slice parallel to the match, with an empty string at slot 0.

open as a page

A Go tool indexes FindStringSubmatch's result at m[1] and panics mid-rewrite — what went wrong?

level: seniorimportance: should knowfreq 30%

basics

~20 s

A line did not match, so FindStringSubmatch returned nil and indexing it panicked with index out of range, length 0. Guard the result for nil before indexing, and write rewritten files atomically so a panic cannot leave half the tree edited.

open as a page

In Go, what is the difference between Regexp.ReplaceAllString and ReplaceAllLiteralString?

level: middleimportance: nice to knowfreq 38%

basics

~10 s

ReplaceAllString interprets dollar signs in the replacement text, expanding $1 and ${name} to captured groups. ReplaceAllLiteralString inserts the replacement byte for byte with no expansion, so a dollar sign stays a dollar sign.

open as a page