How do you extract multiple groups and their positions from text in Java, including across multiple matches?
answer
- Pattern (reuse, thread-safe) vs Matcher (per-input, stateful)
- while(m.find()) iterates non-overlapping matches
- group(n) + start(n)/end(n); end is exclusive
- matches()=whole input, find()=next anywhere; groupCount() excludes group 0
basics
~10 sCompile the pattern, get a Matcher, loop with while(m.find()), and for each match read m.group(n) and m.start(n)/m.end(n). find() advances to the next match each call; matches() only checks the whole input once.
solid answer
~40 sTo pull structured fields out of text, compile a Pattern once with capturing groups, then call pattern.matcher(input) to get a Matcher. For repeated occurrences, loop with while (m.find()): each call advances to the next non-overlapping match. Within each iteration, m.group(n) gives the captured text of group n, and m.start(n)/m.end(n) give its character offsets in the input (end is exclusive; group 0 covers the whole match). Use matches() instead of find() only when the entire input must conform to the pattern — it doesn't iterate. Reuse the compiled Pattern (it's immutable and thread-safe) but not the Matcher (stateful, not thread-safe). For null groups (non-participating optional groups), guard before using the value. groupCount() tells you how many capturing groups exist, excluding group 0.
code
java · 9 linesPattern p = Pattern.compile("(\\w+)=(\\d+)");
Matcher m = p.matcher("a=1 b=22 c=333");
while (m.find()) {
System.out.printf("%s -> %s [%d..%d]%n",
m.group(1), m.group(2), m.start(2), m.end(2));
}
// a -> 1 [2..3]
// b -> 22 [6..8]
// c -> 333 [11..14]go deeper
Can loop with while(find()) and read group(n).
Distinguishes matches/lookingAt/find, uses start/end positions, and reuses Pattern vs Matcher correctly.
Handles null/-1 for non-participating groups, knows groupCount semantics, reset(), and thread-safety implications.
Designs for performance (compiled-pattern reuse, ReDoS-safe patterns) and concurrency, and chooses extraction strategy at scale.
## The two-object model Java regex separates the **pattern** from the **matching engine**: - `Pattern.compile("...")` parses the regex once into an immutable, thread-safe `Pattern`. Compilation is relatively expensive, so do it once and reuse it (e.g. a `static final` field). - `pattern.matcher(input)` creates a `Matcher` — a stateful cursor over one specific input. It is **not** thread-safe; create a fresh one per input/thread. ## Single match: matches() vs lookingAt() vs find() - `matches()` — true only if the **entire** input matches the pattern. Anchors the whole string. - `lookingAt()` — true if the input matches starting at the **beginning** (but need not reach the end). - `find()` — searches for the **next** match anywhere in the remaining input; returns true and positions the matcher on it. ## Iterating all matches The idiom for every occurrence is: ```java while (m.find()) { // process this match } ``` Each `find()` continues *after* the previous match, yielding **non-overlapping** matches left to right. This is how you tokenize a log line, pull every `key=value`, etc. ## Reading groups and positions per match Once positioned on a match (after a true `find()`/`matches()`): - `m.group()` / `m.group(0)` — the whole matched text. - `m.group(n)` — text of capturing group n. - `m.group("name")` — text of a named group. - `m.start(n)` — index where group n started (inclusive). - `m.end(n)` — index just past where group n ended (exclusive). So `input.substring(m.start(n), m.end(n))` equals `m.group(n)`. - `m.start()`/`m.end()` with no arg = bounds of the whole match. - `m.groupCount()` — number of capturing groups in the pattern, **not** counting group 0. ## Null and non-participating groups If a group is optional and didn't participate in the current match, `group(n)` is `null` and `start(n)`/`end(n)` return `-1`. Always guard optional groups before using their text. ## Worked example: key=value pairs ```java Pattern p = Pattern.compile("(\\w+)=(\\w+)"); Matcher m = p.matcher("a=1 b=2 c=3"); while (m.find()) { System.out.println(m.group(1) + " -> " + m.group(2) + " @" + m.start()); } // a -> 1 @0 ; b -> 2 @4 ; c -> 3 @8 ``` ## Resetting and reuse `m.reset()` rewinds the matcher to the start (optionally with new input via `reset(CharSequence)`), letting you reuse a Matcher object rather than allocating a new one. The Pattern itself never needs resetting. ## Common pitfalls - Calling `group()` before a successful match → `IllegalStateException`. - Sharing a `Matcher` across threads → corruption; share the `Pattern` instead. - Recompiling the same pattern in a hot loop → wasted CPU; hoist it. - Forgetting `find()` returns *non-overlapping* matches, so overlapping occurrences are missed.
- What is the difference between matches() and find()?matches() returns true only if the entire input conforms to the pattern and does not iterate; find() locates the next match anywhere in the input and can be called repeatedly in a loop to walk all non-overlapping matches.
- Why reuse the Pattern but not the Matcher across threads?Pattern is immutable and thread-safe, so one compiled instance can be shared. Matcher holds mutable per-input state (current position, group bounds) and is not thread-safe, so each thread needs its own.
The Pattern is a cookie-cutter (make it once, reuse forever); the Matcher is the dough you press it through (fresh batch each time, and you walk it across the sheet with find()).
saying these in an interview costs you the question
- Using matches() expecting it to find a substring (it requires the whole input)
- Sharing one Matcher across threads
- Recompiling the pattern inside a loop instead of hoisting it
- Treating end(n) as inclusive (it is exclusive)
- Including group 0 in groupCount() (it isn't counted)