What is the difference between Matcher.matches() and Matcher.find() in Java regex?
answer
- matches = whole input, anchored both ends
- find = next match anywhere, advances position, loop with while
- lookingAt = anchored at start only
- find() is stateful → reset() to restart
- String.matches delegates to Matcher.matches (whole string)
basics
~20 smatches() returns true only if the regex matches the ENTIRE input from start to end. find() looks for the next place the pattern matches ANYWHERE in the input, and can be called again to find more.
solid answer
~40 smatches() succeeds only when the whole input is consumed by the pattern, as if the pattern were anchored at both ends. find() searches for the next subsequence that matches anywhere in the input; it returns true if found, advances an internal position, and a subsequent find() continues from after the previous match, so a loop over find() iterates all matches. There is also lookingAt(), which anchors only at the start (the match must begin at index 0 but need not reach the end). A common bug is using matches() expecting a substring search. For example, Pattern.compile("\\d+").matcher("abc123").matches() is false (the letters aren't matched), but .find() is true and returns "123". Because find() is stateful, you call reset() to restart from the beginning.
code
java · 10 linesPattern p = Pattern.compile("\\d+");
Matcher m = p.matcher("abc123def456");
m.matches(); // false -> whole input is not all digits
m.lookingAt(); // false -> input does not START with digits
while (m.find()) { // true twice
System.out.println(m.group()); // "123", then "456"
}
m.reset(); // search the same input again from the startgo deeper
Can state the one-line difference: matches() = entire input, find() = next match anywhere, and knows the while(find()) loop.
Explains lookingAt() too and the stateful nature of find()/reset(); knows String.matches delegates to whole-string matches().
Discusses when to use each in real validation vs scanning code, the anchoring equivalence (\A..\z), and group extraction across iterations.
Frames the API choice in terms of correctness/perf (avoiding accidental matches() validation bugs, reusing compiled patterns), and can reason about region() effects on these methods.
## First principles A **regular expression** (regex) is a pattern describing a set of strings. In Java you compile a regex into a `Pattern`, then create a `Matcher` that applies that pattern to a specific input string (a `CharSequence`). The `Matcher` is the object that actually performs matching and holds the *result state* (where a match started/ended, capture groups, current search position). Three matching methods differ ONLY in **where the pattern is allowed to start and whether it must reach the end**: | Method | Must start at index 0? | Must reach the end of input? | Meaning | |---|---|---|---| | `matches()` | Yes | Yes | The ENTIRE input is the match (anchored both ends). | | `lookingAt()` | Yes | No | The match must begin at the start, but can stop early. | | `find()` | No | No | Find the NEXT match anywhere; advances position each call. | ### matches() `matches()` returns `true` only if the *whole* region (by default the whole input) is consumed by the pattern. It is equivalent to wrapping your pattern in `\A...\z` anchors. Example: pattern `\d+` against `"abc123"` → `false`, because `a` is not a digit so the entire string cannot be the match. Against `"123"` → `true`. ### find() `find()` scans forward looking for the *next subsequence* that matches. It returns `true` when it finds one and records the match (`start()`, `end()`, `group()`). Crucially it is **stateful**: each call resumes from the end of the previous match. This is why the idiomatic way to get all matches is: ```java while (matcher.find()) { /* matcher.group() is the next match */ } ``` After the loop ends `find()` has returned `false` and the matcher is exhausted; call `reset()` to search the same input again from the start. ### lookingAt() In between: it anchors at the start but, like `find()`, doesn't require consuming the whole input. `\d+` against `"123abc"` → `lookingAt()` is `true` (matches the leading `123`), but `matches()` is `false` (the `abc` tail is left over). ### String convenience methods `String.matches(regex)`, `String.split`, `String.replaceAll` internally compile the pattern and call `matches()`/`find()` for you. `"abc123".matches("\\d+")` is `false` for the same reason as above — `String.matches` delegates to `Matcher.matches()` (whole-string), NOT to `find()`. ### The classic bug Developers reach for `matches()` expecting "does this contain the pattern?" That's `find()`'s job. If you want whole-string validation (e.g. "is this a valid phone number?"), use `matches()`. If you want "does this text contain a URL somewhere?", use `find()`.
- How do you iterate over every match in a string?Loop while (matcher.find()) and read matcher.group() (and group(n) for captures) inside the loop; each iteration advances to the next non-overlapping match.
- After find() returns false at the end, how do you search the same input again?Call matcher.reset() (optionally reset(newInput) to swap the input), which clears the match state and position back to the start.
saying these in an interview costs you the question
- Thinking matches() does a substring/contains search
- Believing find() restarts from the beginning on every call
- Confusing lookingAt() with matches() (lookingAt does NOT require reaching the end)
- Assuming String.matches("\\d+") is true for "abc123"