skip to content

What do the ^ and $ anchors match in a Java regular expression, and why are they called zero-width?

level: juniorimportance: must knowfreq 72%

answer

  1. Anchor = position, not character
  2. Zero-width assertion
  3. ^ start, $ end (or before trailing terminator)
  4. Default = whole input, single-line
  5. matches() already requires whole input

basics

~10 s

^ matches the start of the text and $ matches the end. They are zero-width because they match a position, not an actual character, so they consume no part of the string.

solid answer

~40 s

In Java regex, ^ asserts the start of the input and $ asserts the end (by default, the very end, allowing for a trailing line terminator). They are anchors, also called zero-width assertions: instead of matching a character, they match a position between characters. That means ^abc$ requires the whole string to be exactly 'abc' when used with Matcher.matches semantics or when the pattern spans the input. Because they consume nothing, you can combine ^ and $ to require the pattern to cover the entire line or input. By default Java is in single-line mode, so ^ and $ refer to the start and end of the whole input, not of each line. You switch to per-line behavior with the MULTILINE flag, and DOTALL changes . but not the anchors.

code

java · 9 lines
java
import java.util.regex.*;

Pattern p = Pattern.compile("^\\d+$");
System.out.println(p.matcher("123").find());   // true
System.out.println(p.matcher("12a3").find());  // false
System.out.println(p.matcher("123\n").find()); // true ($ tolerates trailing \n)

// matches() already requires the whole input, anchors optional:
System.out.println(Pattern.matches("\\d+", "123")); // true

go deeper

for a junior

Knows ^ means start, $ means end, and that they don't consume characters.

for a middle

Explains zero-width assertions, the default single-line behavior, and that $ also matches before a trailing line terminator.

for a senior

Contrasts anchors with matches()/find() semantics and knows when anchors are redundant vs load-bearing.

for a principal

Can reason about anchor behavior across modes for validation design, and articulates pitfalls like trailing-terminator tolerance in security-relevant input validation.

## What an anchor is A **regular expression** (regex) is a pattern used to find or validate text. Most parts of a pattern match **characters**: the letter `a` matches the character `a`. An **anchor** is different: it matches a **position** in the text rather than a character. Because it occupies no characters, it is called a **zero-width assertion** — an assertion (a true/false check) about a position that has **width zero**. Think of a string as having gaps between every character, plus one gap at the very start and one at the very end. For `cat` the positions are: `|c|a|t|`. An anchor tests one of those `|` gaps and either succeeds (the match continues) or fails (this attempt is rejected). It never advances the cursor. ## `^` — start `^` succeeds at the **start of the input**. So the pattern `^c` matches `cat` (there is a `c` right after the start) but `^a` does not match `cat` (the start is not followed by `a`). ## `$` — end `$` succeeds at the **end of the input**. By default Java is forgiving: `$` matches at the very end **or** just before a final line terminator (e.g. a trailing `\n`). So `t$` matches `cat` and also matches `cat\n`. ## Why combine them Because both are zero-width, `^...$` forces the pattern between them to span the whole line/input. `^\d+$` means 'the entire thing is one or more digits'. Without anchors, `\d+` would match digits **anywhere** inside a larger string. ## Default mode is single-line By default, `^` and `$` refer to the start and end of the **entire input**, even if the input contains newline characters. So in the string `"a\nb"`, `^` matches only before the first `a`, and `$` matches only after `b` (or before a trailing terminator). To make `^` and `$` match at the start and end of **each line**, you turn on the `MULTILINE` flag (covered in a related question). ## `matches()` vs `find()` Knowing the difference matters. `Pattern.matches`/`Matcher.matches` require the pattern to match the **whole** input, so `Pattern.matches("\\d+", "123")` is true without anchors. `Matcher.find` searches for the pattern **anywhere**, so there you need `^`/`$` if you want whole-string semantics. A frequent confusion is adding `^...$` inside a `matches()` call (harmless but redundant) versus omitting them in a `find()` loop (changes meaning). ## Escaping To match a literal `$` or `^` character, escape it: `\$`, `\^`. Inside a Java `String` literal you write the backslash twice: `"\\$"`.

  • Does ^abc$ guarantee the whole string equals abc when used with Matcher.find()?
    In default (single-line) mode it requires the whole input to be abc (modulo a trailing terminator for $), since ^ is only at input start and $ only at input end. With MULTILINE it would instead match any line equal to abc.
  • How do you match a literal dollar sign?
    Escape it as \$ in the pattern, written "\\$" in a Java string literal.

saying these in an interview costs you the question

  • Thinking ^/$ match an actual newline character (they match a position)
  • Believing ^ and $ work per-line by default (only with MULTILINE)
  • Forgetting that $ also matches just before a final \n by default
  • Confusing redundant anchors in matches() with required anchors in find()

context