skip to content

Anchors & Boundaries

The anchors that match positions rather than characters: ^ and $ (whose meaning changes under MULTILINE), the word boundary \b, and the input-level \A, \z and \Z. Interviewers use MULTILINE to check you know ^ is not always the start of input.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

5

What do the ^ and $ anchors match in a Java regular expression, and why are they called zero-width?

level: juniorimportance: must knowfreq 72%

answer

  1. Anchor = position, not character
  2. Zero-width assertion
  3. ^ start, $ end (or before trailing terminator)
  4. Default = whole input, single-line
  5. matches() already requires whole input

basics

~10 s

^ matches the start of the text and $ matches the end. They are zero-width because they match a position, not an actual character, so they consume no part of the string.

solid answer

~40 s

In Java regex, ^ asserts the start of the input and $ asserts the end (by default, the very end, allowing for a trailing line terminator). They are anchors, also called zero-width assertions: instead of matching a character, they match a position between characters. That means ^abc$ requires the whole string to be exactly 'abc' when used with Matcher.matches semantics or when the pattern spans the input. Because they consume nothing, you can combine ^ and $ to require the pattern to cover the entire line or input. By default Java is in single-line mode, so ^ and $ refer to the start and end of the whole input, not of each line. You switch to per-line behavior with the MULTILINE flag, and DOTALL changes . but not the anchors.

code

java · 9 lines
java
import java.util.regex.*;

Pattern p = Pattern.compile("^\\d+$");
System.out.println(p.matcher("123").find());   // true
System.out.println(p.matcher("12a3").find());  // false
System.out.println(p.matcher("123\n").find()); // true ($ tolerates trailing \n)

// matches() already requires the whole input, anchors optional:
System.out.println(Pattern.matches("\\d+", "123")); // true

go deeper

for a junior

Knows ^ means start, $ means end, and that they don't consume characters.

for a middle

Explains zero-width assertions, the default single-line behavior, and that $ also matches before a trailing line terminator.

for a senior

Contrasts anchors with matches()/find() semantics and knows when anchors are redundant vs load-bearing.

for a principal

Can reason about anchor behavior across modes for validation design, and articulates pitfalls like trailing-terminator tolerance in security-relevant input validation.

## What an anchor is A **regular expression** (regex) is a pattern used to find or validate text. Most parts of a pattern match **characters**: the letter `a` matches the character `a`. An **anchor** is different: it matches a **position** in the text rather than a character. Because it occupies no characters, it is called a **zero-width assertion** — an assertion (a true/false check) about a position that has **width zero**. Think of a string as having gaps between every character, plus one gap at the very start and one at the very end. For `cat` the positions are: `|c|a|t|`. An anchor tests one of those `|` gaps and either succeeds (the match continues) or fails (this attempt is rejected). It never advances the cursor. ## `^` — start `^` succeeds at the **start of the input**. So the pattern `^c` matches `cat` (there is a `c` right after the start) but `^a` does not match `cat` (the start is not followed by `a`). ## `$` — end `$` succeeds at the **end of the input**. By default Java is forgiving: `$` matches at the very end **or** just before a final line terminator (e.g. a trailing `\n`). So `t$` matches `cat` and also matches `cat\n`. ## Why combine them Because both are zero-width, `^...$` forces the pattern between them to span the whole line/input. `^\d+$` means 'the entire thing is one or more digits'. Without anchors, `\d+` would match digits **anywhere** inside a larger string. ## Default mode is single-line By default, `^` and `$` refer to the start and end of the **entire input**, even if the input contains newline characters. So in the string `"a\nb"`, `^` matches only before the first `a`, and `$` matches only after `b` (or before a trailing terminator). To make `^` and `$` match at the start and end of **each line**, you turn on the `MULTILINE` flag (covered in a related question). ## `matches()` vs `find()` Knowing the difference matters. `Pattern.matches`/`Matcher.matches` require the pattern to match the **whole** input, so `Pattern.matches("\\d+", "123")` is true without anchors. `Matcher.find` searches for the pattern **anywhere**, so there you need `^`/`$` if you want whole-string semantics. A frequent confusion is adding `^...$` inside a `matches()` call (harmless but redundant) versus omitting them in a `find()` loop (changes meaning). ## Escaping To match a literal `$` or `^` character, escape it: `\$`, `\^`. Inside a Java `String` literal you write the backslash twice: `"\\$"`.

  • Does ^abc$ guarantee the whole string equals abc when used with Matcher.find()?
    In default (single-line) mode it requires the whole input to be abc (modulo a trailing terminator for $), since ^ is only at input start and $ only at input end. With MULTILINE it would instead match any line equal to abc.
  • How do you match a literal dollar sign?
    Escape it as \$ in the pattern, written "\\$" in a Java string literal.

saying these in an interview costs you the question

  • Thinking ^/$ match an actual newline character (they match a position)
  • Believing ^ and $ work per-line by default (only with MULTILINE)
  • Forgetting that $ also matches just before a final \n by default
  • Confusing redundant anchors in matches() with required anchors in find()

context

open as a page

What does \b match in Java regex, how does \B differ, and what counts as a word character at a boundary?

level: middleimportance: must knowfreq 66%

basics

~20 s

\b matches a position where a word character is next to a non-word character (or the edge of the string) — a word boundary. \B matches the opposite: a position that is not a word boundary. Both are zero-width.

open as a page

What is the difference between $, \A, \Z, and \z in Java regex, and when does it matter for input validation?

level: seniorimportance: must knowfreq 54%

basics

~10 s

\A always matches the very start of the input. \z matches only the very end. \Z matches the end but allows one trailing line terminator. Unlike $, the input anchors ignore the MULTILINE flag.

open as a page

How does the MULTILINE flag change the meaning of ^ and $ in Java regex, and how do you enable it?

level: middleimportance: should knowfreq 58%

basics

~20 s

By default ^ and $ match the start and end of the whole input. With the MULTILINE flag they match at the start and end of every line, so they fire after each line terminator inside the text.

open as a page

Anchors are zero-width assertions. How do they relate to lookahead/lookbehind, and how would you build a custom positional assertion in Java?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Anchors like ^, $, and \b are built-in zero-width assertions that test a fixed position. Lookahead (?=...) and lookbehind (?<=...) are general zero-width assertions you define yourself, letting you build custom 'boundaries' that the standard anchors can't express.

open as a page