skip to content

Regular Expressions

java.util.regex end to end: the compile-and-match model, character classes, quantifiers, groups, lookaround, anchors, flags, replacement and splitting, and the backtracking performance cliff. Interviewers use regex both for practical parsing and for the ReDoS security angle.

part ofJavaoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

What do the ^ and $ anchors match in a Java regular expression, and why are they called zero-width?

level: juniorimportance: must knowfreq 72%

answer

  1. Anchor = position, not character
  2. Zero-width assertion
  3. ^ start, $ end (or before trailing terminator)
  4. Default = whole input, single-line
  5. matches() already requires whole input

basics

~10 s

^ matches the start of the text and $ matches the end. They are zero-width because they match a position, not an actual character, so they consume no part of the string.

solid answer

~40 s

In Java regex, ^ asserts the start of the input and $ asserts the end (by default, the very end, allowing for a trailing line terminator). They are anchors, also called zero-width assertions: instead of matching a character, they match a position between characters. That means ^abc$ requires the whole string to be exactly 'abc' when used with Matcher.matches semantics or when the pattern spans the input. Because they consume nothing, you can combine ^ and $ to require the pattern to cover the entire line or input. By default Java is in single-line mode, so ^ and $ refer to the start and end of the whole input, not of each line. You switch to per-line behavior with the MULTILINE flag, and DOTALL changes . but not the anchors.

code

java · 9 lines
java
import java.util.regex.*;

Pattern p = Pattern.compile("^\\d+$");
System.out.println(p.matcher("123").find());   // true
System.out.println(p.matcher("12a3").find());  // false
System.out.println(p.matcher("123\n").find()); // true ($ tolerates trailing \n)

// matches() already requires the whole input, anchors optional:
System.out.println(Pattern.matches("\\d+", "123")); // true

go deeper

for a junior

Knows ^ means start, $ means end, and that they don't consume characters.

for a middle

Explains zero-width assertions, the default single-line behavior, and that $ also matches before a trailing line terminator.

for a senior

Contrasts anchors with matches()/find() semantics and knows when anchors are redundant vs load-bearing.

for a principal

Can reason about anchor behavior across modes for validation design, and articulates pitfalls like trailing-terminator tolerance in security-relevant input validation.

## What an anchor is A **regular expression** (regex) is a pattern used to find or validate text. Most parts of a pattern match **characters**: the letter `a` matches the character `a`. An **anchor** is different: it matches a **position** in the text rather than a character. Because it occupies no characters, it is called a **zero-width assertion** — an assertion (a true/false check) about a position that has **width zero**. Think of a string as having gaps between every character, plus one gap at the very start and one at the very end. For `cat` the positions are: `|c|a|t|`. An anchor tests one of those `|` gaps and either succeeds (the match continues) or fails (this attempt is rejected). It never advances the cursor. ## `^` — start `^` succeeds at the **start of the input**. So the pattern `^c` matches `cat` (there is a `c` right after the start) but `^a` does not match `cat` (the start is not followed by `a`). ## `$` — end `$` succeeds at the **end of the input**. By default Java is forgiving: `$` matches at the very end **or** just before a final line terminator (e.g. a trailing `\n`). So `t$` matches `cat` and also matches `cat\n`. ## Why combine them Because both are zero-width, `^...$` forces the pattern between them to span the whole line/input. `^\d+$` means 'the entire thing is one or more digits'. Without anchors, `\d+` would match digits **anywhere** inside a larger string. ## Default mode is single-line By default, `^` and `$` refer to the start and end of the **entire input**, even if the input contains newline characters. So in the string `"a\nb"`, `^` matches only before the first `a`, and `$` matches only after `b` (or before a trailing terminator). To make `^` and `$` match at the start and end of **each line**, you turn on the `MULTILINE` flag (covered in a related question). ## `matches()` vs `find()` Knowing the difference matters. `Pattern.matches`/`Matcher.matches` require the pattern to match the **whole** input, so `Pattern.matches("\\d+", "123")` is true without anchors. `Matcher.find` searches for the pattern **anywhere**, so there you need `^`/`$` if you want whole-string semantics. A frequent confusion is adding `^...$` inside a `matches()` call (harmless but redundant) versus omitting them in a `find()` loop (changes meaning). ## Escaping To match a literal `$` or `^` character, escape it: `\$`, `\^`. Inside a Java `String` literal you write the backslash twice: `"\\$"`.

  • Does ^abc$ guarantee the whole string equals abc when used with Matcher.find()?
    In default (single-line) mode it requires the whole input to be abc (modulo a trailing terminator for $), since ^ is only at input start and $ only at input end. With MULTILINE it would instead match any line equal to abc.
  • How do you match a literal dollar sign?
    Escape it as \$ in the pattern, written "\\$" in a Java string literal.

saying these in an interview costs you the question

  • Thinking ^/$ match an actual newline character (they match a position)
  • Believing ^ and $ work per-line by default (only with MULTILINE)
  • Forgetting that $ also matches just before a final \n by default
  • Confusing redundant anchors in matches() with required anchors in find()

context

open as a page

What is a character class in a Java regular expression, and what does [abc] match?

level: juniorimportance: must knowfreq 78%

basics

~10 s

A character class is a set in square brackets that matches exactly ONE character from that set. [abc] matches a single 'a', 'b', or 'c' (not the word "abc").

open as a page

How do you make a Java regex match case-insensitively, and what is the difference between the CASE_INSENSITIVE flag and the (?i) inline flag?

level: juniorimportance: must knowfreq 70%

basics

~10 s

Pass Pattern.CASE_INSENSITIVE when compiling the pattern, or put (?i) at the start of the pattern string. Both make letters match regardless of upper or lower case, so "abc" matches "ABC".

open as a page

What is a capturing group in a Java regular expression, and how do you retrieve what it matched?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A capturing group is the part of a pattern wrapped in parentheses (). After a successful match you read what it captured with matcher.group(n), where n is the group's number. group(0) is the whole match.

open as a page

What is a lookaround assertion in a Java regular expression, and what does "zero-width" mean?

level: juniorimportance: must knowfreq 60%

basics

~20 s

A lookaround is a regex condition that checks whether text around the current spot matches a pattern, without including that text in the match. "Zero-width" means it consumes no characters - the position does not move.

open as a page

What is the difference between Matcher.matches() and Matcher.find() in Java regex?

level: juniorimportance: must knowfreq 80%

basics

~20 s

matches() returns true only if the regex matches the ENTIRE input from start to end. find() looks for the next place the pattern matches ANYWHERE in the input, and can be called again to find more.

open as a page

What are the basic regex quantifiers in Java (* + ? and {n}/{n,}/{n,m}) and what does each one match?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Quantifiers say how many times the thing before them may repeat. * means zero or more, + means one or more, ? means zero or one. {n} means exactly n, {n,} means n or more, {n,m} means between n and m.

open as a page

Why does a Java regex for a digit use "\\d" with two backslashes, and how does this differ from the regex itself?

level: juniorimportance: must knowfreq 60%

basics

~20 s

The regex needs the text \d, but in a Java String a single backslash starts an escape sequence. So you write "\d" — the first backslash escapes the second, and the String value handed to the regex is actually \d.

open as a page

In Java, what is the difference between String.replaceAll and String.replaceFirst, and how do they differ from String.replace?

level: juniorimportance: must knowfreq 70%

basics

~10 s

replaceAll changes every match of a regex; replaceFirst changes only the first match. Both treat their first argument as a regular expression. String.replace is different: it works on plain literal text, not regex.

open as a page

What does \b match in Java regex, how does \B differ, and what counts as a word character at a boundary?

level: middleimportance: must knowfreq 66%

basics

~20 s

\b matches a position where a word character is next to a non-word character (or the edge of the string) — a word boundary. \B matches the opposite: a position that is not a word boundary. Both are zero-width.

open as a page

What does the dot (.) match in Java regex, and how does it differ from a character class with respect to line terminators and the DOTALL flag?

level: middleimportance: must knowfreq 71%

basics

~20 s

By default the dot matches any single character EXCEPT a line terminator like \n. With the DOTALL flag it matches every character including newlines. A character class like [^x] matches newlines unless you exclude them.

open as a page

How do ranges [a-z] and negation [^abc] work in a Java character class, and how do you include a literal hyphen or caret?

level: middleimportance: must knowfreq 74%

basics

~20 s

[a-z] matches any one character whose code falls between 'a' and 'z'. [^abc] matches any one character that is NOT a, b, or c. Put a hyphen first or last (or escape it) to mean a literal dash.

open as a page

Explain the difference between Pattern.MULTILINE (?m) and Pattern.DOTALL (?s) in Java regex. Which anchors/metacharacters does each affect?

level: middleimportance: must knowfreq 65%

basics

~20 s

MULTILINE makes ^ and $ match at the start and end of every line, not just the whole input. DOTALL makes the dot '.' match newline characters too. They control different things and are often confused.

open as a page

Explain the difference between positive and negative lookahead/lookbehind, with a Java example of each.

level: middleimportance: must knowfreq 58%

basics

~20 s

Positive means "require this pattern here" - (?=X) require X after, (?<=X) require X before. Negative means "require this pattern NOT here" - (?!X) X must not follow, (?<!X) X must not precede. Either way nothing is consumed.

open as a page

What is the difference between greedy and reluctant (lazy) quantifiers in Java regex, and when does it matter?

level: middleimportance: must knowfreq 68%

basics

~20 s

Greedy quantifiers (the default) grab as much text as possible, then give some back if needed. Lazy quantifiers, written with a trailing ?, grab as little as possible and only take more if forced. They differ in how much text each match consumes.

open as a page

What is backtracking in Java's regex engine, and why can it become a performance problem?

level: middleimportance: must knowfreq 55%

basics

~20 s

Java's regex tries one way to match; if it fails, it goes back and tries another. With certain patterns the number of ways to try explodes, so matching can take a huge amount of time on some inputs.

open as a page

What is the difference between $, \A, \Z, and \z in Java regex, and when does it matter for input validation?

level: seniorimportance: must knowfreq 54%

basics

~10 s

\A always matches the very start of the input. \z matches only the very end. \Z matches the end but allows one trailing line terminator. Unlike $, the input anchors ignore the MULTILINE flag.

open as a page

What is a ReDoS attack, and how could a user-supplied regex or input take down a Java service?

level: seniorimportance: must knowfreq 50%

basics

~20 s

ReDoS (Regular-expression Denial of Service) is when an attacker sends an input that makes a vulnerable regex take a huge amount of time to evaluate, freezing the thread and starving the service of CPU until it can't serve real users.

open as a page

How does the limit argument to String.split (and Pattern.split) work, and what is the default trailing-empty-string behavior?

level: seniorimportance: must knowfreq 62%

basics

~20 s

split(regex) with no limit (or limit 0) removes trailing empty strings. A positive limit caps the number of pieces and keeps trailing empties; a negative limit keeps ALL trailing empties with no cap. Limit 0 is the default and the source of most surprises.

open as a page

How does the MULTILINE flag change the meaning of ^ and $ in Java regex, and how do you enable it?

level: middleimportance: should knowfreq 58%

basics

~20 s

By default ^ and $ match the start and end of the whole input. With the MULTILINE flag they match at the start and end of every line, so they fire after each line terminator inside the text.

open as a page

What do the predefined classes \d, \w, \s mean in Java regex, how do they relate to POSIX classes like \p{Alpha}, and what are their ASCII vs Unicode caveats?

level: middleimportance: should knowfreq 68%

basics

~10 s

\d is a digit, \w is a word character (letters, digits, underscore), \s is whitespace. Their uppercase forms (\D, \W, \S) are the negations. By default they match ASCII only unless you enable Unicode.

open as a page

How do you combine multiple Pattern flags in Java, and how does combining compile-time flags compare to combining inline flags? Give the bitwise mechanism.

level: middleimportance: should knowfreq 35%

basics

~10 s

Combine compile-time flags with the bitwise OR operator |, e.g. Pattern.CASE_INSENSITIVE | Pattern.MULTILINE. Inline, just list the letters together, e.g. (?im). Both ways enable several flags at once.

open as a page

What does Pattern.COMMENTS / (?x) do, and what are the rules and pitfalls when using it to write a readable regex?

level: middleimportance: should knowfreq 30%

basics

~20 s

COMMENTS (inline (?x)) lets you spread a regex over multiple lines: it ignores unescaped whitespace and treats # as a comment to end of line. To match a real space you must escape it (\ ) or use \s.

open as a page

How do named capturing groups work in Java regex, and how do you reference them?

level: middleimportance: should knowfreq 55%

basics

~10 s

Write (?<name>...) to give a capturing group a name. Read it back with matcher.group("name") instead of a number, and refer to it inside the pattern with \k<name>. Names make patterns readable.

open as a page

What is a non-capturing group (?:...) in Java regex, and why would you use one instead of a plain ()?

level: middleimportance: should knowfreq 62%

basics

~20 s

A non-capturing group (?:...) groups part of a pattern so a quantifier or alternation applies to it, but it does not capture or get a group number. Use it when you only need grouping, not the captured text.

open as a page

How do you extract multiple groups and their positions from text in Java, including across multiple matches?

level: middleimportance: should knowfreq 58%

basics

~10 s

Compile the pattern, get a Matcher, loop with while(m.find()), and for each match read m.group(n) and m.start(n)/m.end(n). find() advances to the next match each call; matches() only checks the whole input once.

open as a page

Why should you reuse a compiled Pattern instead of calling String.matches() repeatedly?

level: middleimportance: should knowfreq 62%

basics

~10 s

Compiling a regex is expensive. String.matches() recompiles the pattern on every call. If you use a regex many times, compile it once with Pattern.compile() and a static final field, then reuse it.

open as a page

How do capturing-group back-references like $1 work in a Java regex replacement string, and how do you insert a literal $ or backslash?

level: middleimportance: should knowfreq 58%

basics

~20 s

Inside the replacement string, $1, $2, ... insert the text captured by the matching parentheses (groups) of the pattern. To put a literal $ or \ in the output instead, escape it as $ / \, or wrap the whole replacement in Matcher.quoteReplacement.

open as a page

What does Pattern.quote do, and when must you use it when building a Java regex?

level: middleimportance: should knowfreq 50%

basics

~20 s

Pattern.quote(s) returns a regex that matches the string s exactly, treating every character as literal. Use it whenever you build a pattern from text that might contain regex metacharacters like . * + ( ) [ ], especially user input.

open as a page

How do class union and the Java-specific class intersection (&&) and subtraction work, e.g. [a-z&&[^aeiou]]?

level: seniorimportance: should knowfreq 42%

basics

~10 s

Listing members in a class unions them: [a-d[m-p]] is a..d or m..p. Java adds intersection with &&: [a-z&&[^aeiou]] means lowercase letters AND not a vowel, i.e. consonants. It is a set-difference trick.

open as a page

showing 1–30 of 44