skip to content

Lookahead & Lookbehind

Lookahead and lookbehind assert what surrounds a position without consuming characters, which is how you express password rules or match-but-not-followed-by patterns. Java restricts lookbehind to bounded length, a detail interviewers occasionally test.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

5

What is a lookaround assertion in a Java regular expression, and what does "zero-width" mean?

level: juniorimportance: must knowfreq 60%

answer

  1. Cursor sits between chars; consume = advance + include
  2. Zero-width = checks condition, moves nothing, adds nothing to match
  3. Four kinds: (?=) (?!) (?<=) (?<!)
  4. Lookahead asserts what FOLLOWS; lookbehind what PRECEDES
  5. Not capturing groups - no group(n)

basics

~20 s

A lookaround is a regex condition that checks whether text around the current spot matches a pattern, without including that text in the match. "Zero-width" means it consumes no characters - the position does not move.

solid answer

~40 s

A lookaround assertion is a regex construct that tests whether a sub-pattern matches at the current position, but does not add those characters to the matched text. It is "zero-width" because the regex engine's cursor stays put: nothing is consumed. There are four kinds - positive lookahead (?=X) requires X to follow, negative lookahead (?!X) requires X not to follow, positive lookbehind (?<=X) requires X to precede, and negative lookbehind (?<!X) requires X not to precede. They are used to assert context (e.g. a digit only if a letter follows) without capturing that context, which keeps matches and replacements clean. In Java you write them inside Pattern/Matcher just like any other group.

code

java · 6 lines
java
import java.util.regex.*;

Matcher m = Pattern.compile("\\d(?=px)").matcher("5px 9em 7px");
while (m.find()) {
    System.out.println(m.group()); // prints 5 then 7 (digits followed by px); 9 is skipped
}

go deeper

for a junior

Can define zero-width and name the four lookarounds, and explain that the asserted text is not consumed.

for a middle

Can pick the right lookaround for a given context-sensitive match and explain why it keeps group(0) clean.

for a senior

Can reason about cursor mechanics, combine multiple lookarounds at one position, and explain replacement behavior.

for a principal

Can weigh lookaround readability/performance against alternatives (captured-group rebuild, split, parser) and set team conventions.

## First principles: how a regex engine moves When a regular-expression engine matches text, imagine a **cursor** sitting between characters. As the engine matches ordinary characters like `a` or `\d`, it **consumes** them: the cursor moves forward and those characters become part of the matched substring. "Consuming" simply means "this character is now part of what we matched, advance past it." ## What "zero-width" means A **zero-width** construct is one that checks a condition at the cursor's current position **without moving the cursor and without adding any character to the match**. The width of the thing it matches is zero. Familiar zero-width constructs you may already know are `^` (start of line) and `$` (end of line) - they assert a position, they don't eat a character. A **lookaround assertion** is a zero-width construct that lets you assert *an entire sub-pattern* matches (or fails to match) just before or just after the cursor, while still consuming nothing. ## The four lookarounds - **Positive lookahead** `(?=X)` - succeeds if pattern `X` matches starting at the current position. The cursor does **not** advance. - **Negative lookahead** `(?!X)` - succeeds if pattern `X` does **not** match at the current position. - **Positive lookbehind** `(?<=X)` - succeeds if pattern `X` matches the text **immediately ending** at the current position (i.e. just behind the cursor). - **Negative lookbehind** `(?<!X)` - succeeds if `X` does **not** match immediately behind the cursor. ## Why this is useful Because lookarounds consume nothing, the characters they test are **not part of `group(0)`** (the matched text). This matters most in two cases: 1. **Replacement** - `replaceAll` only replaces the consumed text. If you assert context with a lookaround, that context survives the replacement. Example: insert a comma between digit groups without deleting any digit. 2. **Tokenizing/validation** - you can require surrounding context ("a word only when followed by a colon") without capturing it. ## A concrete trace Pattern `\d(?=px)` against `5px`: 1. Engine matches `\d` -> matches `5`, cursor now sits after `5`. 2. Engine reaches `(?=px)`. It tries to match `px` from the current position **without moving**. The next chars are `px` -> success. Cursor stays after `5`. 3. Overall match succeeds, `group(0)` is `"5"` only - the `px` was looked at but not consumed. Contrast with `\dpx` (no lookahead): that would match `"5px"` entirely. ## Java specifics Java's `java.util.regex` supports all four lookarounds. You use them inside any `Pattern`. Lookarounds are **not capturing groups** - they never create a `group(n)`. (A normal `(...)` captures; the `(?...)` family are special non-capturing constructs.) One important Java restriction: **lookbehind must be bounded in length** - more on that in the dedicated question.

  • Does a lookahead create a capturing group you can read with group(1)?
    No. Lookarounds are special non-capturing constructs. If you need to capture text you must put a real (...) group inside or around them.
  • Name a zero-width construct other than a lookaround.
    Anchors like ^, $, \b (word boundary), \A, \z - they assert a position without consuming a character.

saying these in an interview costs you the question

  • Thinking the looked-at text becomes part of group(0) or gets replaced
  • Confusing lookaround with a capturing group
  • Believing lookahead moves the cursor forward past the asserted text

context

open as a page

Explain the difference between positive and negative lookahead/lookbehind, with a Java example of each.

level: middleimportance: must knowfreq 58%

basics

~20 s

Positive means "require this pattern here" - (?=X) require X after, (?<=X) require X before. Negative means "require this pattern NOT here" - (?!X) X must not follow, (?<!X) X must not precede. Either way nothing is consumed.

open as a page

What is the bounded-length restriction on lookbehind in Java, and how do you work around an unbounded requirement?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Java's lookbehind must match text of a limited, known maximum length. So (?<=ab) is fine and bounded quantifiers like (?<=a{1,5}) work, but truly unbounded ones like (?<=a+) or (?<=a*) are not allowed because the engine must know how far back to look.

open as a page

Using lookarounds, how would you insert thousands separators into a number string in Java, and why does this work without deleting any digits?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Match the empty positions between digits where exactly a multiple of three digits remain, using (?<=\d)(?=(\d{3})+$), and replace each with a comma. It matches zero-width positions, so no digit is consumed or removed - commas are just inserted.

open as a page

How can lookahead/lookbehind be used with String.split or Pattern.split to keep delimiters, and what is the trade-off versus a normal split?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

split breaks on the matched delimiter and throws it away. If you split on a zero-width lookaround position instead, no characters are consumed, so the delimiter stays attached to a piece. For example split on (?<=,) keeps the comma at the end of each chunk.

open as a page