Anchors are zero-width assertions. How do they relate to lookahead/lookbehind, and how would you build a custom positional assertion in Java?
answer
- Anchors = hard-coded zero-width assertions
- (?=) (?!) lookahead; (?<=) (?<!) lookbehind
- \b ≈ (?<=\w)(?!\w)|(?<!\w)(?=\w)
- Java lookbehind must be bounded length
- Lookaround can cause ReDoS; anchors can't
basics
~20 sAnchors like ^, $, and \b are built-in zero-width assertions that test a fixed position. Lookahead (?=...) and lookbehind (?<=...) are general zero-width assertions you define yourself, letting you build custom 'boundaries' that the standard anchors can't express.
solid answer
~50 sAll anchors — ^, $, \A, \z, \Z, \b, \B — are zero-width assertions: they test a condition at a position and consume nothing. Lookaround generalizes this. A positive lookahead (?=X) asserts that X can match starting here without consuming it; negative lookahead (?!X) asserts it cannot. Positive lookbehind (?<=X) asserts X immediately precedes this position; negative lookbehind (?<!X) asserts it does not. With these you can express custom boundaries the built-ins lack — for example a 'boundary' between a digit and a letter via (?<=\d)(?=\p{Alpha}), or a thousands separator position via (?<=\d)(?=(\d{3})+$). You can even emulate \b as (?:(?<=\w)(?!\w)|(?<!\w)(?=\w)). In Java, lookbehind is effectively bounded (the engine requires a finite maximum length), which is the main practical limitation versus lookahead. Lookaround is the principled tool when the fixed anchors don't capture the position you need.
code
java · 16 linesimport java.util.regex.*;
// Custom boundary: insert commas as thousands separators (zero-width lookahead)
String grouped = "1234567".replaceAll("\\B(?=(\\d{3})+(?!\\d))", ",");
System.out.println(grouped); // 1,234,567
// digit->letter boundary that no standard anchor provides
Matcher m = Pattern.compile("(?<=\\d)(?=\\p{Alpha})").matcher("9a");
System.out.println(m.find() + " @" + (m.find(0) ? m.start() : -1)); // true @1
// Emulating \b with lookaround (educational, not for production)
String b = "(?:(?<=\\w)(?!\\w)|(?<!\\w)(?=\\w))";
System.out.println(Pattern.compile(b + "cat" + b).matcher("the cat").find()); // true
// Java rejects unbounded lookbehind:
// Pattern.compile("(?<=\\d+)x"); // throws PatternSyntaxExceptiongo deeper
Recognizes that anchors and lookaround both match positions, not characters.
Knows the four lookaround forms and that they are zero-width like anchors.
Builds custom boundaries with lookaround and knows Java's bounded-lookbehind limit.
Frames anchors as special-case assertions, weighs lookaround expressiveness against ReDoS/backtracking risk on untrusted input, and sets conventions for safe regex (bounded lookbehind, linear-time engines, timeouts).
## Recap: zero-width assertions A **zero-width assertion** matches a **position**, not characters — it advances the cursor by zero. The built-in **anchors** are special cases: `^`/`$` (line edges), `\A`/`\z`/`\Z` (input edges), `\b`/`\B` (word boundaries). Each tests a **fixed, predefined** condition about the position. ## Lookaround: user-defined assertions **Lookaround** lets you write your **own** zero-width condition using a sub-pattern. There are four forms: - **Positive lookahead** `(?=X)`: succeeds if `X` matches starting at the current position. Nothing is consumed. - **Negative lookahead** `(?!X)`: succeeds if `X` does **not** match here. - **Positive lookbehind** `(?<=X)`: succeeds if `X` matches the text **immediately before** the current position. - **Negative lookbehind** `(?<!X)`: succeeds if `X` does **not** match immediately before. Because they consume nothing, you can stack them. This is the key insight: **anchors are just hard-coded lookarounds**. For instance, a word boundary `\b` is conceptually: ``` (?:(?<=\w)(?!\w) | (?<!\w)(?=\w)) ``` 'either a word char behind and no word char ahead, or no word char behind and a word char ahead.' ## Building custom boundaries The standard anchors only know about line edges, input edges, and the ASCII word class. Lookaround lets you define boundaries they can't express: - **digit→letter boundary:** `(?<=\d)(?=\p{Alpha})` matches the empty position between `9` and `a` in `9a`. - **camelCase split point:** `(?<=\p{Lower})(?=\p{Upper})` finds the seam in `fooBar`. - **thousands separators:** `\B(?=(\d{3})+(?!\d))` marks positions to insert commas, e.g. for `1234567` → `1,234,567` via `replaceAll(..., ",")`. These are all zero-width: they identify positions without eating characters, exactly like an anchor — but on **your** terms. ## Java-specific constraints and tradeoffs 1. **Bounded lookbehind.** Java's engine requires lookbehind to match a **finite maximum length**. `(?<=\d{3})` is fine; `(?<=\d+)` is rejected/limited because it is unbounded. Lookahead has no such restriction. This is the most common practical surprise when porting patterns from engines with unbounded lookbehind. 2. **Performance / catastrophic backtracking.** Lookaround sub-patterns participate in backtracking. A poorly written lookahead containing nested quantifiers can trigger **catastrophic backtracking** (exponential time) — a denial-of-service vector (ReDoS) on attacker-controlled input. Anchors, being constant-time positional checks, never have this risk. At principal level you weigh 'expressive lookaround' against 'cheap, safe built-in anchor', and consider timeouts or non-backtracking engines (e.g. RE2/J) for untrusted input. 3. **Readability and maintenance.** Emulating `\b` by hand is a great learning exercise but a poor production choice; prefer the built-in unless you genuinely need a non-standard boundary. Reserve custom lookaround boundaries for cases the anchors can't express. ## When to choose what - Standard position (line/input edge, ASCII word boundary): **use the anchor**. - Non-standard position (between character classes, format-driven seams): **use lookaround**. - Untrusted input + complex assertions: **mind ReDoS** — keep assertions simple, bound lookbehind, consider a linear-time engine or a matcher timeout.
- How would you express a boundary between a digit and a letter, which no standard anchor provides?Use lookaround: (?<=\d)(?=\p{Alpha}) asserts a digit immediately precedes and a letter immediately follows the current zero-width position.
- What is the main Java-specific restriction on lookbehind, and why does it matter?Java lookbehind must match a bounded (finite maximum) length, so patterns like (?<=\d+) are not allowed. It matters when porting patterns from engines that support unbounded lookbehind, and pushes you to rewrite the assertion with a fixed bound or restructure the regex.
saying these in an interview costs you the question
- Believing lookaround consumes characters (it is zero-width like anchors)
- Using unbounded lookbehind like (?<=\d+) in Java (must be finite length)
- Hand-rolling \b in production instead of the built-in without need
- Ignoring catastrophic backtracking / ReDoS risk on untrusted input