Explain the difference between positive and negative lookahead/lookbehind, with a Java example of each.
answer
- = means must be present; ! means must be absent
- < points backward => lookbehind
- (?=X) after-yes, (?!X) after-no, (?<=X) before-yes, (?<!X) before-no
- Negative = the 'unless' tool
- Password rules stack multiple positive/negative lookaheads
basics
~20 sPositive means "require this pattern here" - (?=X) require X after, (?<=X) require X before. Negative means "require this pattern NOT here" - (?!X) X must not follow, (?<!X) X must not precede. Either way nothing is consumed.
solid answer
~40 sAll four are zero-width context tests at the current position. Positive lookahead (?=X) succeeds when X matches just ahead; negative lookahead (?!X) succeeds when X does not match just ahead. Positive lookbehind (?<=X) succeeds when X matches immediately behind the cursor; negative lookbehind (?<!X) succeeds when X does not. Positive forms assert presence of context, negative forms assert its absence. A classic use is matching a number only if it is preceded by a currency symbol: (?<=\$)\d+ matches the digits in "$50" but yields group(0) = "50". Negative lookahead is the standard tool for "match X unless followed by Y", e.g. \bcat\b(?! food) matches "cat" but not "cat" in "cat food". Because all are zero-width, the surrounding symbol/word is tested but never captured or replaced.
code
java · 14 linesimport java.util.regex.*;
// positive lookbehind: amount after $
Matcher a = Pattern.compile("(?<=\\$)\\d+(?:\\.\\d+)?").matcher("price $19.99");
a.find(); System.out.println(a.group()); // 19.99
// negative lookahead: 'red' not before ' apple'
Matcher b = Pattern.compile("\\bred\\b(?! apple)").matcher("red car, red apple");
int n = 0; while (b.find()) n++;
System.out.println(n); // 1 (only 'red car')
// stacked lookaheads for validation
boolean ok = "abc12345".matches("^(?=.*\\d)(?=.*[a-z]).{8,}$");
System.out.println(ok); // truego deeper
Knows the four symbols and that = is presence, ! is absence.
Can choose positive vs negative correctly for 'only if' vs 'unless' requirements and write currency/password patterns.
Understands backtracking pitfalls of negative lookahead and how lookarounds stack at one position for validation.
Sets guidance on when stacked-lookahead validation is appropriate versus explicit checks, considering readability and maintenance.
## Recap: zero-width A lookaround tests a sub-pattern at the cursor position and **consumes nothing**. The only question each kind answers is *direction* (ahead vs behind) and *polarity* (must match vs must not match). ## The 2x2 grid | | Lookahead (text after) | Lookbehind (text before) | |---|---|---| | **Positive** (must match) | `(?=X)` | `(?<=X)` | | **Negative** (must NOT match) | `(?!X)` | `(?<!X)` | Mnemonic: `=` means "equals/yes, must be there", `!` means "not, must be absent". `<` points backward, so it is a *behind* construct. ## Positive lookahead `(?=X)` Succeeds only if `X` would match starting at the cursor. Use for "only if followed by". Example - a digit only when `%` follows: `\d+(?=%)` on `"30% off"` -> matches `"30"`, the `%` is asserted but not consumed. ## Negative lookahead `(?!X)` Succeeds only if `X` would **fail** at the cursor. Use for "unless followed by". Example - `\d+(?!%)` is subtle because of backtracking, so the common pattern is whole-word exclusion: `\bfoo\b(?!bar)` matches `foo` not followed by `bar`. The famous use is a password rule: `(?=.*\d)` (must contain a digit anywhere) combined with `(?!.*\s)` (must contain no whitespace). ## Positive lookbehind `(?<=X)` Succeeds only if `X` matches the text **ending exactly at** the cursor. Use for "only if preceded by". Example - the amount after a dollar sign: `(?<=\$)\d+(\.\d+)?` on `"$19.99"` -> matches `"19.99"`, the `$` is not part of the match. ## Negative lookbehind `(?<!X)` Succeeds only if `X` does **not** end at the cursor. Use for "unless preceded by". Example - match `cat` not preceded by `bob`: `(?<!bob)cat`. ## Why polarity matters for replacement Negative lookarounds are powerful guards in `replaceAll`. To thousand-separate digits, the pattern `(?<=\d)(?=(\d{3})+$)` matches an empty position that has a digit behind it and a multiple-of-three digit run ahead - it consumes nothing, so replacing with `,` inserts commas between digits without deleting any. ## Java notes - All four are supported by `java.util.regex`. - They never capture (no `group(n)` for the lookaround itself; a `(...)` inside one still captures). - **Lookbehind (both positive and negative) must be bounded in length** in Java - covered in its own question. - A negative lookbehind that contains capturing groups: those groups are unset if the lookbehind succeeds (because it didn't match anything).
- How would you require a string to contain at least one digit and one uppercase letter using lookaheads?Anchor at start and stack positive lookaheads: ^(?=.*\d)(?=.*[A-Z]).+$ - each lookahead scans the whole string from the start position without consuming, then .+ does the actual match.
- Why is \d+(?!%) tricky for 'a number not followed by percent'?Backtracking: \d+ can give back a digit so the lookahead sees another digit (not %) and still succeeds. You usually need a boundary like \b or to assert the full run, e.g. \d+\b(?!%).
saying these in an interview costs you the question
- Mixing up the arrow: writing (?=X) when you meant behind
- Assuming negative lookahead consumes the rejected text
- Thinking (?<!X) captures or removes the preceding text
- Forgetting that combined lookaheads all test the SAME position