skip to content

Explain lookahead and lookbehind assertions in JavaScript regular expressions — what does "zero-width" mean, and give a case where a lookahead is the only clean solution.

level: middleimportance: must knowfreq 66%

answer

  1. it checks, then rewinds the cursor
  2. examined text stays out of the match
  3. four forms: ahead/behind, positive/negative
  4. several conditions at the same position
  5. lookbehind matches right to left

basics

~20 s

Lookaround tests whether a sub-pattern matches at the current position without consuming characters, so the matched text excludes it. (?=x) and (?!x) look forward, (?<=x) and (?<!x) look backward, and each can be positive or negative.

solid answer

~40 s

A lookaround is an assertion: the engine tries the inner pattern at the current position and then rewinds, keeping only the yes/no answer. Nothing it inspected becomes part of the match, which is what "zero-width" means. There are four forms: `(?=…)` positive lookahead, `(?!…)` negative lookahead, `(?<=…)` positive lookbehind, `(?<!…)` negative lookbehind. The classic use is a multi-condition rule at one position — a password check like `/^(?=.*[a-z])(?=.*[A-Z])(?=.*\d)\S{8,}$/` stacks three independent lookaheads that each scan from the start, which no single linear pattern can express. The other everyday use is a context filter: `/\d+(?= USD)/` matches the number only when the unit follows, but returns just the digits. Lookahead is original ECMAScript; lookbehind landed in ES2018.

code

javascript · 3 lines
javascript
const strong = /^(?=.*[a-z])(?=.*[A-Z])(?=.*\d)\S{8,}$/;
console.log(strong.test('Passw0rdX')); // true
console.log(strong.test('password1')); // false

go deeper

for a junior

Recognise the four lookaround forms and say plainly that they test what is around the cursor without adding those characters to the match.

for a middle

Explain zero-width mechanically — the engine matches the inner pattern and rewinds — and produce the stacked-lookahead or context-filter example unprompted.

for a senior

Show judgement about cost and readability: each assertion is a sub-match that can rescan, and a dense multi-lookahead validator is often worse for users than several small explicit checks.

for a principal

Own the policy question — where declarative pattern matching stops paying for itself, what a validation rule should report back to a user, and how to keep a shared pattern library reviewable.

## The idea: assert, do not consume Everything else in a regex advances a cursor. `a` matches an `a` and moves past it; `\d{3}` consumes three digits. A lookaround does not. The engine runs the inner pattern starting at the current position, records whether it succeeded, then puts the cursor back exactly where it was. That is what **zero-width** means, and it is why `\b`, `^` and `$` are in the same family — they are assertions about position rather than pieces of the match. The practical consequence: characters examined inside a lookaround are not part of `m[0]` and remain available to be matched again by whatever comes next. ## The four forms - `(?=x)` — positive lookahead: succeeds if `x` matches starting here. - `(?!x)` — negative lookahead: succeeds if `x` does **not** match starting here. - `(?<=x)` — positive lookbehind: succeeds if `x` matches ending here. - `(?<!x)` — negative lookbehind: succeeds if `x` does not match ending here. Lookahead has been in the language since the beginning. Lookbehind was added in ES2018; browser support is now universal, though Safari only shipped it in 16.4, which is why older code often works around its absence. ## Case one: stacking independent conditions The strongest argument for lookahead is that it lets you assert several unrelated things about the same string at the same position. ```js const ok = /^(?=.*[a-z])(?=.*[A-Z])(?=.*\d)\S{8,}$/; ok.test('Passw0rdX'); // true ok.test('password1'); // false — no uppercase ``` Each `(?=.*X)` starts at position 0, scans forward looking for `X`, then rewinds to 0. Three assertions, three independent scans, all anchored at the same spot; only after all three succeed does `\S{8,}$` actually consume the string. Without lookahead you would have to enumerate every interleaving of the required character classes, which is combinatorially hopeless. (In production, expressing each rule as its own small test is usually clearer and gives better error messages than one dense pattern — but the interviewer is asking whether you know the mechanism.) ## Case two: context without capture The other everyday use is matching something only in a context you do not want in the result. ```js '42 USD'.match(/\d+(?= USD)/)[0]; // '42' '42 EUR'.match(/\d+(?= USD)/); // null '$1250'.match(/(?<=\$)\d+/)[0]; // '1250' ``` A capturing group could give the same value, but then the consumer has to reach for `m[1]` and the overall match spans text it does not care about. With lookaround, `m[0]` is exactly the payload. Negative forms filter: `/(?<!\$)\b\d+\b/` finds numbers not preceded by a dollar sign. ## Captures inside a lookaround Groups inside a lookaround still capture, and their captures survive if the assertion succeeded — `/(?=(\d+))/` is the standard trick for capturing text without consuming it. Inside a **negative** lookaround the situation is different: the assertion only succeeds when the inner pattern failed, so groups nested in `(?!…)` never end up with a value. JavaScript's lookbehind is unusual in a good way: it is **variable-length**. Many engines require a fixed-width lookbehind; JavaScript allows quantifiers, so `(?<=\w+\s)` is legal. The engine implements this by matching the inner pattern **right to left** from the current position. That direction is observable when a lookbehind contains two greedy quantified groups: `/(?<=(\d+)(\d+))$/.exec('1053')` gives group 1 `'1'` and group 2 `'053'`, because the rightmost group is matched first and takes greedily. Reading it left to right gives the opposite intuition. ## Costs and misconceptions A lookaround is not free — it runs a sub-match — and one placed inside a quantified section can be re-run at every position. The stacked-password pattern scans the input once per assertion, which is fine for a field and unwise for a megabyte. Two misconceptions to head off. First, a lookahead is not a capture: `/(?=\d)/` produces an empty match, not the digit; you still need to consume or capture it. Second, a zero-width assertion cannot be usefully quantified — `(?=a)+` asserts the same thing over and over at the same position and adds nothing. Also beware `(?<!)` versus a plain negated class: `/(?<!a)b/` matches a `b` at the very start of the string, because there is no preceding `a` there, while `/[^a]b/` requires some character to be present. Assertions succeed at boundaries where a consuming pattern would fail.

  • Why does the stacked password pattern need lookaheads rather than one linear pattern?
    Because the required character classes can appear in any order and interleaved with anything else. A linear pattern would have to enumerate every ordering. Each `(?=.*X)` instead rewinds to position 0 and scans independently, so the conditions compose without interacting — then a single consuming pattern checks length.
  • Do capture groups inside a negative lookahead ever hold a value?
    No. A negative lookahead only succeeds when its inner pattern fails to match, and a failed match sets no captures. So a group nested inside `(?!…)` is always `undefined` in the result. Groups inside a *positive* lookaround do keep their captures, which is the usual capture-without-consuming trick.
  • What is the difference between /(?<!a)b/ and /[^a]b/?
    The lookbehind is zero-width, so it succeeds when no character precedes at all — `/(?<!a)b/` matches the `b` in `'b'`. `[^a]` must consume some character that is not `a`, so it fails at the start of the string and also makes that character part of the match.

saying these in an interview costs you the question

  • Thinks a lookahead consumes or captures the text it inspects
  • Believes JavaScript lookbehind must be fixed-length
  • Says (?!x) means "match any character except x"
  • Cannot state that lookarounds can stack at one position
  • Assumes lookaround is free and never rescans the input

context