Given the JavaScript string "<a><b>", what does the regex /<.+>/ match and what does /<.+?>/ match, and what rule explains the difference?
answer
- default appetite of a quantifier
- one form takes all, then gives back
- an extra symbol flips the preference
- leftmost start still wins
- a negated class avoids the whole dance
basics
~20 s/<.+>/ matches the whole "<a><b>" because + is greedy and takes as much as it can. /<.+?>/ matches only "<a>" because the ? after the quantifier makes it lazy, taking the fewest characters that still let the rest of the pattern succeed.
solid answer
~40 sQuantifiers in JavaScript are greedy by default: `+` consumes as many characters as possible, then hands control back and backtracks only as far as it must. On `"<a><b>"`, `.+` first swallows `a><b>`, then gives back one character at a time until a `>` can match — and the last `>` works, so the match is the whole string. Appending `?` to the quantifier makes it lazy: `.+?` takes the minimum, one character, and expands only when the rest of the pattern fails, so the first `>` ends the match at `"<a>"`. Laziness changes how much the quantifier eats, not where matching starts — the leftmost starting position still wins. In practice a negated class like `/<[^>]+>/` expresses the same intent more directly and does far less backtracking.
code
javascript · 9 linesconst s = "<a><b>";
console.log(s.match(/<.+>/)[0]); // "<a><b>" greedy
console.log(s.match(/<.+?>/)[0]); // "<a>" lazy
console.log(s.match(/<[^>]+>/)[0]); // "<a>" negated class
// Laziness changes the length, not the starting position:
console.log("xaybzb".match(/a.*b/)[0]); // "aybzb"
console.log("xaybzb".match(/a.*?b/)[0]); // "ayb"go deeper
Know that quantifiers grab as much as they can by default, and that adding ? right after one flips it to taking as little as possible. Be able to predict the two matches on "<a><b>".
Trace the backtracking step by step: greedy consumes to the end then gives characters back, lazy starts minimal then expands. Say plainly that laziness affects length, not the starting position.
Show that you reach for a negated character class instead, and explain that JavaScript has no possessive quantifiers or atomic groups, so every quantifier remains backtrackable and ambiguity has a real cost on hostile input.
Frame it as a maintainability and risk call: patterns whose correctness depends on backtracking order are hard for reviewers to verify, so set the expectation that shared patterns are unambiguous, tested against long inputs, and replaced by a parser when the grammar is genuinely nested.
## Greedy is the default JavaScript's quantifiers — `*` (zero or more), `+` (one or more), `?` (zero or one) and the counted forms `{n}`, `{n,}`, `{n,m}` — are **greedy**. A greedy quantifier consumes as much input as it possibly can before letting the rest of the pattern try to match, then gives characters back one at a time only when forced. Walk `/<.+>/` over `"<a><b>"`: 1. `<` matches the `<` at index 0. 2. `.+` is greedy, so it consumes everything it can: `a><b>` — to the end of the string. 3. The pattern still needs `>`, but there is no input left. Failure. 4. The engine **backtracks**: `.+` gives back one character, now holding `a><b`, leaving `>` as the next input character. 5. `>` matches. The overall match is `"<a><b>"` — the whole string. The result surprises people who read `<.+>` as "a tag". It is really "a `<`, then as much as possible, then the *last* `>` that still works". ## Lazy quantifiers Putting `?` immediately after a quantifier makes it **lazy** (also called non-greedy or reluctant): `*?`, `+?`, `??`, `{n,m}?`. A lazy quantifier consumes the minimum allowed, then expands one character at a time only when the remainder of the pattern fails. Walk `/<.+?>/` over the same input: 1. `<` matches at index 0. 2. `.+?` takes the minimum for `+`, which is one character: `a`. 3. The pattern needs `>`; the next character *is* `>`. It matches. 4. Match complete: `"<a>"`. Note that `?` is doing double duty in the syntax. On its own it is the quantifier "zero or one". Directly after another quantifier it is the laziness modifier. So `a??` means "zero or one `a`, preferring zero". ## What laziness does not do The single most common misconception is that lazy means "find the shortest match in the string". It does not. Match position is decided first: the engine tries to match starting at index 0, then index 1, and so on, and the **leftmost** successful start wins regardless of length. Laziness only influences how much a quantifier consumes once a start position is being tried. A second misconception is that lazy quantifiers never backtrack. They backtrack in the opposite direction — expanding rather than contracting — and can do just as much work. `/^.*?x$/` on a long string with no `x` still explores every expansion before failing. A third: laziness is not automatically faster. On text where the target is near the start, lazy wins. On text where the intended match runs to the end, greedy wins. Neither is a performance rule. ## The better tool most of the time When you write `<.+?>` you are usually expressing "characters up to but not including the delimiter". A **negated character class** says exactly that, without relying on backtracking at all: ```js "<a><b>".match(/<[^>]+>/)[0]; // "<a>" ``` `[^>]` matches any character that is not `>`, so the quantifier physically cannot run past the delimiter and never has to give anything back. The class version is clearer about intent and does far less work, which also makes it much harder to write a pattern that degrades badly on adversarial input. The equivalence breaks down only for multi-character delimiters, where a lazy quantifier really is the natural spelling. Also worth knowing: JavaScript has **no possessive quantifiers** (`a++`) and **no atomic groups**, the two features other flavours offer for saying "take this and never give it back". Every quantifier in a JavaScript regex is backtrackable, which is exactly why a negated class is the tool of choice for cutting off runaway matching. ## A worked contrast to keep ```js const s = 'name="x" id="y"'; s.match(/".*"/)[0]; // '"x" id="y"' - greedy, spans both quotes s.match(/".*?"/)[0]; // '"x"' - lazy, stops at the first close s.match(/"[^"]*"/)[0];// '"x"' - class, same result, no backtracking ``` Being able to produce that three-way comparison on demand — and to say why the third line is the one you would ship — is what a strong answer looks like.
- Does adding ? to a quantifier ever change where a match starts, not just how long it is?No. The engine picks the leftmost position at which the whole pattern can succeed, and only then does the quantifier's greediness decide how much it consumes there. That is why `/a.*?b/` and `/a.*b/` on `"xaybzb"` both start at the `a`; they differ only in whether the match ends at the first or last `b`.
- Why is /<[^>]+>/ usually preferable to /<.+?>/ for the same job?The negated class cannot cross the delimiter at all, so the engine never consumes past `>` and never has to backtrack to give it up. It states the intent directly — characters other than `>` — and it degrades far better on long or hostile input. The lazy form is still the natural choice for a multi-character delimiter.
- Are lazy quantifiers a performance optimisation?No — they change which match you get, not how much work the engine does. Lazy is faster when the target sits near the start and slower when the match genuinely extends to the end, because it expands one character at a time. Treat the choice as semantic; if you want a real speed guarantee, remove the ambiguity with a character class.
saying these in an interview costs you the question
- Says lazy means the shortest match anywhere in the string
- Expects /<.+>/ to match a single tag
- Believes lazy quantifiers never backtrack
- Claims lazy is always faster than greedy
- Thinks JavaScript supports possessive quantifiers or atomic groups