skip to content

Why does an unanchored regular expression accept input that its author meant to reject?

level: middleimportance: must knowfreq 58%

answer

  1. search, not a whole-subject test
  2. a match may begin anywhere
  3. anchors delete candidate matches
  4. zero-width: consumes no characters
  5. `$` and end-of-input can differ

basics

~20 s

Searching asks whether a match exists anywhere in the subject, not whether the whole subject matches. Without a start and end anchor, a pattern for four digits happily matches the four digits buried inside a longer string.

solid answer

~40 s

A pattern is matched against a subject by looking for the leftmost position at which a match exists; unless you say otherwise, everything before and after that span is simply ignored. So a validator built on `[0-9]{4}` accepts `abc1234def`, because the digits are in there. Anchoring fixes it: `^[0-9]{4}$` requires the match to begin at the start of the subject and end at its end, so nothing may hang off either side. Anchors are **zero-width** — `^`, `$` and `\b` assert something about a position and consume no characters, which is why adding them never changes the captured text, only whether the match is allowed at all.

go deeper

for a junior

Know that a pattern finds a match anywhere unless you pin it. Be able to say why a four-digit pattern accepts a string that also contains letters, and how to fix it.

for a middle

Explain anchors as zero-width assertions that remove candidate matches without consuming characters, and show why a start anchor alone still lets trailing junk through.

for a senior

Demonstrate the validator review habit: check both ends, check the empty value, and check what a trailing line terminator does before the pattern guards anything that matters.

for a principal

Treat 'is this a search or a whole-value rule?' as a contract to be stated once for the codebase, since an unanchored validator fails open and that failure mode is invisible in the happy-path tests everyone writes.

## Searching and testing are different questions There are two questions you can ask with a pattern, and they have different answers: - **Does a match exist somewhere in this subject?** This is a search. It tries the leftmost position at which a match can begin, and everything outside the matched span is irrelevant. - **Does this whole subject match?** This is a full-subject test, where the match must cover every character. Most interfaces default to the first question, or offer both under names that are easy to confuse. That is the root of the classic validation bug: a pattern written as a description of a *whole* value is evaluated as a *search* for a fragment. `[0-9]{4}` is a perfectly good description of a four-digit code and a useless validator, because `abc1234def` contains a four-digit run. ## What an anchor actually is An anchor is a **zero-width assertion**: it tests a property of a position between characters and consumes nothing. - `^` asserts the position at the start of the subject (or, in line mode, the start of any line). - `$` asserts the position at the end of the subject (or, in line mode, the end of any line). - `\b` asserts a **word boundary**: one side is a word character and the other is not, counting the string edges as non-word. - Lookahead and lookbehind generalise this — they test whether a sub-pattern matches at the current position without moving it. Because anchors consume nothing, adding them never changes the text a group captures. They only remove candidate matches. `^[0-9]{4}$` has no match at all in `abc1234def`, which is exactly what the validator wanted. ## Why the failure is easy to miss 1. The pattern is tested against well-formed samples, where it matches and the span happens to be the whole value. 2. The first malformed value still matches, because the good part is a substring of it. 3. The bug surfaces later as data that never should have been stored, usually with junk on one side of a value that looks fine. The failure also survives review well, because the pattern *reads* like a description of the whole value. Only the evaluation is a search. ## Word boundaries are not spaces `\bcat\b` matching in `the cat sat` finds `cat`, and the spaces are **not** part of the match — the boundary asserts a position between a non-word character and a word character. That is why `\bcat\b` also matches in `(cat)` and at the very start or end of the subject, where there is no space at all. Candidates who think `\b` consumes a space are surprised when a replacement removes the surrounding punctuation, or when the match offsets are one character off from what they expected. ## The two anchors are not symmetric everywhere This is a place where conventions genuinely differ, and the difference bites validators: - Some conventions make `$` match not only at the very end of the subject but also immediately before a trailing line terminator, so a value with a newline glued to the end still passes `^[0-9]{4}$`. - Some offer a separate absolute end-of-input anchor with no such concession. - Line mode — where `^` and `$` match at every internal line boundary — is opt-in in some conventions and configured elsewhere in others. The portable habit is to state which behaviour you rely on rather than to assume it: if a trailing line terminator must be rejected, use an absolute end-of-input assertion where the convention offers one, or reject line terminators explicitly inside the pattern. ## A checklist for a pattern used as a validator 1. Decide whether you are searching or testing the whole value, and make the pattern say so rather than relying on the call. 2. Anchor both ends; one anchor alone still lets junk hang off the other side. 3. Ask what a line terminator inside or at the end of the value would do. 4. Check the empty subject: `^[0-9]*$` accepts it, `^[0-9]+$` does not, and that difference is usually the real requirement. | pattern | `1234` | `abc1234def` | `1234\n` | |---|---|---|---| | `[0-9]{4}` | accepted | accepted | accepted | | `^[0-9]{4}` | accepted | rejected | accepted | | `^[0-9]{4}$` | accepted | rejected | convention-dependent | The third column is the one that turns a description into a validator; the fourth is the one worth knowing about before it becomes a defect.

  • What does a word boundary assert, given that it matches no characters?
    It asserts that of the two sides of the current position, exactly one is a word character — the string edges count as non-word. So it holds between a space and a letter, between punctuation and a letter, and at the start of a subject beginning with a letter, without consuming anything.
  • An anchored validator still accepts a two-line payload. What happened?
    Line-anchor mode is in effect, so `^` and `$` also assert at internal line boundaries and the match only has to cover one line. Either turn that mode off, use an absolute end-of-input assertion where the convention provides one, or exclude line terminators inside the pattern itself.
  • Is anchoring only a start anchor enough?
    No. `^[0-9]{4}` still accepts `1234abc`, because the match ends where the pattern runs out and the remaining characters are ignored. A whole-value rule needs both ends pinned.

saying these in an interview costs you the question

  • Thinks a pattern must cover the entire subject by default
  • Believes `^` consumes the first character of the subject
  • Wraps the pattern in unrestricted runs believing that anchors it
  • Assumes a word boundary consumes the space between two words
  • Anchors only the start and calls the value validated
  • Treats `$` as always meaning the final character of the input