In Dart, why does RegExp(r'[A-Z]{3}-\d{4}').hasMatch('xxLIS-0042yy') return true, and how do you validate and extract parts of a booking reference?
answer
- RegExp searches, it does not full-match
- anchor with ^ and $
- raw strings keep backslashes
- firstMatch and namedGroup
- JavaScript regex syntax and semantics
basics
~20 shasMatch searches for the pattern anywhere in the input, so an unanchored pattern accepts text with extra characters around it. Anchor it with ^ and $, then use firstMatch and named groups to pull out the parts.
solid answer
~40 sDart's `RegExp` has JavaScript's syntax and semantics, and its common methods search: `hasMatch`, `firstMatch` and `allMatches` look for the pattern starting at any position, so `[A-Z]{3}-\d{4}` finds `LIS-0042` inside `xxLIS-0042yy`. For validation, anchor it: `^...$`. Write patterns as raw strings, `r'...'`, so backslashes reach the regex engine unescaped. To extract parts, call `firstMatch`, which returns `RegExpMatch?`, and read `group(n)` or, with `(?<name>...)` groups, `namedGroup('name')`. The constructor takes `multiLine`, `caseSensitive`, `unicode` and `dotAll`, all defaulting to `false` except `caseSensitive`, which defaults to `true`. An invalid pattern throws a `FormatException`. Create a `RegExp` once, for example in a top-level `final`, rather than on every call.
code
dart · 17 linesfinal _reference = RegExp(r'^(?<hotel>[A-Z]{3})-(?<number>\d{4})$');
({String hotel, int number})? parseReference(String input) {
final match = _reference.firstMatch(input.trim().toUpperCase());
if (match == null) return null;
return (
hotel: match.namedGroup('hotel')!,
number: int.parse(match.namedGroup('number')!),
);
}
void main() {
final loose = RegExp(r'[A-Z]{3}-\d{4}');
print(loose.hasMatch('xxLIS-0042yy')); // true: a search, not a full match
print(parseReference(' lis-0042 ')); // (hotel: LIS, number: 42)
print(parseReference('xxLIS-0042yy')); // null
}go deeper
Know RegExp, hasMatch and firstMatch, and that you should write patterns as raw strings.
Explain search-versus-full-match, anchoring with ^ and $, named groups via namedGroup, and the four constructor flags.
Keep validators anchored and simple, compile them once, escape user text with RegExp.escape, and prefer dedicated parsers for dates and URIs.
Decide where format validation lives, client for feedback and server for truth, so a regex in the app never becomes the only guard.
## Search, not full match The `RegExp` documentation explains that the most common use of a regular expression is to **search** for a match in the input, and that `firstMatch` finds the first position where the pattern matches. `hasMatch`, `allMatches` and `stringMatch` search the same way. Nothing requires the match to cover the whole string. So `RegExp(r'[A-Z]{3}-\d{4}').hasMatch('xxLIS-0042yy')` is `true`: the engine finds `LIS-0042` in the middle. For **validation**, anchor the pattern: - `^` matches the start of the input and `$` the end. - With `multiLine: true`, they also match at the start and end of each line, which is usually not what a validator wants. - The documentation also warns that an unanchored pattern that begins with `.*` can make the search quadratic; anchors avoid that. ## The API you use every day | Member | Returns | |---|---| | `hasMatch(input)` | `bool`: is there a match anywhere? | | `firstMatch(input)` | `RegExpMatch?`: the first match, or `null` | | `allMatches(input, [start])` | `Iterable<RegExpMatch>`: every non-overlapping match | | `stringMatch(input)` | `String?`: the text of the first match | | `match.group(n)` / `match[n]` | text of capture group `n`; group 0 is the whole match | | `match.namedGroup('name')` | text of a `(?<name>...)` group | | `RegExp.escape(text)` | text with regex special characters escaped | Strings also accept a `RegExp` wherever they take a `Pattern`: `contains`, `split`, `replaceAll`, `replaceAllMapped` and `startsWith`. ## Constructor options `RegExp(source, {multiLine = false, caseSensitive = true, unicode = false, dotAll = false})`: 1. **`multiLine`**: `^` and `$` match at line boundaries too. 2. **`caseSensitive`**: set to `false` to ignore letter case. 3. **`unicode`**: treat the pattern as a Unicode pattern per ECMAScript, needed for `\p{...}` classes. 4. **`dotAll`**: `.` also matches line terminators. An invalid pattern, such as an unclosed group, throws a **`FormatException`** when the `RegExp` is constructed. ## Writing patterns in Dart source - Use **raw strings**, `r'\d{4}'`. In a normal string, `'\d'` must be written `'\\d'`, and a stray `$` would start interpolation. - Build a `RegExp` **once**, as a top-level or static `final`, instead of inside a function that runs per keystroke. - To match user-supplied text literally, pass it through `RegExp.escape` first. ## Validating a booking reference Suppose references look like `LIS-0042`: three letters for the property, a dash, four digits. 1. Normalise the input: `trim()` and `toUpperCase()`, or construct with `caseSensitive: false`. 2. Match with an anchored pattern using named groups: `^(?<hotel>[A-Z]{3})-(?<number>\d{4})$`. 3. If `firstMatch` returns `null`, show a validation message. 4. Otherwise read `namedGroup('hotel')` and `namedGroup('number')`, and convert the digits with `int.parse`, which is safe here because the pattern already guaranteed four digits. Returning a record such as `({String hotel, int number})?` keeps the parsed result typed. ## Working with several matches - `allMatches(input)` returns every non-overlapping match lazily, so a `for` loop over it reads each `RegExpMatch` with its groups. - `match.start` and `match.end` give positions in the input, useful for highlighting. - `input.replaceAllMapped(regExp, (m) => ...)` rewrites each match with a function, for example masking all but the last two digits of a reference. - `input.split(RegExp(r'[,;\s]+'))` splits on any run of separators, which cleans a pasted list of references. - `matchAsPrefix(input, start)` tests whether the pattern matches exactly at one position, the primitive the other methods build on. ## Limits worth stating - Regular expressions validate **shape**, not meaning: `XXX-0000` passes the pattern even if no such property exists. - For structured formats such as dates, URIs or JSON, use the dedicated parsers (`DateTime.parse`, `Uri.parse`, `jsonDecode`) rather than a hand-written pattern. - Complex patterns with nested quantifiers can backtrack badly on non-matching input; keep validators simple and anchored.
- Why write regex patterns as raw strings in Dart?In a normal string literal, backslash starts an escape and `$` starts interpolation, so `'\d'` has to be written `'\\d'` and a `$` anchor needs escaping. A raw string, `r'^\d{4}$'`, passes every character to the regex engine unchanged, which keeps patterns readable and correct.
- How do you find every booking reference in a pasted email body?Use an unanchored pattern with `allMatches`, which returns an `Iterable<RegExpMatch>` of non-overlapping matches, and read `group(0)` or named groups from each. Add word boundaries, `\b`, so references glued to other letters are not picked up.
saying these in an interview costs you the question
- hasMatch checks whether the whole string matches the pattern
- Dart regular expressions use Java or PCRE semantics
- An invalid pattern returns false from hasMatch instead of failing
- firstMatch returns an empty match object when nothing matches
- multiLine: true is needed for ^ and $ to work at all