What do the predefined classes \d, \w, \s mean in Java regex, how do they relate to POSIX classes like \p{Alpha}, and what are their ASCII vs Unicode caveats?
answer
- \d digit, \w word(letters+digits+_), \s whitespace
- Uppercase = negation: \D \W \S
- Shorthands work inside [ ] to union sets
- \p{Alpha}/\p{Digit}... = POSIX names
- Default ASCII; UNICODE_CHARACTER_CLASS for Unicode
basics
~10 s\d is a digit, \w is a word character (letters, digits, underscore), \s is whitespace. Their uppercase forms (\D, \W, \S) are the negations. By default they match ASCII only unless you enable Unicode.
solid answer
~40 sJava's predefined shorthands are \d (a digit 0-9), \w (a word char: letters, digits, or underscore), and \s (whitespace: space, tab, newline, etc.). The uppercase variants \D, \W, \S negate them. They are themselves character classes, so you can combine them inside brackets, e.g. [\w.] for word chars or a dot. Java also offers POSIX classes via \p{...}, such as \p{Alpha}, \p{Digit}, \p{Alnum}, \p{Punct} — by default these and the shorthands are ASCII-only. To make \w, \d, \s and POSIX classes Unicode-aware you compile with the UNICODE_CHARACTER_CLASS flag (or the (?U) inline flag), after which \d matches any Unicode decimal digit and \w any Unicode word character. Negations like [^\d] and the uppercase \D both exclude digits but differ subtly in combinations.
go deeper
Recognizes \d, \w, \s and their meanings and that uppercase forms negate.
Combines shorthands inside classes, knows POSIX \p{...} names, and remembers the ASCII default.
Enables UNICODE_CHARACTER_CLASS for international input, distinguishes \D from [^\d] in combinations, and avoids over-broad \w in validators.
Reasons about correctness/security of input validation across locales, performance of Unicode classes, and consistent regex conventions across a codebase.
## The shorthand classes Writing common sets out by hand is tedious, so regex provides **predefined (shorthand) character classes**. In Java the core three are: - `\d` — a **digit**. By default equivalent to `[0-9]`. - `\w` — a **word character**. By default `[a-zA-Z0-9_]` (letters, digits, underscore). - `\s` — a **whitespace** character: space, tab `\t`, newline `\n`, carriage return `\r`, form feed `\f`, vertical tab `\x0B`. Each has an **uppercase negation**: - `\D` = not a digit, `\W` = not a word char, `\S` = not whitespace. These shorthands are themselves classes matching exactly one character, and they can appear **inside** a bracketed class to be unioned with other members: `[\w.]` matches a word char or a literal dot; `[\d\s]` matches a digit or whitespace. ## POSIX classes Java also supports **POSIX-style** named classes through the `\p{Name}` syntax (POSIX is an old Unix standard whose names Java borrows): - `\p{Digit}` = `[0-9]`, `\p{Alpha}` = `[a-zA-Z]`, `\p{Alnum}` = `[a-zA-Z0-9]`, `\p{Upper}`, `\p{Lower}`, `\p{Space}`, `\p{Punct}`, `\p{Blank}` (space/tab), `\p{XDigit}` (hex digits), etc. - Capitalize-negate is not used here; to negate use `\P{...}` or wrap in a negated class `[^\p{Alpha}]`. ## ASCII vs Unicode — the key caveat By default, both the shorthands and the POSIX classes are **ASCII-only**. So `\d` does NOT match the Arabic-Indic digit `٠` or a fullwidth digit, and `\w` does not match accented letters like `é` by default. This surprises people processing international text. To opt in to Unicode semantics, compile with the flag `Pattern.UNICODE_CHARACTER_CLASS` (or prepend the inline flag `(?U)`): ```java Pattern.compile("\\w+", Pattern.UNICODE_CHARACTER_CLASS); ``` Now `\d` matches any Unicode decimal digit, `\w` any Unicode word character, `\s` any Unicode whitespace, and the POSIX classes follow Unicode definitions. Note this is **separate** from `UNICODE_CASE` (which only affects case-insensitive matching). ## Unicode property classes vs POSIX The `\p{...}` syntax also reaches **true Unicode properties** (covered more in its own area): `\p{L}` is "any letter" across all scripts, `\p{Lu}` uppercase letters, `\p{N}` numbers. These are always Unicode-defined regardless of the flag, unlike the POSIX names which are ASCII until UNICODE_CHARACTER_CLASS is set. ## Negation nuances `\D` and `[^\d]` both mean "not a digit" on their own. The difference appears in **combination**: inside a single class `[^\d\s]` means "neither digit nor whitespace", whereas naively chaining `\D\S` would match two characters. Prefer a single negated class to combine exclusions. ## Java source escaping Every regex backslash doubles in a Java string literal: the regex `\d` is `"\\d"`, and `\p{Alpha}` is `"\\p{Alpha}"`. Forgetting this is a frequent compile-time or behavior bug. ## Deriving the answer Hold three facts: (1) `\d\w\s` are shorthands with uppercase negations, usable inside `[...]`; (2) `\p{...}` gives POSIX names; (3) all default to ASCII and need `UNICODE_CHARACTER_CLASS` for Unicode. From those you can answer any question about coverage and international input.
- What is the difference between UNICODE_CHARACTER_CLASS and UNICODE_CASE?UNICODE_CHARACTER_CLASS makes \d, \w, \s and POSIX classes use Unicode definitions (membership). UNICODE_CASE only affects how case-insensitive matching folds case; it does not change which characters \w or \d match.
- How would you match a single non-word character?Use \W (the negation of \w), or equivalently the negated class [^\w]. Both match one character that is not a letter, digit, or underscore.
saying these in an interview costs you the question
- Assuming \d matches all Unicode digits by default (it is ASCII-only)
- Thinking \w includes accented letters without the Unicode flag
- Confusing UNICODE_CHARACTER_CLASS with UNICODE_CASE
- Believing \p{...} negates by uppercasing (use \P{...})