skip to content

What do the predefined classes \d, \w, \s mean in Java regex, how do they relate to POSIX classes like \p{Alpha}, and what are their ASCII vs Unicode caveats?

level: middleimportance: should knowfreq 68%

answer

  1. \d digit, \w word(letters+digits+_), \s whitespace
  2. Uppercase = negation: \D \W \S
  3. Shorthands work inside [ ] to union sets
  4. \p{Alpha}/\p{Digit}... = POSIX names
  5. Default ASCII; UNICODE_CHARACTER_CLASS for Unicode

basics

~10 s

\d is a digit, \w is a word character (letters, digits, underscore), \s is whitespace. Their uppercase forms (\D, \W, \S) are the negations. By default they match ASCII only unless you enable Unicode.

solid answer

~40 s

Java's predefined shorthands are \d (a digit 0-9), \w (a word char: letters, digits, or underscore), and \s (whitespace: space, tab, newline, etc.). The uppercase variants \D, \W, \S negate them. They are themselves character classes, so you can combine them inside brackets, e.g. [\w.] for word chars or a dot. Java also offers POSIX classes via \p{...}, such as \p{Alpha}, \p{Digit}, \p{Alnum}, \p{Punct} — by default these and the shorthands are ASCII-only. To make \w, \d, \s and POSIX classes Unicode-aware you compile with the UNICODE_CHARACTER_CLASS flag (or the (?U) inline flag), after which \d matches any Unicode decimal digit and \w any Unicode word character. Negations like [^\d] and the uppercase \D both exclude digits but differ subtly in combinations.

go deeper

for a junior

Recognizes \d, \w, \s and their meanings and that uppercase forms negate.

for a middle

Combines shorthands inside classes, knows POSIX \p{...} names, and remembers the ASCII default.

for a senior

Enables UNICODE_CHARACTER_CLASS for international input, distinguishes \D from [^\d] in combinations, and avoids over-broad \w in validators.

for a principal

Reasons about correctness/security of input validation across locales, performance of Unicode classes, and consistent regex conventions across a codebase.

## The shorthand classes Writing common sets out by hand is tedious, so regex provides **predefined (shorthand) character classes**. In Java the core three are: - `\d` — a **digit**. By default equivalent to `[0-9]`. - `\w` — a **word character**. By default `[a-zA-Z0-9_]` (letters, digits, underscore). - `\s` — a **whitespace** character: space, tab `\t`, newline `\n`, carriage return `\r`, form feed `\f`, vertical tab `\x0B`. Each has an **uppercase negation**: - `\D` = not a digit, `\W` = not a word char, `\S` = not whitespace. These shorthands are themselves classes matching exactly one character, and they can appear **inside** a bracketed class to be unioned with other members: `[\w.]` matches a word char or a literal dot; `[\d\s]` matches a digit or whitespace. ## POSIX classes Java also supports **POSIX-style** named classes through the `\p{Name}` syntax (POSIX is an old Unix standard whose names Java borrows): - `\p{Digit}` = `[0-9]`, `\p{Alpha}` = `[a-zA-Z]`, `\p{Alnum}` = `[a-zA-Z0-9]`, `\p{Upper}`, `\p{Lower}`, `\p{Space}`, `\p{Punct}`, `\p{Blank}` (space/tab), `\p{XDigit}` (hex digits), etc. - Capitalize-negate is not used here; to negate use `\P{...}` or wrap in a negated class `[^\p{Alpha}]`. ## ASCII vs Unicode — the key caveat By default, both the shorthands and the POSIX classes are **ASCII-only**. So `\d` does NOT match the Arabic-Indic digit `٠` or a fullwidth digit, and `\w` does not match accented letters like `é` by default. This surprises people processing international text. To opt in to Unicode semantics, compile with the flag `Pattern.UNICODE_CHARACTER_CLASS` (or prepend the inline flag `(?U)`): ```java Pattern.compile("\\w+", Pattern.UNICODE_CHARACTER_CLASS); ``` Now `\d` matches any Unicode decimal digit, `\w` any Unicode word character, `\s` any Unicode whitespace, and the POSIX classes follow Unicode definitions. Note this is **separate** from `UNICODE_CASE` (which only affects case-insensitive matching). ## Unicode property classes vs POSIX The `\p{...}` syntax also reaches **true Unicode properties** (covered more in its own area): `\p{L}` is "any letter" across all scripts, `\p{Lu}` uppercase letters, `\p{N}` numbers. These are always Unicode-defined regardless of the flag, unlike the POSIX names which are ASCII until UNICODE_CHARACTER_CLASS is set. ## Negation nuances `\D` and `[^\d]` both mean "not a digit" on their own. The difference appears in **combination**: inside a single class `[^\d\s]` means "neither digit nor whitespace", whereas naively chaining `\D\S` would match two characters. Prefer a single negated class to combine exclusions. ## Java source escaping Every regex backslash doubles in a Java string literal: the regex `\d` is `"\\d"`, and `\p{Alpha}` is `"\\p{Alpha}"`. Forgetting this is a frequent compile-time or behavior bug. ## Deriving the answer Hold three facts: (1) `\d\w\s` are shorthands with uppercase negations, usable inside `[...]`; (2) `\p{...}` gives POSIX names; (3) all default to ASCII and need `UNICODE_CHARACTER_CLASS` for Unicode. From those you can answer any question about coverage and international input.

  • What is the difference between UNICODE_CHARACTER_CLASS and UNICODE_CASE?
    UNICODE_CHARACTER_CLASS makes \d, \w, \s and POSIX classes use Unicode definitions (membership). UNICODE_CASE only affects how case-insensitive matching folds case; it does not change which characters \w or \d match.
  • How would you match a single non-word character?
    Use \W (the negation of \w), or equivalently the negated class [^\w]. Both match one character that is not a letter, digit, or underscore.

saying these in an interview costs you the question

  • Assuming \d matches all Unicode digits by default (it is ASCII-only)
  • Thinking \w includes accented letters without the Unicode flag
  • Confusing UNICODE_CHARACTER_CLASS with UNICODE_CASE
  • Believing \p{...} negates by uppercasing (use \P{...})

context