skip to content

What does \b match in Java regex, how does \B differ, and what counts as a word character at a boundary?

level: middleimportance: must knowfreq 66%

answer

  1. \b = one side word, one side non-word
  2. \w default = [A-Za-z0-9_]
  3. \B = NOT a boundary (both sides same kind)
  4. ASCII default → accented letters look like non-word
  5. [\b] means backspace, not boundary

basics

~20 s

\b matches a position where a word character is next to a non-word character (or the edge of the string) — a word boundary. \B matches the opposite: a position that is not a word boundary. Both are zero-width.

solid answer

~50 s

\b is a word boundary: a zero-width position where one side is a word character and the other side is a non-word character or the start/end of the input. A word character by default is [A-Za-z0-9_]. So \bcat\b matches 'cat' as a whole word but not the 'cat' inside 'category'. \B is the negation: it matches any position that is not a word boundary, i.e. between two word characters or between two non-word characters. A classic use is \bword\b for whole-word search and \Bword\B for 'word' embedded inside another token. Because the definition of a word character is ASCII-centric by default, boundaries around accented or non-Latin letters can surprise you; the UNICODE_CHARACTER_CLASS flag (or \b{g} grapheme boundaries in newer APIs) makes \b respect Unicode word characters. Like all anchors, \b and \B consume no characters.

code

java · 12 lines
java
import java.util.regex.*;

Pattern whole = Pattern.compile("\\bcat\\b");
System.out.println(whole.matcher("the cat sat").find()); // true
System.out.println(whole.matcher("category").find());    // false

Pattern embedded = Pattern.compile("\\Bcat\\B");
System.out.println(embedded.matcher("scatter").find());  // true

// ASCII default treats é as non-word -> boundary inside "café":
System.out.println(Pattern.compile("f\\b").matcher("café").find()); // true
System.out.println(Pattern.compile("(?U)f\\b").matcher("café").find()); // false

go deeper

for a junior

Knows \b finds whole words and \B is its opposite.

for a middle

Defines word char as [A-Za-z0-9_], explains both boundary directions, and uses \bword\b for whole-word search.

for a senior

Articulates the ASCII default pitfall, the UNICODE_CHARACTER_CLASS fix, and the [\b]=backspace gotcha.

for a principal

Reasons about grapheme boundaries, locale/Unicode correctness in tokenization, and performance/correctness tradeoffs in large-scale text processing.

## The idea of a word boundary When we search text for a whole word, we don't want `cat` to match inside `category` or `concatenate`. A **word boundary** captures exactly that intuition: it is the **position** between a 'word' part of the text and a 'non-word' part. ## What is a 'word character'? By default in Java regex, a **word character** is `[A-Za-z0-9_]` — ASCII letters, ASCII digits, and the underscore. This is the same set as the shorthand `\w`. Everything else (spaces, punctuation, and, by default, accented letters like `é`) is a **non-word character**. ## `\b` — word boundary `\b` is a **zero-width assertion** that succeeds at a position where **exactly one** of the two sides is a word character. Concretely it matches: - between a non-word char (or string start) and a word char — the **start** of a word, and - between a word char and a non-word char (or string end) — the **end** of a word. So in `"a cat!"`, `\b` matches before `c` and after `t`. The pattern `\bcat\b` matches the standalone word but fails inside `category` because after `t` there is another word char (`e`), so there is no boundary there. ## `\B` — non-boundary `\B` is the **negation**: it matches every position that is **not** a word boundary — i.e. where **both** sides are word chars, or **both** sides are non-word chars (including the edges, depending on what's adjacent). For example `\Bcat\B` matches the `cat` inside `local cat` ... no — it matches `cat` only when it is surrounded by word chars on both sides, e.g. the `cat` in `scatter`. Useful for finding embedded substrings. ## The ASCII default and its surprises Because the default word class is ASCII-only, a letter like `é` is treated as a **non-word** character. So in `"café"`, there is a `\b` between `f` and `é` (word→non-word), which is rarely what you want for natural-language text. To fix this, enable **UNICODE_CHARACTER_CLASS** (`Pattern.UNICODE_CHARACTER_CLASS`, inline `(?U)`), which redefines `\w` (and therefore `\b`) to include Unicode letters and digits. Newer JDKs also offer `\b{g}` for **grapheme cluster** boundaries, useful for emoji and combining marks. ## Zero-width, like all anchors `\b` and `\B` match positions and consume **no** characters. They are independent of the MULTILINE flag (that flag only affects `^`/`$`). Inside a Java string literal you must double the backslash: `"\\bcat\\b"`. Note also that inside a **character class** `[\b]` means the backspace character, not a word boundary — a notorious gotcha.

  • Why does \bcat\b fail to match in 'category'?
    After 't' comes 'e', a word character, so there is no word boundary there — both sides are word chars, so the trailing \b assertion fails.
  • How do you make \b treat accented letters as word characters?
    Enable UNICODE_CHARACTER_CLASS (Pattern.UNICODE_CHARACTER_CLASS or inline (?U)), which redefines \w and thus \b to use Unicode word characters.

saying these in an interview costs you the question

  • Thinking \b matches a space or an actual character (it is zero-width)
  • Assuming accented/Unicode letters are word chars by default (they aren't without (?U))
  • Using [\b] expecting a word boundary (that's backspace inside a class)
  • Believing MULTILINE affects \b (it does not)

context