skip to content

What does `\b` match in a Python `re` pattern, and what does `re.ASCII` change about it?

level: seniorimportance: should knowfreq 40%

answer

  1. It matches a position, not a character
  2. Defined entirely in terms of `\w`
  3. `\w` is Unicode by default for `str`
  4. One flag shrinks the alphabet to ASCII
  5. `\B` is its exact negation

basics

~20 s

\b is a zero-width assertion: it matches at a position where exactly one side is a word character. ‘Word character’ means \w, which for str patterns is Unicode-aware by default; re.ASCII narrows it to [A-Za-z0-9_].

solid answer

~40 s

`\b` consumes nothing — it asserts a boundary between a `\w` character and a non-`\w` character (or the edge of the string). So `\bcat\b` matches `cat` in `"the cat."` but not inside `"category"` or `"concat"`. `\B` is its exact negation. Everything therefore hinges on the definition of `\w`, and for a `str` pattern that definition is the full Unicode set of letters, digits and underscore — accented and non-Latin letters included. Passing `re.ASCII` restricts `\w`, `\W`, `\b`, `\B`, `\d`, `\D`, `\s` and `\S` to ASCII, which quietly turns `é` into a non-word character and makes `\bcafé\b` stop matching `café`. `bytes` patterns are ASCII-only already, so the flag is redundant there. Reach for `re.ASCII` only when the input really is ASCII protocol text, never on user-supplied prose.

code

pycon · 9 lines
pycon
>>> import re
>>> re.findall(r"\bcat\b", "cat category the cat.")
['cat', 'cat']
>>> re.findall(r"\w+", "naïve café")
['naïve', 'café']
>>> re.findall(r"\w+", "naïve café", re.ASCII)
['na', 've', 'caf']
>>> print(re.search(r"\bcafé\b", "café", re.ASCII))
None

go deeper

for a junior

Recall that \b marks a word boundary and matches no characters of its own, and that \bcat\b finds the standalone word but not the cat inside category. That distinction alone answers the screening version of this question.

for a middle

Explain the mechanics: \b succeeds where exactly one neighbour is a \w character, \B is its negation, and \w for a str pattern covers Unicode letters and digits. Be able to say what re.ASCII narrows and what it leaves alone.

for a senior

Demonstrate the production judgement: re.ASCII on protocol text, never on human text, because the failure is silent — a \b rule simply stops firing on names with accented letters. Mention that re.IGNORECASE folds the full Unicode range and what that means for a check you rely on.

for a principal

Own the policy on how text is normalised and validated before it reaches a rule engine at all, and whether pattern rules should be expressed in terms of \b or an explicit delimiter class. Silent under-matching in a scanning rule is a whole class of incident you want designed out.

### Zero-width means it matches a position, not a character `\b` is an assertion. It succeeds or fails at a position between two characters (or at the very start or end of the subject) and consumes nothing, so it never appears in the matched text and never advances the engine. Concretely, `\b` succeeds where exactly one of the two neighbouring positions holds a word character: ```python import re re.findall(r"\bcat\b", "cat category the cat.") # ['cat', 'cat'] re.search(r"\bcat", "concat") # None re.findall(r"\Bcat\B", "concatenate") # ['cat'] ``` `\B` asserts the opposite: a position where both sides are word characters, or neither is. The pair is exhaustive, which is why `\Bcat\B` finds the embedded occurrence that `\bcat\b` refuses. One trap worth knowing: **inside a character class, `\b` is not a boundary at all**. `[\b]` means the backspace character, U+0008. Character classes match characters, and a zero-width assertion has no meaning there, so the escape is reused. ### The whole answer depends on `\w` Since `\b` is defined in terms of word characters, the interesting question is what counts as one. For a `str` pattern, `\w` means the Unicode word characters: anything Unicode classifies as a letter or a decimal digit, plus the underscore. That includes `é`, `ï`, Greek, Cyrillic, CJK ideographs and much more. ```python re.findall(r"\w+", "naïve café") # ['naïve', 'café'] re.findall(r"\w+", "naïve café", re.ASCII) # ['na', 've', 'caf'] ``` That second line is the failure mode in one image. Under `re.ASCII` the accented characters become non-word characters, so a single word is split into fragments — and, correspondingly, `\b` starts asserting boundaries in the *middle* of words. A blocklist rule written as `\bfree\b` behaves fine either way, but `\bcafé\b` matches `café` normally and matches nothing at all with `re.ASCII` set, because there is no boundary after `é` when `é` is not a word character to begin with. ### What `re.ASCII` actually covers `re.ASCII` (alias `re.A`, inline `(?a)`) makes `\w`, `\W`, `\b`, `\B`, `\d`, `\D`, `\s` and `\S` perform ASCII-only matching. It changes nothing else: literals, `.` and explicit classes such as `[a-zé]` are unaffected. It exists because ASCII-only semantics are occasionally what you want — parsing an identifier in a wire protocol, a hex token, a header name — and because they are cheaper and more predictable in exactly those cases. `bytes` patterns are ASCII-only unconditionally: `re.findall(rb"\w+", "naïve".encode())` yields `[b'na', b've']` with no flag at all, because the encoded accent is a pair of non-ASCII bytes. Passing `re.UNICODE` with a `bytes` pattern is not merely redundant, it raises `ValueError: cannot use UNICODE flag with a bytes pattern`. ### The `re.IGNORECASE` half of the same story The same Unicode-by-default posture makes `re.IGNORECASE` stronger than most people expect on `str` patterns: it performs full Unicode case folding. `re.search(r"s", "ſ", re.IGNORECASE)` matches, because LATIN SMALL LETTER LONG S case-folds to `s`; `re.search(r"k", "K", re.IGNORECASE)` matches the KELVIN SIGN for the same reason. If you are using a case-insensitive regular expression as a security or de-duplication check, those equivalences are either a feature (they close a bypass) or a surprise (two ‘different’ keys collide), and you should know which. Combining `re.IGNORECASE | re.ASCII` restores plain ASCII folding. ### Deciding, in production The judgement an interviewer is probing is this: whose text is it? * **Machine text you control the grammar of** — identifiers, tokens, header names, generated codes — `re.ASCII` is a reasonable tightening. It makes `\w` mean what the grammar says it means instead of whatever Unicode has grown to include, and it makes the pattern reject exotic look-alike input rather than silently accepting it. * **Human text from users** — names, addresses, memo fields, product descriptions — do not use `re.ASCII`. Any rule built on `\b` will misbehave on perfectly ordinary names the moment they contain a letter outside ASCII, and the failure is silent: the rule simply stops firing. And for word-boundary rules generally, remember that `\b` is a *lexical* notion, not a semantic one. It knows nothing about hyphenation, apostrophes or scripts that do not delimit words with spaces at all. `\bcat\b` will happily match inside `cat-nap`, because `-` is not a word character. If your rule must respect a real notion of a token, state that boundary explicitly — for instance a class of the delimiters you actually allow — rather than trusting `\b` to mean what a reader assumes it means. ### What to say in an interview Define `\b` as zero-width and relative to `\w`; show that `\bcat\b` skips `category`; state that `\w` is Unicode for `str` and ASCII for `bytes`; then explain `re.ASCII` as narrowing the alphabet, and give the accented-word example as the reason not to set it on human text.

  • How does `re.IGNORECASE` behave on non-ASCII text in a `str` pattern?
    It case-folds using full Unicode rules, so a pattern `s` matches LATIN SMALL LETTER LONG S and `k` matches the KELVIN SIGN. That is usually what you want for human text, but it means case-insensitive comparison merges characters you may not have considered equal. Adding `re.ASCII` restricts folding to the ASCII letters.
  • Do `bytes` patterns need `re.ASCII` to get ASCII-only class escapes?
    No — they are ASCII-only unconditionally, so the flag is redundant there. Going the other way is an outright error: passing `re.UNICODE` with a `bytes` pattern raises `ValueError: cannot use UNICODE flag with a bytes pattern`. The practical consequence is that matching encoded text as bytes silently splits multi-byte characters.
  • What does `[\b]` mean inside a character class?
    The backspace character, U+0008. A character class matches characters, so a zero-width assertion has no meaning inside it and the escape is reused for the control character. This is a real source of silently-wrong patterns: `[\bx]` is a class of backspace and `x`, not ‘a boundary or an x’.
  • Does `\bcat\b` match inside the string `cat-nap`?
    Yes. `-` is not a word character, so there is a genuine boundary after `cat`. `\b` is a lexical assertion about `\w` neighbours, not a notion of a semantic token, and it knows nothing about hyphenation or apostrophes. If your rule needs a specific set of delimiters, spell that class out instead of relying on `\b`.

saying these in an interview costs you the question

  • Says `\b` matches a space or a literal delimiter character
  • Thinks `\b` consumes a character and appears in the match
  • Assumes `\w` is `[A-Za-z0-9_]` for `str` patterns
  • Sets `re.ASCII` on user-supplied prose containing accents
  • Treats `\b` and `\B` as the same assertion
  • Believes `re.IGNORECASE` folds only the 26 ASCII letters

context