How do class union and the Java-specific class intersection (&&) and subtraction work, e.g. [a-z&&[^aeiou]]?
answer
- Juxtapose members = union
- && = intersection (both sides)
- Intersect with [^...] = subtraction
- [a-z&&[^aeiou]] = consonants
- && is Java/ICU-only, not portable
basics
~10 sListing members in a class unions them: [a-d[m-p]] is a..d or m..p. Java adds intersection with &&: [a-z&&[^aeiou]] means lowercase letters AND not a vowel, i.e. consonants. It is a set-difference trick.
solid answer
~40 sInside a Java character class, simply juxtaposing members forms a union: [a-d[m-p]] (or just [a-dm-p]) matches a-d or m-p. Java extends POSIX with set intersection using &&: [a-z&&[def]] matches characters in BOTH operands, here d, e, f. Combined with negation you get subtraction: [a-z&&[^aeiou]] is "lowercase letters intersected with non-vowels" = the lowercase consonants. Nesting is allowed, so [\p{L}&&[^\p{Lu}]] approximates lowercase letters. The && binds the union of everything on its left against the class on its right, so [a-z&&[^bc]] removes b and c. This intersection syntax is a Java/ICU-style feature; many other regex flavors do not support it, which matters for portable patterns.
code
java · 9 linesimport java.util.regex.Pattern;
var consonant = Pattern.compile("[a-z&&[^aeiou]]"); // lowercase letters minus vowels
System.out.println(consonant.matcher("b").matches()); // true
System.out.println(consonant.matcher("e").matches()); // false (vowel removed)
var both = Pattern.compile("[a-z&&[def]]"); // intersection
System.out.println(both.matcher("d").matches()); // true
System.out.println(both.matcher("a").matches()); // false (not in {d,e,f})go deeper
May not know intersection exists; can at least union ranges in one class.
Unions ranges and shorthands and recognizes && produces an intersection.
Uses [a-z&&[^...]] for subtraction, understands && precedence, and flags portability when copying patterns to other engines.
Weighs maintainability/portability of && idioms, documents the Java dependency, and chooses clearer alternatives for cross-engine code.
## Building bigger sets from smaller ones A character class describes a **set** of single characters. Java lets you build that set with three operations: **union** (default), **intersection** (`&&`), and, via negation, **subtraction**. ## Union (the default) Putting members next to each other unions them. Nested brackets are allowed but optional: - `[a-dm-p]` and `[a-d[m-p]]` are identical: match `a`-`d` **or** `m`-`p`. - You can union ranges, loose chars, and shorthands: `[\d\sA-F]` = a digit, whitespace, or `A`-`F`. ## Intersection with && Java (following the ICU/Unicode-style extension) adds **intersection** with the `&&` operator: a character must be in **both** sides. - `[a-z&&[def]]` = characters in `a-z` AND in `{d,e,f}` = `d`, `e`, `f`. - The left operand is everything to the left of `&&` (treated as a union); the right operand is the class after `&&`. ## Subtraction = intersection with a negated class There is no dedicated minus operator; you express **set difference** by intersecting with a **negated** class: - `[a-z&&[^aeiou]]` = lowercase letters AND not a vowel = the lowercase **consonants**. - `[a-z&&[^bc]]` = lowercase letters except `b` and `c`. - `[\p{L}&&[^\p{Lu}]]` ≈ letters that are not uppercase (an approximation of lowercase letters across Unicode; `\p{Lu}` is the Unicode "uppercase letter" property). ## How precedence reads Think of `&&` as splitting the class into left-union `&&` right-class. So in `[a-z0-9&&[^5]]`, the left side is `a-z` or `0-9`, intersected with "not 5", giving alphanumerics except `5`. To intersect more than two sets, nest: `[A&&B&&C]` parses left-to-right. ## Portability warning This `&&` intersection is a **Java-specific (ICU-family)** feature. Classic POSIX/PCRE and JavaScript do **not** support it; a pattern relying on `&&` will not behave the same when copied to those engines (in some, `&` is literal). If a regex must be portable, express subtraction another way (e.g. enumerate, or use lookarounds) or document the Java dependency. ## Empty and degenerate intersections An intersection with no common members matches nothing (it can never succeed), e.g. `[a-c&&[x-z]]`. That is legal but usually a bug — verify operands overlap. ## Java source escaping The `&&` needs no escaping, but any backslash shorthand inside doubles in source: the regex `[\p{L}&&[^\p{Lu}]]` is `"[\\p{L}&&[^\\p{Lu}]]"`. ## Deriving the answer Remember: juxtapose = union; `&&` = intersection; intersect with `[^...]` = subtraction; and `&&` is Java-only. With those four facts you can read or build any composite class, including the consonants idiom.
- Why is [a-z&&[^aeiou]] preferable to listing every consonant manually?It is shorter, self-documenting ("letters minus vowels"), and less error-prone than enumerating 21 consonants. The subtraction intent is explicit and easy to adjust.
- What happens if you copy a && pattern into JavaScript?JavaScript's regex engine does not support && intersection; it would treat & as a literal ampersand, changing the meaning silently. You must rewrite the pattern for portability.
saying these in an interview costs you the question
- Assuming && works in JavaScript/PCRE (Java-specific)
- Treating & as literal and getting an unexpected match
- Expecting a dedicated minus/subtraction operator (there isn't one)
- Writing a non-overlapping intersection that can never match