How do char ranges like `'a'..'z'` work for character classification, and how would you test if a Char is an ASCII letter or digit using ranges?
answer
- `'a'..'z'` is a CharRange ordered by code point
- ranges are ASCII-only; `isLetter()`/`isDigit()` are Unicode-aware
- compose with `||` for letters-or-digits
- 'a'=97, 'A'=65, '0'=48
- `'é' in 'a'..'z'` is false but `'é'.isLetter()` is true
basics
~10 s'a'..'z' is a range of characters in code-point order. c in 'a'..'z' is true for lowercase letters. Combine ranges with || to test letters or digits.
solid answer
~40 s`'a'..'z'` produces a `CharRange` whose membership follows the Unicode code-point order of the chars. `c in 'a'..'z'` returns true exactly when `c`'s code is between 'a' (97) and 'z' (122) inclusive — same `contains` machinery as numeric ranges. For classification you compose ranges: `c in 'a'..'z' || c in 'A'..'Z'` for ASCII letters, `c in '0'..'9'` for ASCII digits. This only covers ASCII; for full Unicode use the stdlib predicates `Char.isLetter()`, `Char.isDigit()`, `Char.isLetterOrDigit()`, `Char.isWhitespace()`, which delegate to `Character` and handle all Unicode categories. Char ranges are convenient and explicit but ASCII-only, so they're ideal for parsers/validators where you deliberately want ASCII semantics (e.g. validating an identifier or hex char).
code
kotlin · 9 linesfun isValidIdentifierChar(c: Char, first: Boolean): Boolean {
val letter = c in 'a'..'z' || c in 'A'..'Z' || c == '_'
val digit = c in '0'..'9'
return if (first) letter else letter || digit
}
println(isValidIdentifierChar('A', first = true)) // true
println(isValidIdentifierChar('7', first = true)) // false (cannot start with digit)
println(isValidIdentifierChar('7', first = false)) // truego deeper
Knows c in 'a'..'z' tests for a lowercase letter and can combine with || for digits.
Explains code-point ordering, ASCII-only scope, and that stdlib isLetter()/isDigit() handle Unicode.
Chooses ranges vs. predicates deliberately based on ASCII vs. Unicode requirements and performance.
Reasons about correctness/locale/security implications of ASCII classification in parsers and validators across a codebase.
## What a `CharRange` is `'a'..'z'` calls `Char.rangeTo(Char)` and yields a **`CharRange`** — a closed, inclusive interval over `Char` values. A `Char` in Kotlin is a UTF-16 code unit, and ordering is by its numeric **code point**: `'a'` is 97, `'z'` is 122, `'0'` is 48, `'A'` is 65. So `c in 'a'..'z'` is true iff `97 <= c.code <= 122`. ```kotlin println('m' in 'a'..'z') // true println('M' in 'a'..'z') // false (uppercase M is code 77, below 'a') println('5' in '0'..'9') // true ``` ## Classifying characters with ranges Compose ranges with boolean operators: ```kotlin fun isAsciiLetter(c: Char) = c in 'a'..'z' || c in 'A'..'Z' fun isAsciiDigit(c: Char) = c in '0'..'9' fun isHexDigit(c: Char) = c in '0'..'9' || c in 'a'..'f' || c in 'A'..'F' ``` This is explicit and fast (each check is an O(1) comparison), and it is intentionally **ASCII-only**. ## Range classification vs. stdlib predicates The Kotlin stdlib also offers Unicode-aware predicates: - `Char.isLetter()` — true for any Unicode letter, including `'ä'`, `'ж'`, `'中'`. - `Char.isDigit()` — any Unicode decimal digit. - `Char.isLetterOrDigit()`, `Char.isWhitespace()`, `Char.isUpperCase()`, etc. **Key difference:** ```kotlin println('é' in 'a'..'z') // false — outside ASCII a-z println('é'.isLetter()) // true — Unicode letter ``` Use **char ranges** when you specifically want ASCII semantics (parsers, hex, identifiers in many DSLs); use the **`isXxx()`** predicates when you want correct Unicode classification. ## `when` with char ranges ```kotlin fun category(c: Char) = when (c) { in '0'..'9' -> "digit" in 'a'..'z', in 'A'..'Z' -> "letter" else -> "other" } ``` Multiple `in` conditions can sit on one branch separated by commas.
- Why does `'é' in 'a'..'z'` return false even though é is a letter?The char range only covers ASCII code points 97-122; 'é' has a higher code point. For Unicode-correct results use `'é'.isLetter()`.
- How do you test for a hex digit using ranges?`c in '0'..'9' || c in 'a'..'f' || c in 'A'..'F'`.
Char ranges are like a ruler marked only with ASCII; the isLetter() predicates are a full multilingual dictionary.
saying these in an interview costs you the question
- Believing `c in 'a'..'z'` matches accented or non-Latin letters.
- Confusing `CharRange` membership ('a'..'z') with `String.contains`.
- Using `'a'..'Z'` (lowercase to uppercase) and expecting all letters — that range is empty because 'a'(97) > 'Z'(90).
- Not knowing `isLetter()`/`isDigit()` exist for Unicode classification.