Can Java identifiers contain Unicode (non-ASCII) characters? Explain what 'letter' means in the identifier rules.
answer
- 'Letter' = Unicode letter, not just A–Z
- isJavaIdentifierStart / isJavaIdentifierPart are the rule
- $ = currency symbol, _ = connector punctuation
- \uXXXX escapes resolved before tokenizing
- Legal but teams keep names ASCII
basics
~20 sYes. Java allows Unicode letters in identifiers, so you can use characters from other languages, like accented letters or non-Latin scripts. A 'letter' is not limited to A–Z; many Unicode letter characters and currency symbols are allowed too.
solid answer
~40 sJava identifiers are not restricted to ASCII. The language defines the legal characters in terms of Unicode, not just A–Z, so a name can include letters from many scripts — accented Latin letters, Greek, Cyrillic, CJK, and so on. Formally, the first character must satisfy Character.isJavaIdentifierStart and later characters Character.isJavaIdentifierPart, which accept Unicode letters, the connector punctuation (like underscore), currency symbols (like $ and €), and digits where appropriate. So names like resumé, naïve, or 名前 compile. In practice teams stick to ASCII for portability and tooling reasons, but the rule itself is Unicode-aware. Note Java source can also use Unicode escapes (\uXXXX), which the compiler resolves very early, so an escape for a letter can even form part of an identifier.
go deeper
Knows that Java is not limited to A–Z and that other-language letters can appear in names.
Explains that 'letter' is Unicode-defined and names the isJavaIdentifierStart/Part methods.
Distinguishes Unicode characters from Unicode escapes, knows escapes are resolved pre-tokenization, and can justify the ASCII-only convention.
Discusses homoglyph/security and tooling/portability tradeoffs and can reason about lexer phases and JLS character-category definitions.
## The key idea: 'letter' means Unicode letter When people first learn the identifier rules they hear 'must start with a letter,' and they assume *letter* means `a`–`z` and `A`–`Z`. In Java that is too narrow. The Java Language Specification defines the allowed characters using **Unicode**, the global character standard that assigns a number (code point) to every character in essentially every human writing system. The formal definition is given by two methods on the `Character` class: - **`Character.isJavaIdentifierStart(c)`** — true if `c` may be the *first* character of an identifier. This includes any Unicode **letter** (Latin, Greek, Cyrillic, Han/Chinese, Hiragana, etc.), the **letter-number** category, **currency symbols** (`$`, `€`, `¥`, …), and **connecting punctuation** such as the underscore `_`. - **`Character.isJavaIdentifierPart(c)`** — true if `c` may appear *after* the first character. This is everything `isJavaIdentifierStart` allows, **plus** digits, plus a few non-printing 'ignorable' control characters. So `$` and `_` are legal not as special cases but because they fall into the currency-symbol and connector-punctuation categories that the rule admits. ## Concrete consequences Because the rule is Unicode-based, all of these are legal identifiers: - `café`, `naïve`, `resumé` (accented Latin letters) - `π`, `λ` (Greek letters — sometimes used in math code) - `名前`, `имя`, `이름` (CJK / Cyrillic / Hangul names) - `€uros` (a currency symbol can even start a name) The compiler accepts them all. What it will *not* accept is, say, an emoji used where a letter is required, a space, or a punctuation mark that isn't a connector/currency symbol. ## Unicode escapes vs Unicode characters There is a subtle related point. Java source can contain **Unicode escapes** of the form `\uXXXX`, where `XXXX` is a hex code point. The compiler translates these escapes into the actual characters in a very early lexical step, *before* tokenizing. That means `age` is exactly the same identifier as `age` (because `a` is the letter `a`). This also means a stray `\u` sequence in a comment can break compilation — a classic gotcha — because the translation happens before comments are recognized. ## Why teams usually avoid non-ASCII identifiers Even though it is legal, most production code restricts identifiers to ASCII. Reasons: keyboards and editors vary, some tools and diff viewers mishandle non-ASCII, visually similar characters from different scripts (homoglyphs) can hide bugs or be a security concern, and English-based naming is the de facto convention in most codebases. So the practical answer is 'yes it is allowed, but we keep names ASCII anyway.' ## Summary The identifier rule is defined over Unicode, not ASCII. 'Letter' includes letters from every script, plus currency symbols and connector punctuation; digits are allowed after the first position. `Character.isJavaIdentifierStart`/`Part` are the authoritative tests, and Unicode escapes are resolved before tokenizing so they can form identifiers too.
- Is café a legal variable name?Yes — é is a Unicode letter, so café satisfies the identifier rules and compiles. Teams may still avoid it for portability.
- How does age relate to the identifier age?They are identical. a is the letter 'a', and Unicode escapes are resolved before tokenizing, so both spellings denote the same identifier.
saying these in an interview costs you the question
- Claiming only ASCII letters are allowed
- Confusing Unicode escapes (\uXXXX) with String escapes in literals
- Saying emoji are allowed as identifier letters (most are not 'letters')
- Believing non-ASCII identifiers are forbidden by the compiler