When is whitespace actually required between tokens in Java, and what rule of the lexer makes it necessary?
answer
- Maximal munch = longest match, greedy lexer
- Word tokens (keyword/identifier/number) need separating
- Punctuation self-separates: 'a+b' is fine
- Comments act as whitespace separators
- '++' vs '+ +' shows operator-side munch
basics
~20 sYou need whitespace when two pieces of code would otherwise stick together into one word, like 'int x' (without the space it becomes the single name 'intx'). The compiler reads the longest possible word at a time, so it can't split joined words by itself.
solid answer
~40 sWhitespace is required only to keep two adjacent tokens from merging when both are made of identifier-type characters. The lexer uses 'maximal munch' (longest-match): it greedily consumes the longest run of valid identifier/keyword/number characters, so 'intx' is read as one identifier, not 'int' + 'x'. Therefore keyword-then-identifier ('return value'), identifier-then-identifier, and number-then-identifier need a separator. Tokens made of punctuation (operators, braces, commas, parentheses) do NOT need separating from words or each other, because the lexer can tell where they start - 'a+b', 'list.get(i)', and '){' are all fine. Edge cases: a digit-letter boundary in a literal must be careful ('0x1Ep2' is one token), and '+ +' vs '++' differ, but those are about literal/operator grammar more than free whitespace.
go deeper
Recognizes that you need a space between words like 'int x' and that operators don't need spaces, even without naming the rule.
Names maximal munch / longest-match and explains it forces separators between word-like tokens while punctuation self-delimits.
Extends the rule to operator merging ('++' vs '+ +'), numeric-literal atomicity, and comments-as-whitespace, and connects a missing space to 'cannot find symbol' errors.
Discusses maximal munch as a lexer design choice (vs longest-match alternatives), its interaction with grammar ambiguity, and implications for code generators and minifiers.
## Tokens and the lexer Before Java compiles, a **lexer** (scanner) turns the raw character stream into **tokens** - the atomic units: keywords (`int`, `return`), identifiers (names like `count`), literals (`42`, `"hi"`), operators (`+`, `==`), and separators (`{`, `}`, `(`, `;`, `,`, `.`). Whitespace between tokens is thrown away; the question is *when can two tokens sit with no whitespace at all between them?* ## Maximal munch (the longest-match rule) The lexer is **greedy**: at each position it consumes the **longest sequence of characters that forms a valid token** before moving on. This rule is called **maximal munch** (or longest-match). It is why some boundaries need a space and others don't. Identifiers, keywords, and number literals are all built from the same **"word" character class** - letters, digits, `_`, `$`. So when two such tokens are adjacent, maximal munch swallows them into a single longer word: - `int x` -> tokens `int`, `x`. Remove the space: `intx` -> a **single identifier** `intx`. The keyword `int` is lost. - `return value` needs the space; `returnvalue` is one identifier. - A digit run followed by letters would also be misread. **Punctuation tokens** (operators and separators) are **not** word characters, so the lexer knows they end the current word and start a new token. That is why these are all fine with zero whitespace: ```java a+b // identifier, +, identifier list.get(i) // identifier . identifier ( identifier ) }else{ // '}' and '{' are separators that bound the keyword 'else' ``` ## The required-separator rule, stated precisely > A separator (whitespace **or** a punctuation token) is required between two tokens **only if** removing it would let maximal munch read them as a single, different token. In practice this means: keep whitespace between any two **word-like** tokens (keyword/identifier/number literal) that are directly adjacent. ## Subtle cases - **Operator merging:** `i+ +j` vs `i++j`. `++` is a single operator under maximal munch, so spacing changes the parse. `a- -b` vs `a--b` similarly. This is the operator-side analogue of the word rule. - **Numeric literals:** a literal is one greedy token. `0x1Ep2`, `1_000`, `3.14f` are each a single token; you cannot insert whitespace inside them, and adjacent letters can be misread as a suffix. - **Comments count as whitespace:** `int/*c*/x` is legal because the comment acts as a separator - the lexer treats a comment like whitespace. ## Why it matters Understanding maximal munch explains *why* you can compress operator-heavy code (`for(int i=0;i<n;i++)`) but must keep `int i` apart, and it demystifies confusing minified or generated code. It is also the reason a missing space sometimes yields a baffling 'cannot find symbol: intx' error instead of a syntax error.
- Why does 'int/*comment*/x' compile even with no spaces?Because a comment is treated as whitespace by the lexer. It separates 'int' and 'x' just as a space would, so the two tokens don't merge into 'intx'.
- Does 'a+++b' parse, and how?Yes, as '(a++) + b' via maximal munch: the lexer takes the longest operator first ('++'), then '+', so it reads a++, +, b. Spacing it as 'a++ + b' is clearer but identical.
The lexer is a greedy reader sounding out the longest word it can each time. 'intx' looks like one word to it, so you insert a space to say 'these are two words'. Punctuation is like a hyphen the reader already sees as a break.
saying these in an interview costs you the question
- Claiming whitespace is required around operators to compile
- Not knowing the longest-match/maximal-munch rule by name or concept
- Thinking the compiler can split 'intx' back into 'int x'
- Forgetting comments serve as token separators