skip to content

Why is a comment delimiter like // or /* ignored when it appears inside a string literal?

level: seniorimportance: should knowfreq 45%

answer

  1. Lexer is mode/state based
  2. String mode only seeks closing quote
  3. // in "http://..." is literal
  4. Quote inside a comment is literal too
  5. Strip comments with a lexer, not regex

basics

~20 s

Inside a string the characters are just data, not source syntax. The lexer knows it is reading a string, so it does not treat // or /* as comment starters. So System.out.println("http://x") prints the whole URL.

solid answer

~50 s

Comment delimiters have no special meaning inside a string (or character) literal because the lexer is *context-sensitive*: when it has entered a string literal it scans for the closing quote, not for comment starters. So `//`, `/*`, and `*/` inside `"..."` are ordinary characters that become part of the string value. The canonical example is a URL: `String s = "http://example.com";` — the `//` after `http:` does not start a comment and the rest of the line is preserved. The same applies in reverse: a quote character inside a comment is meaningless, because the lexer is in comment mode and only looks for `*/` (or end of line for `//`). This mutual exclusivity is a fundamental tokenizer property; a naive find-and-strip of `//` would corrupt strings, which is why tools that process Java source must use a real lexer rather than regex.

go deeper

for a junior

Recognize that // inside "..." is part of the string, e.g. a URL prints in full.

for a middle

Explain it via lexer string-mode vs code-mode and give the symmetric case (quote inside a comment).

for a senior

Articulate the tokenizer-state model, escape handling, and why regex comment-stripping is unsafe for tooling.

for a principal

Reason about correct source transformation pipelines (lexer-accurate codemods), text blocks, and the Unicode-escape pre-pass edge case.

### The phenomenon This compiles and prints a full URL, comment-looking `//` and all: ```java System.out.println("Visit http://example.com /* not a comment */ now"); ``` Neither the `//` nor the `/* ... */` is treated as a comment. They are just characters in the string. ### Why — lexer context (modes) The **lexer** (scanner) reads the source one character at a time and is always in some *mode/state*: - **Normal code mode:** here `//` starts a line comment and `/*` starts a block comment. - **String-literal mode:** entered when it reads an opening `"`. In this mode it only watches for the **closing `"`** (and for escape sequences like `\"` or `\n`). It does **not** look for comment starters. Every other character, including `/`, `*`, is appended to the string value. - **Char-literal mode:** same idea for `'x'`. - **Comment mode:** entered at `//` (until end of line) or `/*` (until `*/`). In this mode a `"` is just a character — it does **not** start a string. Because these modes are mutually exclusive, the meaning of `//` depends entirely on *where* it appears: - In code → comment. - In a string → literal text. ### Symmetric consequence A quote inside a comment does nothing: ```java // he said "hello" and left ``` The `"` characters do not open a string; the whole line is a comment. ### Escapes still apply inside strings String mode also handles escape sequences, so `"a\"b"` is the 3-character string `a"b` — the escaped quote does not end the string. Comment delimiters get no such treatment because they were never special in string mode to begin with. ### Why it matters in practice - **URLs and file paths** (`http://`, `C:/foo`) in strings are safe — no accidental comment. - **Source-processing tools must use a real lexer.** A regex that strips everything after `//` will destroy URLs inside strings, break code formatters, minifiers, and codemods. The correct approach tracks string/char/comment state exactly as the compiler does. - **Text blocks** (`""" ... """`, Java 15+) follow the same principle: delimiters inside the block are literal text. ### One subtlety: Unicode escapes Java processes `\uXXXX` Unicode escapes very early — even before tokenizing. So a `*/` could form `*/`. This is an edge case worth knowing but rarely encountered; the normal rule (context decides) holds for ordinary characters.

  • Why is using a regex to strip // comments from Java source dangerous?
    Because // inside a string literal is legitimate text (e.g. a URL). A naive regex can't tell code from string context and will corrupt strings; you need a state-tracking lexer.
  • Does a double-quote inside a /* */ comment start a string literal?
    No. While in comment mode the lexer only looks for */ ; quote characters are just text and have no effect.

saying these in an interview costs you the question

  • Claiming "http://x" gets truncated at // — it does not
  • Thinking a regex //.*$ is a safe way to strip comments
  • Forgetting that a quote inside a comment does NOT start a string

context