skip to content

In a scannerless grammar with no token stage, why must every rule that may be followed by spacing say so explicitly?

level: seniorimportance: should knowfreq 36%

answer

  1. one stage, reading raw characters
  2. no pass discards the gaps
  3. unconsumed space breaks the next literal
  4. one spacing rule, one consistent position
  5. boundaries need a predicate too

basics

~20 s

Nothing discards whitespace for a scannerless parser: its rules match raw characters, so any space, tab or newline a rule does not consume is still sitting there and makes the next literal fail. Spacing is threaded through the rules by hand.

solid answer

~50 s

A scannerless parser has one stage, and that stage reads characters. There is no earlier pass that grouped input and quietly dropped the gaps, so whitespace is ordinary input: if `"@" Name "("` is written with nothing between the elements, then `@note (x)` simply fails at the space. The usual discipline is a rule such as `Spacing <- (" " / "\t" / "\r" / "\n")*` invoked after every literal that may be followed by a gap — often by wrapping each literal in a helper so it is impossible to forget. The compensation is real: because layout is visible to the grammar, whitespace-sensitive syntax becomes expressible instead of awkward, and a text document with annotations embedded in it can be described by one grammar rather than by a pre-pass and a parser that must agree with each other.

go deeper

for a junior

Recall that without a separate grouping stage the rules match raw characters, so a space nobody consumes will break the next literal in the sequence.

for a middle

Explain the discipline: one spacing rule that can match zero characters, applied in a consistent position, plus a predicate wherever a literal could match the front of a longer word.

for a senior

Diagnose the real symptom — a grammar that accepts compact input and rejects the same input formatted — and fix it by convention rather than by patching the one rule that failed.

for a principal

Decide whether a single-stage grammar is the right shape for the syntax at all: it buys layout sensitivity and nesting inside free text, and it costs a spacing and boundary duty on every rule the team writes.

## One stage, so no free gaps In a design with a separate first stage, that stage groups characters into units and discards the runs between them, so the rules that follow never see a space. A **scannerless** grammar has no such stage: its rules match characters directly. Whitespace is therefore ordinary input that some rule must consume, exactly like a bracket or a name. The symptom is unmistakable once you know it. A grammar reading ``` Annotation <- "@" Name "(" Arg ")" ``` accepts `@note(x)` and rejects `@note (x)`, because between `Name` and `"("` there is nothing that can consume the space. The rule is not wrong about the syntax; it is silent about the gaps, and silence means *no gap allowed*. ## The discipline that keeps it manageable ``` Spacing <- (" " / "\t" / "\r" / "\n")* token(e) <- e Spacing # match e, then absorb any trailing gap ``` 1. Define one `Spacing` rule and nothing else that consumes gaps, so the definition of "blank" lives in one place. 2. Apply it in **one consistent position** — conventionally after each element rather than before — so that every rule can assume the input starts at a non-blank character. 3. Consume leading spacing exactly once, at the entry rule, since the after-each-element convention never handles the very beginning. 4. Wrap literals in a helper so adding a new literal cannot forget the spacing; a grammar that spells `Spacing` out by hand at fifty call sites will miss some. Because `Spacing` uses repetition, it succeeds on zero characters too, so writing it in a place where no gap appears costs nothing. ## Boundaries are your job as well The same absence explains a second duty. With a separate stage, a run of letters is grouped before any comparison happens, so a keyword can never match the front of a longer word. Here it can, and the fix is a predicate: ``` Keyword <- "note" !Letter Spacing ``` `!Letter` succeeds only when the next character is not a letter and consumes nothing, so `notebook` no longer matches the keyword. Any place where one literal is a prefix of possible input needs that guard written by hand. ## What you get in exchange | Concern | Separate stage first | Scannerless | |---|---|---| | Whitespace between elements | dropped for you | consumed by a rule you write | | Word boundaries | implied by grouping | stated with a predicate | | Layout-significant syntax | awkward: the stage must report it | natural: the rules see the characters | | Foreign syntax nested inside text | two components must agree | one grammar covers both | | Where a spacing bug shows up | in the grouping stage | as an unexpected parse failure | For the frame this leaf lives in — annotations embedded in ordinary prose — the last two rows are the whole reason to be scannerless. The document is mostly free text, and the ordinary-text rule can be written directly: ``` Text <- (!Annotation .)+ ``` This consumes characters one at a time for as long as an annotation does not start here; `!Annotation` looks ahead and consumes nothing, and `.` takes exactly one character. There is no way to express that cleanly when an earlier stage has already decided what counts as a unit, because "everything up to the next thing that would parse as an annotation" is not a category that stage knows about. ## Failure modes worth recognising - **A grammar that works only on compact input**, because spacing was added where the examples needed it and nowhere else. - **Spacing consumed in two places**, before and after an element, which is harmless for correctness but doubles the work and confuses later readers about the convention. - **A gap allowed where it must not be**, such as between the `@` marker and the name, because the helper was applied uniformly without asking where a break is meaningful. - **Comments forgotten.** If the syntax has them, they belong inside `Spacing`; otherwise a comment is a parse error everywhere except the one place someone handled it. - **An error message pointing at the wrong place**, since a missing spacing rule surfaces at the next literal rather than at the gap. The reflex to build: when a rule sequences two visible elements, ask what may legally sit between them, and write that answer into the grammar. Nothing else will.

  • Where should a spacing rule be applied — before each element or after it?
    Pick one and apply it everywhere; after each element is the common choice, because every rule can then assume it starts on a non-blank character. That convention leaves exactly one special case, the leading spacing at the very start of input, which the entry rule consumes once.
  • What does a scannerless grammar make easy that a two-stage design makes awkward?
    Syntax where layout matters, and syntax nested inside free text. The rules see the characters, so indentation or a run of ordinary prose up to the next marker can be described directly, instead of being negotiated between a grouping stage and a parser that must agree on categories the stage invented.
  • Should comments be part of the spacing rule?
    Usually yes, if they may appear wherever a gap may. Folding them into the one spacing rule means every site that already tolerates a gap tolerates a comment, and the definition stays in a single place instead of being repeated at the handful of sites someone remembered.

saying these in an interview costs you the question

  • Expects whitespace to be skipped automatically between rules
  • Adds a spacing rule in one place and assumes it applies everywhere
  • Thinks a scannerless grammar cannot express layout-sensitive syntax
  • Assumes a literal cannot match the front of a longer word
  • Claims explicit spacing rules make the grammar ambiguous