When designing the grammar for an external textual DSL - for instance using a tool like Xtext with EBNF-style grammar rules - what does the grammar actually define, and how does that feed into the rest of the language's tooling such as parsing, the abstract model, and editor support?
answer
- EBNF rules -> parser + model
- grammar rule = metamodel class
- assignment = typed feature
- ambiguity -> LL(*) conflict
- concrete vs abstract syntax separation
basics
~20 sThe grammar is a set of rules describing what a valid sentence in your mini-language looks like, similar to a sentence-structure rule like 'an order has a customer name and a list of items.' Tools like Xtext turn that grammar into both a parser and an in-memory data model, and then automatically generate editor features like syntax highlighting and autocomplete from the same rules.
solid answer
~50 sAn EBNF-style grammar defines the DSL's concrete syntax as production rules made of terminals (keywords, literals) and non-terminals composed recursively, with cardinality operators (`?`, `*`, `+`) and alternatives (`|`). A tool like Xtext parses source text against these rules to build a parse tree, and simultaneously derives a typed metamodel (an Ecore/EMF model, functioning like an AST) directly from the grammar - each rule becomes a class, each assignment becomes a typed feature. That single generated metamodel then backs everything else: syntax highlighting, the outline view, cross-reference resolution (scoping), validation, and code generators all operate on the typed model rather than raw text. Good grammar design keeps the language unambiguous enough for the parser generator to resolve deterministically (e.g. LL(*)-parseable) and separates concrete syntax choices from the underlying concepts so the metamodel stays stable even if surface syntax later changes.
go deeper
Knows, roughly, that a grammar defines what counts as valid syntax in the language.
Understands that grammar rules generate both a parser and a typed model, and can describe basic EBNF constructs (terminals, non-terminals, cardinality operators).
Can discuss ambiguity resolution and LL(*) parsing constraints concretely, and explains why separating concrete from abstract syntax matters for language evolution.
Weighs grammar design decisions against long-term language evolvability - metamodel stability across grammar changes, and migration strategy for existing DSL documents when the grammar changes.
## What the grammar actually defines A DSL grammar's job is to define concrete syntax - the exact textual notation a valid program in the language may take - as a set of production rules. In an EBNF-style notation (the style tools like Xtext and ANTLR use), each rule names a syntactic category and describes what sequence of terminals (literal keywords, punctuation, or lexer tokens like identifiers and numbers) and other rule references may appear, with operators for: - **optionality** (`?`) - **repetition** (`*` for zero-or-more, `+` for one-or-more) - **alternation** (`|` for either-or) A rule such as `Order: 'order' name=ID 'for' customer=[Customer] items+=OrderLine*;` says: 1. the literal keyword `order`, 2. then an identifier assigned to a `name` feature, 3. then the literal `for`, 4. then a cross-reference to a `Customer` element assigned to `customer`, 5. then zero or more `OrderLine`s collected into a list-valued `items` feature. This single line does two jobs at once: it tells the parser exactly what text is legal, and it implicitly defines a class `Order` with a string `name`, a reference `customer`, and a list `items` - which is exactly what a generative tool like Xtext turns into an actual typed metamodel class behind the scenes, with no separate hand-written AST needed. ## Why grammars exist Grammars exist because a DSL's usefulness depends entirely on precisely, unambiguously distinguishing valid from invalid input and giving a mechanical way to turn valid input into something a program can act on. Without a formal grammar, you're left hand-parsing strings with regexes or ad hoc string splitting, which is brittle, hard to extend, and gives terrible error messages. A grammar-first approach also means the entire downstream toolchain - parser, in-memory model, and later validators and generators - is derived from and stays consistent with a single source of truth, instead of drifting out of sync as the language evolves. ## The trade-off The major trade-off in grammar design is **expressiveness versus parseability**. - **Ambiguity.** A grammar with too much freedom (for instance letting the same token sequence be interpreted two different ways depending on later context) can become ambiguous, meaning the parser generator cannot determine, using only a bounded lookahead, which alternative to commit to. - **Bounded lookahead.** Tools like Xtext, built on ANTLR, typically require the grammar to be LL(*)-parseable: the parser must be able to decide which production to take by looking ahead a finite (if sometimes large) number of tokens, without backtracking arbitrarily. - **Separating concrete from abstract.** Designing a clean grammar also means separating concrete syntax (what keywords and punctuation you chose) from the underlying abstract concepts (what data the language actually needs to capture) - a discipline that pays off later, because if you decide to change surface syntax (rename a keyword, reorder clauses) without changing the underlying concepts, existing generators and validators built against the metamodel don't need to change at all. ## Failure modes 1. **Ambiguity as the language grows.** The most common failure mode in practice is exactly the ambiguity case: as a language grows and new alternatives get added to existing rules, the parser generator starts reporting lookahead conflicts between rules that used to be unambiguous. This isn't a cosmetic warning - Xtext/ANTLR simply refuses to generate a working parser, or silently picks a resolution order that may not match the author's intent, until the grammar is refactored (extracting common prefixes, adding syntactic predicates, or restructuring alternatives). 2. **Metamodel churn.** A second, subtler failure mode is that because the metamodel is derived automatically from the grammar, a seemingly small grammar tweak (renaming an assignment, changing a rule from single-valued to a list) silently changes the shape of every downstream artifact - validators, generators, and any previously saved DSL documents - which can break existing model instances or generator code with no compiler-level warning that ties the two together. ## The whole pipeline, end to end A concrete, well-known real-world instance of this whole pipeline is the Xtext framework itself, used to build languages such as Xtend (a Java-targeting DSL) and countless internal enterprise configuration languages: a team writes one `.xtext` grammar file, and Xtext generates the ANTLR-based parser, the EMF/Ecore metamodel, a default serializer, and a working Eclipse or VS Code editor with syntax highlighting and basic content-assist, all from that single grammar - meaning grammar design decisions ripple directly and immediately into editor usability, not just parsing correctness.
- What's the difference between abstract syntax and concrete syntax here?Concrete syntax is the actual textual notation the grammar rules describe - the specific keywords, punctuation, and layout a user types. Abstract syntax is the underlying typed model (the metamodel, with classes and features) that the concrete syntax populates once parsed. Multiple different concrete syntaxes could, in principle, map onto the same abstract syntax, which is why keeping them conceptually separate makes the language easier to evolve.
- What happens when your grammar becomes ambiguous as you add a new rule alternative?The parser generator reports a lookahead conflict - for example an LL(*) conflict in Xtext/ANTLR - because it cannot decide, within its lookahead budget, which alternative production to commit to for some input. You then have to refactor the grammar: factor out a common prefix, reorder alternatives, or add an explicit disambiguating token, since the tool cannot guess the author's intent.
- How do grammar rules translate into editor features like autocomplete, without extra work?Because Xtext derives both a metamodel and a partial parser from the grammar, it can, at any cursor position, compute the set of grammar elements that would be syntactically valid next and offer them as content-assist proposals - so basic autocomplete comes directly from the grammar definition with no separately hand-written completion engine required.
A DSL grammar is like a blueprint for LEGO instructions: it specifies which pieces can connect to which, in what order and how many times, and a good grammar tool builds both the sorting bins (the typed data model) and the assembly guide (the editor's autocomplete and highlighting) directly from that same blueprint.
saying these in an interview costs you the question
- Conflates a grammar with a database schema with no notion of ordered syntax at all
- Thinks the AST/metamodel classes must be hand-written separately from the grammar
- Doesn't recognize ambiguity/lookahead conflicts as a real, common failure mode
- Believes editor autocomplete requires writing a bespoke completion engine by hand
- Cannot distinguish a parse tree from the derived typed model