Why can a script that parses with no syntax error still be rejected by a later compiler pass?
answer
- shape settled, meaning not yet
- the pass that carries a symbol table
- resolution errors then agreement errors
- two distant points must be correlated
- names are unbounded and author-created
basics
~20 sBecause parsing only proves the text has a legal shape. The pass after it checks meaning against a symbol table — that every name resolves, that it is used as the kind of thing it was declared, and that operand types agree — and those depend on declarations arbitrarily far away.
solid answer
~50 sA parser answers one question: does this token sequence match the language's shape rules? A reference to an undeclared name has a perfectly legal shape, as does adding a duration to a file handle. Both are rejected afterwards, by the pass that carries a **symbol table**. That pass resolves each name to a declaration and then checks agreement: the name exists and is visible, it is used as the kind it was declared (a step invoked, not read as a value), the two sides of an assignment have compatible types, a call's arguments match the declared parameters in number and type. What makes these different in nature from shape rules is that each one correlates two points in the text that may be arbitrarily far apart, over an unbounded set of names the shape rules cannot enumerate. That is why the checks are done by walking the tree with a table rather than by matching the token stream.
go deeper
Recall the distinction in plain terms: one stage checks that the code is written in a legal shape, a later one checks that the names and types in it actually make sense together.
Explain the mechanism: the later pass carries a symbol table, resolves each name to a declaration, then compares kinds and types, which the shape rules cannot do because they hold no memory of names.
Demonstrate diagnostic judgment: separate resolution failures from agreement failures, explain error poisoning that stops one typo cascading, and account for why fixing a syntax error surfaces a new batch.
Consider where to draw the line for a language you own: which mistakes are worth rejecting at this pass versus deferring, and what each extra rejection costs in authoring friction and diagnostic quality.
## Two different questions about one program Compiling a program asks two questions in sequence, and they are not the same question: 1. **Is it shaped legally?** Are the tokens arranged the way the language's rules allow — does every opened block close, does a call have a name followed by a parenthesised argument list, does an assignment have exactly one target? This is what the parser settles, and a failure here is a **syntax error**. 2. **Does it mean anything coherent?** Does every name used correspond to a visible declaration, is each name used as the kind of thing it was declared to be, and do the types of things combined actually fit together? This is settled by the pass that walks the parsed tree carrying a symbol table, and a failure here is a **semantic error**. A script that fails question 2 while passing question 1 is completely ordinary. In a workflow-automation script, `retry_after = output_file` has textbook shape — identifier, assignment operator, identifier — and is nonsense if one names a duration and the other a file handle. Shape carries no information about what either name was declared as. ## What this pass actually compares The checks divide cleanly into two families, and mixing them up is a common interview stumble: **Resolution errors** — about whether a name means anything here: - the name has no visible declaration on the scope chain; - the name is declared later in a scope that requires declaration first; - the name is declared twice in one scope, so a reference to it would be ambiguous; - the name is used as the wrong kind of entity: a variable invoked as a step, a step read as a value. **Agreement errors** — about whether the resolved things fit together: - the two sides of an assignment have incompatible types; - a call passes the wrong number of arguments, or arguments whose types do not match the declared parameters; - a condition is given a value that is not usable as a truth value; - a result is used where a definition declares none is produced. Every item in the first family needs the symbol table before it can be evaluated at all. Every item in the second needs the first family to have finished, because the type of an operand *is* the type recorded on the declaration it resolved to. ## Why this is not folded into the shape rules The distinguishing property of these checks is that each correlates **two points in the program that can be arbitrarily far apart**, over a set of names that is unbounded and created by the program itself. Deciding whether `retry_after = output_file` is acceptable requires the declarations of both names — which may be hundreds of lines away, in enclosing scopes, in a different definition entirely — and requires remembering, for every name in scope, which kind and type it holds. The shape rules a parser matches carry no such memory: they describe the arrangement of tokens, not a growing table of facts about names invented at author time. So the pipeline splits the labour. The parser produces a tree cheaply from the shape rules alone; the pass after it walks that tree and maintains exactly the memory the shape rules cannot: a table per scope, chained, with an entry per declaration. That split is also what makes the diagnostics good, because the table holds declaration positions and a semantic error can therefore point at **both** ends of the disagreement — the use and the declaration it disagrees with. ## Consequences an engineer actually feels | Symptom | Which question failed | What the message can name | |---|---|---| | An unclosed block or a stray separator | Shape | The token position, and usually the construct expected there | | A misspelled variable name | Meaning (resolution) | The reference, plus near-miss names taken from the tables in scope | | An argument of the wrong type | Meaning (agreement) | The call site and the declared parameter's position | | Two declarations of one name in one block | Meaning (resolution) | Both declaration positions | Two practical points follow. First, a tool that reports only the first error in each category is doing something reasonable, not lazy: an unresolved name is bound to a poisoned entry that agrees with everything, precisely so the agreement checks do not emit a cascade caused by one typo. Second, the **order** of the messages you see is a pipeline artefact — every shape error is found before any meaning error can be evaluated, because a tree the parser could not build is a tree this pass cannot walk. Fixing the last syntax error in a file routinely reveals a fresh batch of semantic ones, and that is the pipeline working as designed rather than a regression.
- Why does fixing the last syntax error in a file often reveal a batch of new errors?Because the meaning checks can only run over a tree, and a file the parser could not finish produces no usable tree. Those errors were always there; they were simply unreachable. The ordering is a pipeline property, not a sign that the fix caused them.
- Why is a name used as the wrong kind of entity a resolution error rather than a type error?Because the entry records a kind — variable, parameter, step — alongside the type, and the mismatch is found the moment the reference resolves, before any type is compared. Invoking something declared as a value fails on kind even when its type would have been irrelevant.
- What does binding an unresolved name to a poisoned entry buy the run?One misspelling would otherwise disagree with every operand it touches, and each disagreement would be reported. A poisoned entry is defined to agree with anything, so the pass reports the missing declaration once and still surfaces the unrelated real problems elsewhere in the file.
saying these in an interview costs you the question
- calls every compiler rejection a syntax error
- thinks the parser knows what names are declared
- says type checking happens while tokens are being matched
- claims a well-formed tree means a valid program
- treats an undeclared name and a type mismatch as one check
- assumes later errors appearing after a fix means the fix broke something