Why does one stray token make a parser report fifty syntax errors, and how is that cascade suppressed?
answer
- one mistake, reported fifty times
- state disagrees with the text
- quiet window after each report
- deduplicate by position and message
- cap the list, say it was capped
basics
~10 sAfter a failure the parser's state no longer matches the text, so every following token disagrees and becomes another message. Suppression combines real resynchronisation, a quiet window after each report, deduplication, and a cap.
solid answer
~50 sA cascade is one mistake reported many times, not many mistakes. Once a rule has failed, the parser's stack describes a structure the source does not have, so each subsequent token conflicts with it and — if reporting stays open — earns its own message. Four mechanisms suppress this. Recovery that lands on a genuine construct boundary is the main one, because errors reported after a real resynchronisation are usually independent. On top sits a **quiet window**: nothing is reported within a short distance of the previous message, so repairs still happen but silently. Then **deduplication**, since the same message at the same position is always noise. Finally a **cap**, because past some count the tail is worthless and pushes the one useful message out of view. The cost is that a genuine second mistake close to the first can be hidden.
code
pseudocode · 15 lineslast_reported_at = MINUS_INFINITY
QUIET = 3 // tokens of silence after a report
reported = 0
on parse_error(token):
if token.index <= last_reported_at + QUIET or reported >= MAX_ERRORS:
recover(token) // repair or skip, but emit nothing
return
report(token, expected_set_of(current_state))
last_reported_at = token.index
reported = reported + 1
recover(token)
on resynchronised_at_anchor():
last_reported_at = MINUS_INFINITY // the parser is trustworthy againgo deeper
Know that a long list of syntax errors usually comes from a single mistake near the top, and that the first message is the one to act on.
Explain why a desynchronised parser disagrees with every following token, and name the suppressors: real resynchronisation first, then a quiet window, deduplication and a cap.
Show how you would tune and verify suppression — a window measured in tokens, a corpus with two known independent mistakes, and later phases gated on error nodes so they do not cascade in turn.
Own the trade-off between a calm list and a complete one: how much the team is willing to hide to keep the first message trustworthy, and what the tool must say when it truncates.
## One mistake wearing fifty hats The parser's stack is a claim about the shape of the text. The moment a rule fails, that claim is false: the parser believes it is somewhere the source is not. Every token after the failure is now evaluated against the wrong expectation, so almost all of them look wrong too. If the parser is allowed to report freely, the result is a list whose first line describes the author's mistake and whose next forty-nine describe the parser's confusion. Users learn this the hard way and develop the habit of fixing only the first error and re-running — which is a reasonable heuristic against a tool that has not solved the problem. The cascade is worth distinguishing from a *genuine* multi-error file, because the cure for one is fatal for the other. A file with three independent typos should produce three messages. A file with one typo should not produce fifty. Suppression that cannot tell these apart simply trades one bad experience for another. ## The four suppressors 1. **Resynchronise properly.** The strongest suppressor is recovery that reaches a real construct boundary, because from there the remaining input is parsed on its own terms and later failures are genuinely independent. Everything below is a safety net for when recovery lands badly. 2. **A quiet window.** After a report, suppress any further message within a short distance — measured in tokens, not lines or characters, since a line is an arbitrary amount of syntax. Recovery and repair still run inside the window; only reporting is silenced. 3. **Deduplication.** Identical messages at the same position are always noise, and near-identical messages at nearby positions usually are. Collapsing them costs nothing and is safe. 4. **A cap.** Past a bounded count, stop reporting and say so explicitly. An unbounded list buries the one line the reader needs, and a truncated list that announces the truncation is strictly more honest than one that silently stops. ## Tuning the quiet window - **Measure in tokens.** A window of a few tokens kills same-construct repeats while leaving a mistake on the next statement visible. - **Reset on a real boundary.** When recovery reaches an anchor, the window can be cleared: the parser is trustworthy again. - **Verify on a corpus.** Build files with two independent, known mistakes and check both are still reported. This is the only test that distinguishes a well-tuned window from a wide one that hides everything. - **Never widen the window to fix a noisy parser.** If suppression is doing most of the work, the real problem is a recovery strategy that is not reaching construct boundaries. ## What suppression costs | Mechanism | What it removes | What it can hide | |---|---|---| | Proper resynchronisation | The whole cascade, at its source | Nothing, when the anchor set is right; a wrong anchor loses structure instead | | Quiet window | Repeats inside the damaged region | A genuine second mistake that sits close to the first | | Deduplication | Identical or near-identical repeats | Two distinct problems that happen to read alike | | Message cap | The worthless tail of a long list | Real errors past the cap, which is why the truncation must be announced | ## Cascades after parsing The same effect repeats downstream. If the parser fabricated tokens or recorded error nodes, any later phase that walks the tree will encounter structure the author never wrote, and reporting on it produces a second cascade whose messages look authoritative but describe the parser's guess. The rule that prevents it is simple and worth stating explicitly: later phases may **run** over a damaged tree, but must not **report** diagnostics whose position falls inside an error node or derives from a synthetic token. In practice a single flag — does this tree contain any error node — decides whether phases that produce artefacts run at all, and per-node positions decide which individual messages are dropped. ## Reading an error list as a diagnostician When a user shows you fifty syntax errors, the first question is not which one is real but whether the list is a cascade at all. Two symptoms give it away: the messages march forward in source order with no gaps, and they describe increasingly implausible things — a parser complaining about the top-level shape of a file halfway down a function is not describing anything the author did.
- How do you choose the size of the quiet window?By what it costs when it is wrong. Measure it in tokens, not lines, and start with a few — enough to kill same-construct repeats, not enough to swallow the next statement. Then validate on a corpus of files carrying two independent, known mistakes: if the second one disappears, the window is too wide, whatever the noise reduction looks like.
- Why cap the total number of messages instead of letting the list run?Because the tail of a cascade is worthless and pushes the useful first message out of a terminal or a log. Stopping at a bounded count and stating that further errors were suppressed keeps the signal visible and, crucially, tells the reader the file has not been fully checked — which a silently truncated list does not.
- What should phases after parsing do with a subtree the parser marked as an error node?Run if they like, but suppress any diagnostic whose position falls inside it. Structure there is the parser's guess, not the author's text, so messages computed over it describe the guess. Phases that emit artefacts should refuse outright whenever the tree contains any error node at all.
saying these in an interview costs you the question
- Reads the error count as a count of real mistakes
- Believes every message after the first names a distinct problem
- Thinks the only cure is to stop after the very first error
- Sets a quiet window so wide that independent mistakes vanish
- Lets later phases report diagnostics derived from a guessed subtree
- Truncates the list without telling the reader it was truncated