Walk through how a browser turns an incoming stream of HTML bytes into a DOM tree, and explain why it does not wait for the whole document to arrive first.
answer
- two stages, not one
- a state machine over characters
- insertion modes and a stack
- malformed markup is repaired, never fatal
- chunks are parsed as they land
basics
~20 sBytes are decoded to characters, a tokenizer state machine turns those into start-tag, end-tag, text and comment tokens, and a tree constructor turns tokens into nodes and attaches them to the DOM. It runs on each chunk as it arrives so the page can render before the document ends.
solid answer
~50 sThere are two stages. First the response bytes are decoded into characters using the declared encoding, and a tokenizer — a state machine over that character stream — emits tokens: start tags, end tags, character data, comments, doctype. Second, a tree construction stage feeds those tokens through a set of *insertion modes* that decide where each node goes, inserting implied elements like `<tbody>` and closing unclosed ones as needed. The result is the DOM. It is streaming by design: the parser runs over each network chunk as it lands rather than buffering the whole document, so the DOM grows incrementally and the browser can compute style and paint a partially built page long before the last byte arrives. HTML parsing also has no fatal errors — malformed markup is recovered from by spec-defined rules, so every byte sequence produces some DOM.
code
html · 8 lines<!-- What you write -->
<table>
<tr><td>one</td></tr>
</table>
<!-- What the tree construction stage actually builds:
table > tbody > tr > td
The tbody element was never in the source. -->go deeper
Be able to say the browser reads the HTML and builds a tree of nodes as it goes, and that a script in the head cannot find elements that the parser has not reached yet.
Name both stages and what each produces: a tokenizer state machine emitting tags and text, then tree construction with insertion modes that inserts implied elements and repairs bad markup. Explain why parsing is chunk-by-chunk.
Connect incremental parsing to real delivery decisions — streaming versus buffered responses, early flush of the head — and diagnose a DOM-versus-source mismatch as parser repair rather than a framework bug.
Own the architectural consequence: whether the product's rendering strategy exploits the streaming parser at all, and what a team gives up by shipping an empty shell that builds every node from script after the parse has finished.
## Stage 0: bytes to characters The network delivers bytes. Before anything can be tokenized the browser must know the encoding, which it takes from the `Content-Type` response header, a byte order mark, or a `<meta charset>` in the first part of the document. If it guesses wrong and then finds a `charset` declaration later, it may have to restart the parse — which is the practical reason `<meta charset="utf-8">` belongs in the very first bytes of `<head>`. ## Stage 1: tokenization The tokenizer is a state machine. It walks the character stream and, depending on its current state, emits tokens: - doctype - start tag (with its attributes) - end tag - character (text) - comment - end-of-file Seeing `<` moves it into tag-open state; a letter after that moves it into tag-name state; `>` emits the token and returns to data state. The state machine is why context matters so much: inside `<script>` or `<style>` the tokenizer switches to a raw-text state where `<div>` is just characters, not a tag. That is the mechanism behind the classic hazard of an unescaped `</script>` inside a string in an inline script — it ends the element, because the tokenizer is not parsing JavaScript, it is scanning for that sequence. ## Stage 2: tree construction Tokens are handed to the tree construction stage, which is where the DOM appears. It keeps a stack of open elements and operates in an *insertion mode* — "before head", "in head", "in body", "in table", and so on — and each mode has rules for each token type. This stage is far more than a bracket matcher. It: - **Inserts implied elements.** A document with no `<html>`, `<head>` or `<body>` tag still gets all three. A `<tr>` written directly inside `<table>` gets a `<tbody>` created around it, which is why the DOM you inspect rarely matches the HTML you wrote for tables. - **Closes unclosed elements.** A `<li>` still open when the next `<li>` starts is implicitly closed. - **Relocates stray content.** Text or elements that are not allowed inside `<table>` are moved out before it — the "foster parenting" rule — so the node ends up as a previous sibling of the table rather than inside it. - **Reconstructs active formatting elements**, which is how overlapping tags such as `<b><i></b></i>` are repaired into a valid tree. ## No fatal errors Unlike XML, HTML parsing never aborts. Every possible byte sequence has a defined outcome, and browsers agree on it because the recovery rules are specified, not invented per vendor. That is a feature — the web is full of broken markup and a blank page for a stray tag would be unacceptable — but it means a browser silently building a tree you did not intend is a normal failure mode. When a component renders in the wrong place, comparing the source with the DOM inspector is the first diagnostic. ## Why it is incremental The parser is fed each chunk of the response as it arrives; it does not buffer the document. Everything downstream inherits that property: nodes appear in the DOM, style can be computed for them, and the browser can lay out and paint what it has. A user can see and interact with the top of a long page while the server is still generating the bottom. Two consequences follow directly. **Streaming server responses pay off.** If a server flushes the head and the top of the page immediately and streams the rest, the browser can start fetching subresources and painting during the time the server spends assembling slow content. Buffering the whole response until it is complete throws that overlap away. **The DOM you observe mid-parse is provisional.** A script that runs during parsing sees only the nodes above it. Querying for an element that appears later in the document returns `null` — not because the selector is wrong, but because the node does not exist yet. This is the mechanical reason initialisation code either sits at the end of the document or waits for parsing to complete. ```js // In a script placed in <head>, before the element exists in the document: document.querySelector('#app'); // null — the parser has not reached it ``` ## What to say in an interview Name the two stages and use the right words — tokenizer and tree construction — then make the incremental point, because that is what the rest of the loading story hangs on: a streaming parser is what lets the browser overlap network, parsing and painting instead of doing them in sequence.
- Why does the DOM inspector show a tbody element you never wrote?Tree construction inserts implied elements required by the content model. A `tr` token in the "in table" insertion mode causes a `tbody` to be created and pushed onto the stack of open elements first, so the row has a legal parent. The DOM is the parser's output, not a copy of your source.
- Why can an inline script's string containing the characters </script> break the page?Because inside a script element the tokenizer is in a raw-text state, scanning for the end-tag sequence rather than parsing JavaScript. It has no idea the sequence is inside a string literal, so it ends the element there and treats the rest of your code as markup. Escaping the slash avoids it.
- What does incremental parsing buy a server that streams its HTML?Overlap. While the server is still assembling slow parts of the page, the browser is already parsing the head, discovering subresources, and painting the top of the document. If the server buffers the full response instead, all of that work is pushed behind the slowest query.
saying these in an interview costs you the question
- Says malformed HTML makes the parse fail like XML
- Claims the whole document is buffered before parsing
- Thinks the DOM is a literal one-to-one copy of the source
- Cannot separate tokenization from tree construction
- Assumes a head script can query elements further down the page