skip to content

Parsing Model and Error Recovery

HTML never throws — the parser silently repairs your markup, which is why a stray tag can produce a DOM you did not write. Interviewers ask why a <div> inside a <p> ends up as a sibling, and the answer is the tree-construction algorithm.

part ofHTMLoverview, primer and where to startread it →
on this pageshow

questions

5

Given the HTML `<p>Intro <div>Block</div></p>`, what DOM does the browser build, and why is the div not a child of the p?

level: middleimportance: must knowfreq 55%

answer

  1. not a browser bug — a content model
  2. p holds phrasing content only
  3. a start tag can imply an end tag
  4. the stray end tag leaves something behind

basics

~20 s

A p element cannot contain a div, so the div's start tag implicitly closes the open p. The DOM becomes a p holding "Intro", a sibling div holding "Block", and an empty p created by the now-stray end tag.

solid answer

~50 s

The `p` element's content model allows only phrasing content, and `div` is not phrasing content, so the parser never even tries to nest them. During tree construction, a `<div>` start tag checks whether a `p` element is open in button scope and, if so, closes it first — the p's end tag is *implied*. That leaves `<p>Intro </p><div>Block</div>` in the DOM. Then the `</p>` you wrote arrives with no p on the stack of open elements: that is a parse error whose prescribed recovery is to insert a fresh `p` element with no attributes and immediately close it, so you also get an empty `<p></p>` you never typed. Nothing here is a browser bug or a rendering quirk — every browser produces the same three nodes, because the recovery is specified. It is why a selector or a script written against `p > div` silently matches nothing.

code

html · 6 lines
html
<p>Intro <div>Block</div></p>

<div>
  <p>Intro</p>
  <div>Block</div>
</div>

go deeper

for a junior

Know that a paragraph may only contain text-level content, so putting a div inside <p> does not work, and reach for a div or section as the wrapper instead.

for a middle

Explain the mechanism: the div start tag implies the paragraph's end tag because of p's content model, and the leftover </p> materialises an empty paragraph. Be able to predict the resulting node list.

for a senior

Talk about how this reaches production through templates and CMS-injected rich text, what breaks downstream when the DOM shape differs from the source, and how you would spot it during triage.

for a principal

Frame it as a contract problem: which components are allowed to emit block-level content, how that is enforced at the template or validation layer, and why relying on authors to remember content models does not scale.

## What you wrote and what you get Source: ```html <p>Intro <div>Block</div></p> ``` Parsed DOM: ```html <p>Intro </p> <div>Block</div> <p></p> ``` Three children where you expected one, and an empty paragraph nobody asked for. This is the single most-asked question about HTML parsing, and the answer has two halves: an implied end tag, and a stray end tag. ## Half one: content models drive implied end tags Every HTML element declares what it is allowed to contain. The `p` element's content model is *phrasing content* — text and text-level elements such as `<span>`, `<a>`, `<em>`, `<strong>`, `<img>`, `<code>`. A `div` is flow content, not phrasing content, so `<div>` inside `<p>` is not merely discouraged, it is unrepresentable. The HTML parser encodes this directly. In the *in body* insertion mode, the handler for a `<div>` start tag reads, in effect: if the stack of open elements has a `p` element in button scope, close that p element; then insert the div. The same rule is attached to a long list of start tags — `address`, `article`, `aside`, `blockquote`, `details`, `div`, `dl`, `fieldset`, `figcaption`, `figure`, `footer`, `form`, `h1`–`h6`, `header`, `hgroup`, `hr`, `main`, `menu`, `nav`, `ol`, `p`, `pre`, `section`, `table`, `ul` and others. That is why `<p>one<p>two` is valid markup for two paragraphs: the second `<p>` closes the first. Because the closing is implied by the *next start tag*, the paragraph's end tag is genuinely optional in HTML. Omitting it is conforming; the parser fills it in. ## Half two: the stray end tag By the time your `</p>` is tokenized, the div has already been closed by `</div>` and there is no `p` on the stack. The specification's rule for an end tag `p` with no p in button scope is: this is a parse error; insert an HTML element for a `p` start tag with no attributes, then close a p element. The result is a real, empty `<p></p>` node in the document. This surprises people because the intuitive recovery would be "ignore it". HTML chose "materialise it" for historical compatibility, and — crucially — chose *something*, so that every engine agrees. ## The same mechanism elsewhere Implied end tags are not a paragraph special case. They are how HTML has always allowed terse markup: ```html <ul> <li>one <li>two </ul> <dl> <dt>term <dd>definition </dl> <select> <option>a <option>b </select> ``` Each `<li>` closes an open `li`; `<dd>` closes an open `dt`; a new `<option>` closes the previous one; the container's end tag closes the last child. All three produce flat sibling lists, exactly as intended. ## Why it matters in real work The bug usually arrives through templating. A component renders "a paragraph of intro text" and a caller drops a block-level widget inside it, or a rich-text field from a CMS injects `<div>` or `<figure>` into a paragraph wrapper. The page still looks roughly right, so nobody notices until: - a selector or query written for `p > div`, or a wrapper style scoped to the paragraph, stops applying to the moved node; - an empty paragraph appears in the layout and someone "fixes" it with a rule that hides empty paragraphs; - assistive-technology output changes, because the structure being announced is not the structure the author drew. The fix is never to fight the parser. It is to choose a container whose content model accepts what you are putting in it — a `div`, a `section`, or a `figure` — and to reserve `<p>` for actual paragraphs of phrasing content. ## How to diagnose it in seconds View source shows the bytes you shipped; the inspected element tree shows the DOM the parser built. When those two disagree, an implied end tag is nearly always the reason. Confirming that they disagree is faster than reasoning about it, and it is the habit worth demonstrating in an interview. ## What a strong answer sounds like Name the content model, name the implied end tag, predict the empty `p`, and finish with "and every browser does this identically because the recovery is specified". That last clause is what tells an interviewer you understand HTML error recovery as a designed algorithm rather than as browser folklore.

  • Which other elements get closed automatically by the start tag that follows them?
    `li` is closed by the next `<li>`, `dt` and `dd` by the next `<dt>` or `<dd>`, `option` by the next `<option>` or `<optgroup>`, and `td`, `th` and `tr` by the next cell or row. All of their end tags are optional in HTML, which is why `<ul><li>one<li>two</ul>` is conforming and produces two sibling list items rather than nested ones.
  • Does the same repair happen when markup is parsed as a fragment rather than as a whole document?
    Yes. Fragment parsing runs the same tokenizer and the same tree-construction rules, with a context element supplying the insertion mode. So a string containing `<p><div>…</div></p>` is repaired identically wherever it is parsed. The repair belongs to HTML's parsing algorithm, not to how the markup reached the parser.
  • How would you confirm this is happening on a page you did not write?
    Compare the served source with the built element tree. The source shows the bytes; the inspected tree shows what tree construction produced. If the div sits beside the paragraph rather than inside it, and an empty paragraph trails it, you are looking at an implied end tag plus stray-end-tag recovery, not at a styling problem.

saying these in an interview costs you the question

  • Says the browser is buggy or that only Chrome does this
  • Claims the div ends up as a child of the p
  • Thinks any element may nest inside any other
  • Blames CSS when a p > div selector stops matching
  • Assumes the stray </p> is simply discarded

context

open as a page

In an HTML document, what does the trailing slash in `<br />` and `<div />` mean to the HTML parser, and which elements can actually close themselves?

level: juniorimportance: should knowfreq 50%

basics

~20 s

In HTML the trailing slash is ignored. Void elements such as br, img and input never take an end tag and close themselves anyway; on an ordinary element such as div the slash does nothing, so <div /> stays open.

open as a page

You write `<table><tr><td>1</td></tr></table>` with no `<tbody>` in the source. What does the parsed DOM contain that your markup does not, and why?

level: middleimportance: should knowfreq 45%

basics

~20 s

The parser inserts a tbody element for you, so the DOM is table > tbody > tr > td. A row start tag inside a table implies a tbody start tag, which means table rows are never direct children of the table element.

open as a page

A templating partial ships a `<div>` whose end tag is missing. HTML has no fatal parse errors — what does the browser do with the markup that follows, and how would you catch this class of bug before release?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Nothing fails. The div stays on the parser's stack of open elements, so following markup becomes its descendants until an ancestor's end tag or end of file closes it implicitly. Catch it by validating rendered pages in CI, not by eye.

open as a page

How does the HTML parser treat markup inside an inline `<svg>` element differently from ordinary HTML, and what happens if a `<p>` start tag appears inside it?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

Inside <svg> the parser switches to foreign-content rules: a trailing slash genuinely self-closes an element, and known attribute names are case-corrected to spellings such as viewBox. A <p> start tag forces a breakout, popping the parser back into HTML.

open as a page