skip to content

Why is the markup `<p>Intro <div>Block</div></p>` invalid HTML, and what DOM does the browser actually build from it?

level: middleimportance: should knowfreq 50%

answer

  1. p takes phrasing content only
  2. the start tag closes it implicitly
  3. the div ends up as a sibling
  4. a stray end tag leaves something empty
  5. span breaks the rule without reparenting

basics

~20 s

The p element accepts only phrasing content, and div is flow content, so the nesting is invalid. The parser closes the open p when it sees the div start tag, so the div becomes a sibling of the paragraph — and the trailing </p> produces a second, empty paragraph.

solid answer

~40 s

`p` has a content model of phrasing content — text and text-level elements — and `div` is flow content, so a `div` child is invalid. What matters more is that the parser does not fail; it repairs. Seeing a `<div>` start tag while a `p` element is open, the parser implicitly closes the paragraph, so the div is emitted as the paragraph's **sibling**, not its child. Then the stray `</p>` arrives with no open paragraph to close, and the parser inserts an empty `<p></p>` for it. You end up with three siblings: a paragraph containing "Intro", the div, and an empty paragraph. The page usually looks close enough that nobody notices — until a selector, a script, or a tree-diffing tool works against the structure you wrote rather than the one that exists.

code

html · 15 lines
html
<!-- authored -->
<p>Intro <div>Block</div></p>

<!-- resulting DOM: three siblings, including an empty paragraph -->
<!--
<p>Intro </p>
<div>Block</div>
<p></p>
-->

<!-- valid rewrite: the div is the container, the paragraph is inside it -->
<div>
  <p>Intro</p>
  <div>Block</div>
</div>

go deeper

for a junior

Know that a div may not go inside a p, and that the browser silently rearranges the markup rather than reporting an error, so the DOM can differ from what you typed.

for a middle

Explain both halves: p accepts only phrasing content so div is invalid, and the parser implicitly closes the paragraph at the div start tag, leaving the div as a sibling plus an empty paragraph from the stray end tag.

for a senior

Demonstrate how this surfaces in production — selectors and scripts that no longer match, unexplained spacing, server markup disagreeing with the rebuilt client DOM — and how you localise it by comparing source against the inspector.

for a principal

Own prevention at scale: validation in the pipeline, component slots typed so callers cannot pass flow content into a phrasing-only position, and a clear rule for which markup violations block a build.

## Why it is invalid The `p` element's content model in the HTML specification is **phrasing content**: text nodes and the text-level elements — `span`, `a`, `em`, `strong`, `code`, `img`, `br`, `time`, and so on. `div` is **flow content** and is not phrasing content, so it may not be a child of `p`. The same rule makes `<p><ul>…</ul></p>`, `<p><section>…</section></p>` and `<p><h2>…</h2></p>` invalid. Note the direction of the rule. It is not that a div can never sit near a paragraph — a `div` containing a `p` is perfectly valid, because `div` accepts flow content and `p` is flow content. Only the containment `p` → `div` is illegal. ## What the parser does instead of failing HTML has no fatal parse errors for this. The parser is specified to recover, and its recovery here is deterministic and worth knowing precisely. The `p` element is one of the elements the parser closes implicitly. When a `<div>` start tag arrives while a `p` element is open, the parser closes that paragraph first, then inserts the div. The div is therefore a **sibling** of the paragraph, not a descendant. The second half is the part candidates miss. After `</div>`, the `</p>` end tag arrives — but there is no open paragraph any more, because it was closed implicitly. The parser handles this end tag by creating a paragraph element for it and immediately closing it, which yields an **empty `<p></p>`** in the tree. So this source: ```html <p>Intro <div>Block</div></p> ``` produces this tree: ```html <p>Intro </p> <div>Block</div> <p></p> ``` Three siblings where the author wrote one nested structure. You can confirm it in any browser by inspecting the elements panel, or by reading `document.body.children.length`. ## The contrast that proves it is the content model, not "block inside inline" A `span` also accepts only phrasing content, so `<span><div>x</div></span>` is equally invalid. But the parser does **not** implicitly close a `span`, because `span` is not one of the elements subject to that recovery rule. The div simply ends up nested inside the span in the DOM. ```html <!-- invalid, and the parser reparents: div becomes a sibling --> <p>a <div>b</div></p> <!-- equally invalid, but the parser leaves it alone: div stays inside the span --> <span>a <div>b</div></span> ``` Two violations of the same content-model rule, two different repairs. Validity is one question; what the parser does about it is a separate one. ## Why the silent repair costs you Because the page usually still looks approximately right, the bug hides. The damage shows up wherever something walks the tree: - A descendant selector or a `querySelector` written as "the div inside that paragraph" matches nothing — it is not inside it any more. - Code that reads the paragraph's children, or its text, finds only the leading text. - Spacing looks off because there is now an extra empty paragraph in the flow. - Tooling that generates markup on a server and rebuilds the DOM on the client compares two structures that disagree, because one side went through the parser's repair and the other did not. That last case is the nastiest: the source template and the live DOM genuinely differ, so the mismatch is real, not a caching artefact. ## How to fix it Decide what the content actually is: - If the block content belongs with the intro text as one unit, wrap the whole thing in a `div` and keep the paragraph inside it: `<div><p>Intro</p><div>Block</div></div>`. - If the inner content is a run of text you want to mark up, use a `span`, which is phrasing content and legal inside `p`. - If the paragraph really contains several blocks, it was never a paragraph — a `p` holds one run of prose. ## Catching it Run markup through the Nu HTML validator, which names the violation directly ("element div not allowed as child of element p in this context"). When debugging, compare the source file with the elements panel: where they diverge, an invalid nesting is the usual cause. The tell for this specific bug is an empty paragraph you never wrote.

  • Where does the empty <p></p> in the resulting tree come from?
    From the trailing `</p>`. The paragraph was already closed implicitly when the `<div>` start tag arrived, so that end tag finds no open p. The parser's recovery for a `</p>` with no paragraph in scope is to create a p element and close it immediately, leaving an empty paragraph in the tree.
  • Is <span><div>x</div></span> repaired the same way?
    No. It breaks the same content-model rule — span accepts only phrasing content — but the parser does not implicitly close a span, so the div stays nested inside it in the DOM. That contrast is the point: invalidity is decided by the content model, while the repair depends on the parser's per-element rules.
  • Would <p><ul><li>one</li></ul></p> behave the same way?
    Yes, in shape. `ul` is flow content and not phrasing content, so it is invalid inside `p`, and a `<ul>` start tag also implicitly closes an open paragraph. You get the paragraph, then the list as its sibling, then an empty paragraph from the stray end tag.

saying these in an interview costs you the question

  • It is fine because browsers fix it anyway
  • Invalid means the browser throws a parse error
  • p cannot hold a div because div is a block element
  • The DOM matches the source, only validation complains
  • Adding display: inline to the div would make it valid

context