An HTML page's <head> contains a stray <div> before its <meta> and <link> tags. What does the parser do at that point, and what stops working?
answer
- the head has a fixed list of allowed elements
- the parser never reports an error
- an end tag you did not write
- body opens earlier than the source suggests
- the tags are in the DOM, just not in head
basics
~20 sOnly a fixed set of elements may appear in <head>. On reaching a <div> the parser implicitly closes <head>, opens <body>, and puts the div and everything after it in the body — so metadata written later is no longer in the head, and a late encoding declaration is missed.
solid answer
~50 s`<head>` accepts only `base`, `link`, `meta`, `noscript`, `script`, `style`, `template` and `title` (plus whitespace and comments). When the parser meets anything else — a `<div>`, or even a stray non-whitespace character from a broken template — it does not error out. It acts as if `</head>` were written there, opens `<body>` implicitly, and everything from that point on becomes body content. The observable symptom is exactly what people report: "my meta tags disappeared from the head" — in DevTools they are sitting under `<body>`. What actually breaks depends on the tag. A `<link rel="stylesheet">` still applies, so the page looks fine and the bug hides. `<meta charset>` is the dangerous one: it is now almost certainly past the 1024-byte pre-scan window as well as out of `<head>`, so the encoding falls back to a guess. `<base>` after the split no longer precedes the URLs it was meant to govern.
code
html · 12 lines<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<title>Orders</title>
<div class="wrapper">
<meta name="description" content="Order history">
</head>
<body>
<p>Content</p>
</body>
</html>go deeper
Know that <head> holds only metadata elements — meta, title, link, style, script, base — and that page content belongs in <body>. Being able to spot misplaced markup is enough at this level.
Explain the recovery mechanism: an implied </head>, an implicitly opened <body>, and the later explicit tags being ignored. Say why comments and whitespace are safe but a div is not.
Diagnose it from symptoms. Reason about why styles still apply while the encoding declaration silently fails, and show how you would confirm the real head boundary in DevTools rather than trusting view-source.
Treat the head as an owned region with a single author and validation in CI, so independently-shipped template partials cannot append markup that silently relocates every tag beneath it into the body.
## The head's content model `<head>` is not a free-form container. Its content model is metadata content, which in practice means exactly these elements: - `<base>` — the document base URL - `<link>` — stylesheets, icons, resource hints, canonical - `<meta>` — charset, viewport, named metadata - `<noscript>` — but only containing link/style/meta - `<script>` - `<style>` - `<template>` - `<title>` — required, and exactly one Comments and whitespace are allowed anywhere and do **not** end the head. Anything else does. ## What the parser does — no error, a recovery HTML parsing never fails. The tree-construction stage is a state machine, and while it is in the "in head" state, a start tag it does not recognise as head content is handled by an *implied end tag*: the parser behaves as though `</head>` appeared, moves on, and then — still holding the unrecognised token — opens `<body>` and reprocesses the token there. ```html <head> <meta charset="utf-8"> <title>Orders</title> <div class="wrapper"> <!-- head ends here, implicitly --> <meta name="description" content="..."> <!-- this is in <body> now --> <link rel="stylesheet" href="/app.css"> <!-- so is this --> </head> <!-- ignored: head is already closed --> <body> <!-- ignored: body is already open --> ``` The explicit `</head>` and `<body>` tags further down are simply dropped, because the parser has already moved past those states. Nothing in the console tells you this happened. ## What breaks, tag by tag This is where the answer gets interesting, because the consequences are uneven — which is precisely why the bug survives review. - **`<link rel="stylesheet">`** still works. Browsers honour stylesheet links found in the body, so the page renders styled and looks correct. It does change when styles arrive relative to content, which can surface as a restyle of content already painted above. - **`<script>`** still runs. No visible symptom at all. - **`<meta charset>`** is the real casualty. It is now both outside `<head>` and, in any realistic document, beyond the 1024-byte window the browser pre-scans for an encoding declaration. The browser has already committed to a guessed encoding, and non-ASCII text can render as mojibake. - **`<base>`** is meant to precede every element carrying a URL. After an implicit body start there is usually markup above it that has already resolved against the document URL, so resolution is split across two bases. - **Validators and tooling** flag the document, and any tool that reads metadata by looking at `document.head` — or by parsing the head server-side — will miss tags that are technically in the body. ## How this ships in real projects Almost never by someone typing a `<div>` into a head on purpose. The usual causes are a templating partial that emits a wrapper element, a layout that concatenates fragments in the wrong order, a stray character or unescaped text from a CMS field interpolated into the head, or a mis-nested conditional that closes a tag early. Server-side rendering makes it worse, because the head is assembled from several independently-owned partials. ## How to find and prevent it The fastest diagnosis is the DevTools Elements panel: expand `<head>` and see where it actually ends, versus what the source says. `document.head.children` in the console gives the same answer programmatically. Beyond that, run the markup through an HTML validator in CI, and treat the head as an owned, small, single-authored region of the template rather than a place partials append to freely. ## What an interviewer is checking Two things. First, that you know HTML parsing is *error-recovering by specification*, not error-reporting — invalid markup produces a defined, silent DOM rather than a failure. Second, that you can reason about a bug whose most visible symptoms are absent: the page looks right, the styles apply, and only the encoding and the metadata-consuming tools are wrong.
- How would you confirm in the browser that the head ended earlier than the source says?Open the Elements panel and expand `<head>` — the rendered DOM shows where it actually closed, which will not match view-source. Programmatically, `document.head.children` lists what genuinely made it into the head, and comparing that list against the tags you expect pinpoints the first offending element.
- Why does the page still look correct when a stylesheet link ends up in the body this way?Browsers honour `<link rel="stylesheet">` wherever they find it, including in body, so the rules still apply. That is exactly why the bug is dangerous: the most visible check passes. What changes is when the styles arrive relative to painted content, and the metadata tags around the link are silently misplaced too.
- Do comments or blank lines in <head> risk closing it the same way?No. Comments and whitespace are permitted anywhere in the document and leave the parser in the same state, so they never trigger the implied `</head>`. It takes a start tag outside the metadata content model, or non-whitespace text, to end the head early.
saying these in an interview costs you the question
- Says the browser reports a parse error for invalid head content
- Thinks the stray element is simply dropped
- Believes an explicit </head> later in the source still applies
- Assumes nothing breaks because the page renders correctly
- Cannot name which elements <head> actually permits