skip to content

The DOM Tree and Node Types

You will learn what a DOM node actually is and how the tree is typed — elements, text, comments, fragments — plus the live-vs-static collection trap. Interviewers ask because candidates who only know querySelector get surprised the first time a collection mutates underneath a loop.

on this pageshow

questions

5

In the browser DOM, what is the difference between a Node and an Element, and which kinds of node besides elements does a parsed HTML document contain?

level: juniorimportance: must knowfreq 55%

answer

  1. everything in the tree shares a base type
  2. tags are only one kind of node
  3. whitespace between tags is not free
  4. small integers name the kinds
  5. uppercase surprise on nodeName

basics

~20 s

Node is the base type for everything in the DOM tree; Element is the subtype for tags. A parsed HTML document also contains Text nodes (including whitespace between tags), Comment nodes, the Document itself, and its DocumentType node.

solid answer

~40 s

`Node` is the base interface every participant in the tree implements, and it carries `nodeType`, `nodeName`, `nodeValue`, and the parent/child links. `Element` is one subtype of `Node` — the one that corresponds to a tag and adds `tagName`, attributes, `classList`, `id` and the query methods. The other node types you actually meet in an HTML page are `Text` (`nodeType` 3, `nodeName` `"#text"`), `Comment` (8, `"#comment"`), `Document` (9, `"#document"`), `DocumentType` (10, the `<!DOCTYPE html>` node) and `DocumentFragment` (11). The practical consequence is that formatted markup produces text nodes for the whitespace between tags, so a node-level walk over a pretty-printed list hits `"#text"` entries you did not write, and only element nodes have `classList` or `getAttribute`.

code

javascript · 12 lines
javascript
const p = document.createElement('p');
p.innerHTML = 'Hi <!-- note --><b>there</b>';

for (const n of p.childNodes) {
  console.log(n.nodeType, JSON.stringify(n.nodeName), JSON.stringify(n.nodeValue));
}
// 3 "#text"    "Hi "
// 8 "#comment" " note "
// 1 "B"        null

console.log(document.nodeType);              // 9
console.log(document.doctype?.nodeName);     // "html"

go deeper

for a junior

Be ready to say that every tag, text run and comment in the page is a node, that Element is the node kind for tags, and that only elements carry attributes and classList.

for a middle

Explain the type table — nodeType 1/3/8/9/10/11 and their nodeName values — and why pretty-printed markup produces text nodes that a node-level walk must filter out.

for a senior

Show where this bites in production: hand-rolled tree walks and serializers that break on formatter-introduced text nodes, and code that treats nodeName string comparison as a reliable tag test across HTML and XML documents.

for a principal

Frame the guidance for a codebase: prefer element-level APIs and explicit nodeType guards over ad-hoc walks, and decide where the team is allowed to traverse raw nodes at all versus going through a helper that normalises node kinds.

## The interface hierarchy The DOM is a tree of objects, and those objects share a base interface: ``` EventTarget └── Node ├── Document nodeType 9 ├── DocumentType nodeType 10 ├── DocumentFragment nodeType 11 ├── Element nodeType 1 └── CharacterData ├── Text nodeType 3 └── Comment nodeType 8 ``` `EventTarget` at the top is why every node can take an event listener. `Node` adds the things that make an object part of a tree — the parent/child links, `nodeType`, `nodeName`, `nodeValue`, `textContent`, `ownerDocument`. `Element` is just one branch: the branch that corresponds to a tag written in the markup. So "every Element is a Node, not every Node is an Element" is the one-line answer, and the interview follow-up is *what the other nodes are*. ## nodeType and nodeName `nodeType` is a small integer with named constants on `Node`: | node | nodeType | constant | nodeName | |---|---|---|---| | `<p>` | 1 | `Node.ELEMENT_NODE` | `"P"` | | text run | 3 | `Node.TEXT_NODE` | `"#text"` | | `<!-- c -->` | 8 | `Node.COMMENT_NODE` | `"#comment"` | | `document` | 9 | `Node.DOCUMENT_NODE` | `"#document"` | | `<!DOCTYPE html>` | 10 | `Node.DOCUMENT_TYPE_NODE` | `"html"` | | fragment | 11 | `Node.DOCUMENT_FRAGMENT_NODE` | `"#document-fragment"` | Two details worth knowing. For elements in an HTML document, `nodeName` is the tag name **uppercased** — `"DIV"`, not `"div"` — which is why tag comparisons written as `node.nodeName === 'div'` silently never match. And `nodeValue` is the text for `Text` and `Comment` nodes but `null` for elements; the text of an element is reached through `textContent`, which concatenates the text of the whole subtree. Attributes are a special case. In the current DOM standard `Attr` inherits from `Node` with `nodeType` 2, but attributes are *not* children of the element — they hang off it separately and never appear in a child list. ## Where the surprise text nodes come from The parser turns every run of character data between tags into a `Text` node, and whitespace is character data: ```html <ul> <li>one</li> </ul> ``` That `<ul>` has three child nodes, not one: a text node holding the newline and two spaces, the `<li>` element, and another text node holding the trailing newline. Minified markup would produce only the element. This is the single most common junior stumble in this area — a node-level walk that assumes "the first child is the first tag" breaks the moment someone runs a formatter over the HTML. Comments are real nodes too. They are invisible to the user, they hold their text in `nodeValue`, and a walk that does not filter node types will visit them. ## Element is where the HTML API lives Everything you think of as "DOM manipulation" is on `Element` (or `HTMLElement` below it), not on `Node`: `tagName`, `id`, `className`, `classList`, `getAttribute`/`setAttribute`, `matches`, `closest`, `getBoundingClientRect`. A `Text` node has none of them. That is exactly why the DOM offers element-filtered accessors alongside node-level ones — the element-level view skips text and comment nodes so you do not have to filter by hand. When you *do* need to filter by hand, compare against the constant rather than the literal: ```js if (node.nodeType === Node.ELEMENT_NODE) { node.classList.add('seen'); // safe: only elements have classList } ``` ## Document, DocumentType, DocumentFragment `document` itself is a node (type 9) sitting above `<html>`; `document.documentElement` is the `<html>` element and `document.body` the `<body>` element. `document.doctype` is a `DocumentType` node whose `name` is `"html"` for a modern page, or `null` if the markup had no doctype. A `DocumentFragment` is a node that owns children but has no parent and is never rendered — it is a lightweight, parentless container you can build a subtree in. ## What an interviewer is checking They want to see that you think in terms of a typed tree rather than "a bunch of tags". The give-away answers are: naming `Text` and `Comment` as first-class nodes, knowing that whitespace in formatted markup becomes text nodes, knowing that `nodeName` is uppercase for HTML elements, and knowing that element-only APIs such as `classList` simply do not exist on the other node types.

  • Why does node.nodeName === 'div' never match a div in an HTML page?
    Because `nodeName` for an element in an HTML document is the tag name uppercased, so it is `"DIV"`. Compare case-insensitively, compare against the uppercase literal, or better, test `node.nodeType === Node.ELEMENT_NODE` and then use `node.tagName` or `node.matches('div')`, which is case-insensitive for HTML elements.
  • What does nodeValue return for an element, and how do you read its text instead?
    `nodeValue` is `null` for elements — it only carries data for `Text` and `Comment` nodes. Element text comes from `textContent`, which concatenates the text of every descendant text node, or from a specific child text node's `nodeValue` if you need just that run.
  • How does whitespace in your HTML source change the node structure you walk?
    Each run of whitespace between tags becomes its own `Text` node, so a formatted `<ul>` has text nodes surrounding every `<li>`. The same markup minified has none. Any node-level walk must therefore filter by `nodeType` or use the element-only view, or it will process nodes that came from the code formatter.

saying these in an interview costs you the question

  • Uses Node and Element as interchangeable words
  • Thinks whitespace between tags produces no nodes
  • Assumes nodeName for a div is lowercase 'div'
  • Expects classList or getAttribute to exist on a text node
  • Believes comments are stripped and absent from the tree

context

open as a page

In the browser DOM, what is the difference between the live HTMLCollection returned by document.getElementsByClassName() and the static NodeList returned by document.querySelectorAll(), and how does that difference break a loop that removes elements?

level: middleimportance: must knowfreq 65%

basics

~20 s

getElementsByClassName returns a live HTMLCollection that keeps re-reflecting the document, while querySelectorAll returns a static NodeList captured once. Looping forward over the live collection while removing elements shifts the indexes and skips every other match.

open as a page

In the browser DOM, when do node.parentNode and node.parentElement return different things, and which should a loop that climbs toward the root use?

level: middleimportance: should knowfreq 35%

basics

~20 s

They agree whenever the parent is an element. They differ when the parent is not one: for the <html> element, parentNode is the Document but parentElement is null, and inside a DocumentFragment parentNode is the fragment while parentElement is null. Climbing loops should use parentElement.

open as a page

A legacy page renders with subtly wrong element widths, and document.compatMode reports "BackCompat". What has the browser done, what in the HTML caused it, and can you switch it at runtime?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

"BackCompat" means the browser parsed the page in quirks mode and is applying legacy layout rules, including the old box model where width includes padding and border. A missing or unrecognised DOCTYPE causes it, and the mode is fixed at parse time.

open as a page

You take a node out of a same-origin iframe with iframe.contentDocument.querySelector('.row') and append it into the parent page. What happens to the node's ownerDocument, and why can `row instanceof HTMLElement` be false in the parent page?

level: seniorimportance: nice to knowfreq 15%

basics

~20 s

Inserting the node adopts it: the browser changes its ownerDocument to the parent page's document and the append succeeds. Adoption does not change the object's prototype, which still comes from the iframe's realm, so instanceof against the parent window's HTMLElement fails.

open as a page