skip to content

In PHP 8.5, when do you choose SimpleXML, DOMDocument or the Dom\XMLDocument class added in 8.4 for reading XML?

level: middleimportance: must knowfreq 48%

answer

  1. convenience versus full node API
  2. same libxml tree underneath
  3. dom_import_simplexml and back
  4. 8.4 spec-compliant Dom namespace
  5. all three load the whole document

basics

~20 s

SimpleXML is quickest for reading known structures; DOMDocument gives the full node API for editing, precise namespaces and XPath; Dom\XMLDocument, added in 8.4, is the spec-compliant successor with querySelector(). All three hold the whole document in memory.

solid answer

~40 s

`SimpleXMLElement` turns an XML tree into property and array access, which is ideal for reading a feed whose shape you know. `DOMDocument` exposes the full DOM: every node type, creating, moving and removing nodes, namespace-aware methods, schema validation, and `DOMXPath` for queries. The two share libxml's tree, so `dom_import_simplexml()` and `simplexml_import_dom()` switch views without copying. PHP 8.4 added the `Dom` namespace: `Dom\XMLDocument` and `Dom\HTMLDocument` are WHATWG spec-compliant counterparts, built with static factories like `Dom\XMLDocument::createFromString()`, with `querySelector()` and `Dom\XPath`; the old classes remain for compatibility. For HTML specifically, `DOMDocument::loadHTML()` uses an HTML 4 parser, so modern HTML belongs in `Dom\HTMLDocument`. None of them streams; very large files need `XMLReader`.

code

php · 17 lines
php
<?php
declare(strict_types=1);

$xml = '<feed><product sku="A-100"><name>Desk lamp</name></product></feed>';

// SimpleXML: read quickly
$feed = simplexml_load_string($xml);
echo (string) $feed->product['sku'], "\n";            // A-100

// DOM view of the same node, no copy: remove it
$el = dom_import_simplexml($feed->product);
$el->parentNode->removeChild($el);
echo $feed->asXML();                                    // <feed/> after the XML declaration

// PHP 8.4+ spec-compliant API
$doc = Dom\XMLDocument::createFromString($xml);
echo $doc->querySelector('product > name')->textContent, "\n"; // Desk lamp

go deeper

for a junior

Recall that SimpleXML is the quick read API and DOMDocument the full node API, and that both load the whole document.

for a middle

Explain what DOM can do that SimpleXML cannot, how the two share one tree, and what the 8.4 Dom classes change for HTML and CSS selectors.

for a senior

Choose the API per job: DOM or Dom\HTMLDocument for editing and HTML, SimpleXML for reads, XMLReader when size makes any tree unaffordable.

for a principal

Set the guideline for new code between the legacy DOM classes and the 8.4 Dom namespace, weighing consistency with existing code against spec compliance.

## Three tree APIs over one parser PHP's bundled XML extensions all sit on **libxml2**. Three of them build an in-memory tree of the whole document: | API | Style | Best for | |---|---|---| | `SimpleXMLElement` (`simplexml_load_string()`, `simplexml_load_file()`) | properties for children, array access for attributes, `xpath()` | reading a known structure quickly | | `DOMDocument` + `DOMXPath` | the classic W3C DOM: nodes, `createElement()`, `appendChild()`, `removeChild()` | editing, precise namespace work, validation | | `Dom\XMLDocument` / `Dom\HTMLDocument` (PHP 8.4+) | spec-compliant DOM, `querySelector()`, `Dom\XPath` | new code that wants standard behaviour; HTML5 parsing | ## SimpleXML: fast to write, limited to reach SimpleXML shines when the question is "give me every product's SKU and price". It reads naturally, supports XPath through `xpath()` and `registerXPathNamespace()`, and can make light edits with `addChild()`, `addAttribute()` and `asXML()`. Its limits: - it has no concept of moving or removing arbitrary nodes; - text nodes, comments and processing instructions are hard to reach; - every value is a `SimpleXMLElement` that must be cast; - namespaced children need `children($namespaceUri)`. ## DOMDocument: the full node model `DOMDocument` gives you every node type and every operation: - **loading**: `load($filename, $options)` and `loadXML($source, $options)` return `bool`; - **navigation**: `getElementsByTagName()`, `childNodes`, `parentNode`, `DOMXPath::query()` and `evaluate()`; - **editing**: `createElement()`, `appendChild()`, `insertBefore()`, `removeChild()`, then `saveXML()`; - **validation**: `schemaValidate()` and `schemaValidateSource()` against an XSD. The price is verbosity: reading one attribute means `$el->getAttribute('sku')` after finding `$el`. ## Switching views without copying `dom_import_simplexml($sxe)` returns a `DOMElement` for the same underlying node, and `simplexml_import_dom($node)` goes the other way. Because both wrap the same libxml tree, a change made through one view is visible through the other. A common pattern is to navigate with SimpleXML and drop into DOM for the one operation SimpleXML lacks. ## The PHP 8.4 Dom namespace PHP 8.4 added a `Dom` namespace of new classes, counterparts to the old ones (`Dom\Node` for `DOMNode`, and so on), described in the migration guide as HTML 5-compatible and WHATWG spec-compliant, fixing long-standing bugs in the DOM extension. Key differences for everyday use: 1. **Construction by factory**: `Dom\XMLDocument::createFromString($xml)`, `createFromFile($path)` or `createEmpty()`, rather than `new DOMDocument()` plus `loadXML()`. 2. **CSS selectors**: `querySelector()` and `querySelectorAll()` on documents and elements. 3. **XPath**: `new Dom\XPath($doc)`, with `query()` returning a `Dom\NodeList`. 4. **HTML5**: `Dom\HTMLDocument::createFromString()` parses with an HTML5 parser. The manual warns that `DOMDocument::loadHTML()` uses an HTML 4 parser, so the resulting tree can differ from what browsers build, and it cannot be used safely for sanitizing HTML. The old `DOM*` classes are not deprecated; they remain for backward compatibility, and existing code keeps working. ## What none of them does: stream All three build the complete tree before you can read the first product, so memory grows with the document. For a feed of several hundred megabytes or more, use `XMLReader`, which pulls one node at a time, and expand only the current record into a small tree. ## A decision guide - Known, small-to-medium document, read-only: **SimpleXML**. - Editing, reordering, precise namespaces or XSD validation in existing code: **DOMDocument**. - New code on 8.4+, or any HTML parsing: **Dom\XMLDocument** / **Dom\HTMLDocument**. - Very large input: **XMLReader**, optionally expanding each record into one of the above. ## The same read in each API | Task | SimpleXML | DOMDocument | Dom\XMLDocument (8.4+) | |---|---|---|---| | load a string | `simplexml_load_string($xml)` | `$d = new DOMDocument(); $d->loadXML($xml);` | `Dom\XMLDocument::createFromString($xml)` | | first product's SKU | `(string) $x->product['sku']` | `$d->getElementsByTagName('product')->item(0)->getAttribute('sku')` | `$d->querySelector('product')->getAttribute('sku')` | | XPath | `$x->xpath('//product')` | `(new DOMXPath($d))->query('//product')` | `(new Dom\XPath($d))->query('//product')` |

  • Why should new HTML-scraping code use Dom\HTMLDocument instead of DOMDocument::loadHTML()?
    `DOMDocument::loadHTML()` parses with an HTML 4 parser, so modern markup can produce a different tree from the one a browser builds, and the manual says it cannot be used safely for sanitizing HTML. `Dom\HTMLDocument::createFromString()` (PHP 8.4+) uses an HTML5 parser and offers `querySelector()`.
  • Does converting with dom_import_simplexml() duplicate the document?
    No. It returns a `DOMElement` wrapping the same libxml node, so edits through the DOM object are visible through the `SimpleXMLElement` and vice versa. `simplexml_import_dom()` works the same way in the other direction.

saying these in an interview costs you the question

  • Believing SimpleXML streams large files while DOM loads them
  • Thinking the 8.4 Dom classes deprecated DOMDocument
  • Using DOMDocument::loadHTML() as a safe HTML sanitizer
  • Assuming dom_import_simplexml() makes an independent copy
  • Expecting to create Dom\XMLDocument with new and loadXML()