In PHP 8.5, when do you choose SimpleXML, DOMDocument or the Dom\XMLDocument class added in 8.4 for reading XML?
answer
- convenience versus full node API
- same libxml tree underneath
- dom_import_simplexml and back
- 8.4 spec-compliant Dom namespace
- all three load the whole document
basics
~20 sSimpleXML is quickest for reading known structures; DOMDocument gives the full node API for editing, precise namespaces and XPath; Dom\XMLDocument, added in 8.4, is the spec-compliant successor with querySelector(). All three hold the whole document in memory.
solid answer
~40 s`SimpleXMLElement` turns an XML tree into property and array access, which is ideal for reading a feed whose shape you know. `DOMDocument` exposes the full DOM: every node type, creating, moving and removing nodes, namespace-aware methods, schema validation, and `DOMXPath` for queries. The two share libxml's tree, so `dom_import_simplexml()` and `simplexml_import_dom()` switch views without copying. PHP 8.4 added the `Dom` namespace: `Dom\XMLDocument` and `Dom\HTMLDocument` are WHATWG spec-compliant counterparts, built with static factories like `Dom\XMLDocument::createFromString()`, with `querySelector()` and `Dom\XPath`; the old classes remain for compatibility. For HTML specifically, `DOMDocument::loadHTML()` uses an HTML 4 parser, so modern HTML belongs in `Dom\HTMLDocument`. None of them streams; very large files need `XMLReader`.
code
php · 17 lines<?php
declare(strict_types=1);
$xml = '<feed><product sku="A-100"><name>Desk lamp</name></product></feed>';
// SimpleXML: read quickly
$feed = simplexml_load_string($xml);
echo (string) $feed->product['sku'], "\n"; // A-100
// DOM view of the same node, no copy: remove it
$el = dom_import_simplexml($feed->product);
$el->parentNode->removeChild($el);
echo $feed->asXML(); // <feed/> after the XML declaration
// PHP 8.4+ spec-compliant API
$doc = Dom\XMLDocument::createFromString($xml);
echo $doc->querySelector('product > name')->textContent, "\n"; // Desk lampgo deeper
Recall that SimpleXML is the quick read API and DOMDocument the full node API, and that both load the whole document.
Explain what DOM can do that SimpleXML cannot, how the two share one tree, and what the 8.4 Dom classes change for HTML and CSS selectors.
Choose the API per job: DOM or Dom\HTMLDocument for editing and HTML, SimpleXML for reads, XMLReader when size makes any tree unaffordable.
Set the guideline for new code between the legacy DOM classes and the 8.4 Dom namespace, weighing consistency with existing code against spec compliance.
## Three tree APIs over one parser PHP's bundled XML extensions all sit on **libxml2**. Three of them build an in-memory tree of the whole document: | API | Style | Best for | |---|---|---| | `SimpleXMLElement` (`simplexml_load_string()`, `simplexml_load_file()`) | properties for children, array access for attributes, `xpath()` | reading a known structure quickly | | `DOMDocument` + `DOMXPath` | the classic W3C DOM: nodes, `createElement()`, `appendChild()`, `removeChild()` | editing, precise namespace work, validation | | `Dom\XMLDocument` / `Dom\HTMLDocument` (PHP 8.4+) | spec-compliant DOM, `querySelector()`, `Dom\XPath` | new code that wants standard behaviour; HTML5 parsing | ## SimpleXML: fast to write, limited to reach SimpleXML shines when the question is "give me every product's SKU and price". It reads naturally, supports XPath through `xpath()` and `registerXPathNamespace()`, and can make light edits with `addChild()`, `addAttribute()` and `asXML()`. Its limits: - it has no concept of moving or removing arbitrary nodes; - text nodes, comments and processing instructions are hard to reach; - every value is a `SimpleXMLElement` that must be cast; - namespaced children need `children($namespaceUri)`. ## DOMDocument: the full node model `DOMDocument` gives you every node type and every operation: - **loading**: `load($filename, $options)` and `loadXML($source, $options)` return `bool`; - **navigation**: `getElementsByTagName()`, `childNodes`, `parentNode`, `DOMXPath::query()` and `evaluate()`; - **editing**: `createElement()`, `appendChild()`, `insertBefore()`, `removeChild()`, then `saveXML()`; - **validation**: `schemaValidate()` and `schemaValidateSource()` against an XSD. The price is verbosity: reading one attribute means `$el->getAttribute('sku')` after finding `$el`. ## Switching views without copying `dom_import_simplexml($sxe)` returns a `DOMElement` for the same underlying node, and `simplexml_import_dom($node)` goes the other way. Because both wrap the same libxml tree, a change made through one view is visible through the other. A common pattern is to navigate with SimpleXML and drop into DOM for the one operation SimpleXML lacks. ## The PHP 8.4 Dom namespace PHP 8.4 added a `Dom` namespace of new classes, counterparts to the old ones (`Dom\Node` for `DOMNode`, and so on), described in the migration guide as HTML 5-compatible and WHATWG spec-compliant, fixing long-standing bugs in the DOM extension. Key differences for everyday use: 1. **Construction by factory**: `Dom\XMLDocument::createFromString($xml)`, `createFromFile($path)` or `createEmpty()`, rather than `new DOMDocument()` plus `loadXML()`. 2. **CSS selectors**: `querySelector()` and `querySelectorAll()` on documents and elements. 3. **XPath**: `new Dom\XPath($doc)`, with `query()` returning a `Dom\NodeList`. 4. **HTML5**: `Dom\HTMLDocument::createFromString()` parses with an HTML5 parser. The manual warns that `DOMDocument::loadHTML()` uses an HTML 4 parser, so the resulting tree can differ from what browsers build, and it cannot be used safely for sanitizing HTML. The old `DOM*` classes are not deprecated; they remain for backward compatibility, and existing code keeps working. ## What none of them does: stream All three build the complete tree before you can read the first product, so memory grows with the document. For a feed of several hundred megabytes or more, use `XMLReader`, which pulls one node at a time, and expand only the current record into a small tree. ## A decision guide - Known, small-to-medium document, read-only: **SimpleXML**. - Editing, reordering, precise namespaces or XSD validation in existing code: **DOMDocument**. - New code on 8.4+, or any HTML parsing: **Dom\XMLDocument** / **Dom\HTMLDocument**. - Very large input: **XMLReader**, optionally expanding each record into one of the above. ## The same read in each API | Task | SimpleXML | DOMDocument | Dom\XMLDocument (8.4+) | |---|---|---|---| | load a string | `simplexml_load_string($xml)` | `$d = new DOMDocument(); $d->loadXML($xml);` | `Dom\XMLDocument::createFromString($xml)` | | first product's SKU | `(string) $x->product['sku']` | `$d->getElementsByTagName('product')->item(0)->getAttribute('sku')` | `$d->querySelector('product')->getAttribute('sku')` | | XPath | `$x->xpath('//product')` | `(new DOMXPath($d))->query('//product')` | `(new Dom\XPath($d))->query('//product')` |
- Why should new HTML-scraping code use Dom\HTMLDocument instead of DOMDocument::loadHTML()?`DOMDocument::loadHTML()` parses with an HTML 4 parser, so modern markup can produce a different tree from the one a browser builds, and the manual says it cannot be used safely for sanitizing HTML. `Dom\HTMLDocument::createFromString()` (PHP 8.4+) uses an HTML5 parser and offers `querySelector()`.
- Does converting with dom_import_simplexml() duplicate the document?No. It returns a `DOMElement` wrapping the same libxml node, so edits through the DOM object are visible through the `SimpleXMLElement` and vice versa. `simplexml_import_dom()` works the same way in the other direction.
saying these in an interview costs you the question
- Believing SimpleXML streams large files while DOM loads them
- Thinking the 8.4 Dom classes deprecated DOMDocument
- Using DOMDocument::loadHTML() as a safe HTML sanitizer
- Assuming dom_import_simplexml() makes an independent copy
- Expecting to create Dom\XMLDocument with new and loadXML()