skip to content

A supplier's 3 GB XML product feed exhausts memory under simplexml_load_file(); how do you process it in PHP with XMLReader instead?

level: seniorimportance: should knowfreq 32%

answer

  1. pull parser, one node at a time
  2. read() to find, next() to skip subtrees
  3. expand() just the current record
  4. simplexml_import_dom for convenience
  5. XMLReader::fromUri since 8.4

basics

~20 s

simplexml_load_file() builds the whole tree in memory. XMLReader pulls one node at a time: reach each <product> with read() and next('product'), expand() just that element into a small tree, process it and move on, so memory stays flat.

solid answer

~40 s

Tree APIs such as `simplexml_load_file()` and `DOMDocument::load()` parse the whole document into memory before you see the first product, so a multi-gigabyte feed hits `memory_limit`. `XMLReader` is a **pull parser**: open it with `XMLReader::fromUri($path)` (PHP 8.4+) or `open()`, call `read()` until the cursor is on the first `<product>` element, then loop. For each product, `expand($doc)` builds a DOM copy of just that element, which `simplexml_import_dom()` turns into friendly SimpleXML access; process it and call `next('product')`, which jumps to the next sibling of that name without walking the current subtree. Only one product's tree exists at a time. Checks use `nodeType === XMLReader::ELEMENT` and `localName`. The same libxml options and internal-error handling apply, and `XMLWriter` streams output the same way.

code

php · 25 lines
php
<?php
declare(strict_types=1);

$reader = XMLReader::fromUri('supplier-feed.xml');
$doc = new DOMDocument();
$count = 0;

// move to the first <product>
while ($reader->read() && !($reader->nodeType === XMLReader::ELEMENT && $reader->localName === 'product'));

while ($reader->nodeType === XMLReader::ELEMENT && $reader->localName === 'product') {
    $node = $reader->expand($doc);
    if ($node !== false) {
        $product = simplexml_import_dom($node);
        $sku = (string) $product['sku'];
        $price = (float) $product->price;
        // upsert $sku / $price here, then let $product go
        $count++;
    }
    if (!$reader->next('product')) {
        break;
    }
}
$reader->close();
echo $count, ' products, peak ', memory_get_peak_usage(true), " bytes\n";

go deeper

for a junior

Recall that SimpleXML and DOM load the whole file, while XMLReader reads one node at a time.

for a middle

Explain the cursor model, the difference between read() and next(), and how expand() plus simplexml_import_dom() gives easy access to one record.

for a senior

Build a flat-memory import for huge feeds with error collection, namespace-safe checks and peak-memory logging, and stream exports with XMLWriter.

for a principal

Weigh streaming XML imports in PHP against staging the feed elsewhere, considering failure recovery, partial imports and supplier SLAs.

## Why the tree APIs fail on big feeds `simplexml_load_file()`, `DOMDocument::load()` and `Dom\XMLDocument::createFromFile()` all build a complete in-memory tree. Every element, attribute and text node becomes a libxml node plus PHP wrapper objects when you touch them, so memory grows with the whole feed and the job dies with "Allowed memory size exhausted" long before the first record is processed. Raising `memory_limit` postpones the problem until the next, larger feed. ## XMLReader: a cursor over the document `XMLReader` exposes the document as a stream of nodes. You move a cursor forward and inspect the node it is on: - **opening**: `XMLReader::fromUri($uri, $encoding, $flags)`, `fromString()` and `fromStream()` are static factories added in PHP 8.4; the older `XMLReader::open()` and `XMLReader::XML()` still work; - **moving**: `read()` advances to the next node in document order (including into children); `next(?string $name)` skips the current subtree and moves to the next sibling, optionally one with the given name; - **inspecting**: `nodeType` (compare with `XMLReader::ELEMENT`, `XMLReader::END_ELEMENT`, ...), `name`, `localName`, `namespaceURI`, `getAttribute($name)`; - **materialising**: `expand(?DOMNode $baseNode)` returns a `DOMNode` copy of the current node and its subtree; `readOuterXml()` returns it as a string. ## The record-at-a-time pattern For a feed of repeated `<product>` elements: 1. open the reader; 2. `read()` until the cursor is on the first `<product>` element; 3. loop while the current node is a `<product>` element: - `expand($doc)` it, passing a `DOMDocument` so the node belongs to a document; - wrap it with `simplexml_import_dom()` for easy access, and cast values; - handle the record (validate, upsert, count); - `next('product')` to jump to the next product, skipping the one just handled; 4. `close()` the reader. Only the current product's small tree exists at any moment, so memory stays roughly flat regardless of file size. ## Details that trip people up | Pitfall | What happens | Fix | |---|---|---| | calling `read()` after handling a record | the cursor walks into the record's children | use `next('product')` | | `expand()` without a base document | `simplexml_import_dom()` warns "Imported Node must have associated Document" | pass a `DOMDocument` as `$baseNode` | | a default namespace on the feed | `name` may carry a prefix | compare `localName` and `namespaceURI` | | collecting records into an array | memory grows again | process and discard each record | | silent parse errors mid-file | the loop ends early with a warning | enable `libxml_use_internal_errors(true)` and check `libxml_get_errors()` after the loop | The same `LIBXML_*` flags apply through the `$flags` argument, so the entity-safety rules for untrusted feeds are unchanged. ## Writing big XML: XMLWriter The output side mirrors this. `XMLWriter` writes a document incrementally instead of building a tree: - `XMLWriter::toUri($path)` (PHP 8.4+, or `openUri()`), then `startDocument('1.0', 'UTF-8')`; - `startElement('product')`, `writeAttribute('sku', $sku)`, `writeElement('name', $name)`, `endElement()`; - call `flush()` every few hundred records so the buffer is written out; - `endDocument()` and a final `flush()`. XMLWriter escapes text and attribute values itself, so values are passed as plain strings. ## Measuring success Log `memory_get_peak_usage(true)` every few thousand products. With the pattern above the figure should stay level from the first thousand records to the last; a steady climb means something is still accumulating. ## Choosing between expand() and readOuterXml() | Approach | Cost | Convenience | |---|---|---| | `expand($doc)` + `simplexml_import_dom()` | builds the record's DOM once | SimpleXML access, no reparse | | `new SimpleXMLElement($reader->readOuterXml())` | serialises then reparses the record | simple to write | | reading attributes and text with the cursor alone | lowest | verbose for nested records | For records of a few dozen nodes any of them is fine; `expand()` avoids the extra serialise-and-parse round trip.

  • Why use next('product') instead of read() after processing a product?
    `read()` moves to the next node in document order, which is the first child of the product you just handled, so the loop would walk every child node. `next('product')` skips the current subtree and moves to the next sibling named `product`, which is both faster and keeps the loop condition simple.
  • How would you write a large XML export without building a DOM tree?
    Use `XMLWriter`: open it with `XMLWriter::toUri()` (8.4+) or `openUri()`, call `startDocument()`, then `startElement()`, `writeAttribute()`, `writeElement()` and `endElement()` per record, and `flush()` periodically so the buffer is written out. Memory stays proportional to the unflushed buffer, not the document.

saying these in an interview costs you the question

  • Raising memory_limit as the fix for a feed that keeps growing
  • Believing SimpleXML reads files lazily
  • Calling read() after each record and walking into its children
  • Accumulating every expanded record into one array
  • Ignoring parse errors that end an XMLReader loop early