ElementTree.parse OOMs on a 6 GB XML feed; how do you stream it with iterparse?
answer
- The parse call is the problem, not the file
- Handle records as they finish, not after
- Elements complete on one particular event
- Clearing the record is only half of it
- The root still references every emptied element
basics
~20 sET.parse materialises the entire document as objects, so peak memory is a multiple of the file size. ET.iterparse yields (event, element) pairs as parsing proceeds: handle each record on its end event, clear it, and unhook it from the root so memory stays flat.
solid answer
~50 s`ET.parse` builds the whole tree before returning, and an element per tag costs far more than the bytes it came from, so a multi-gigabyte feed is hopeless. `ET.iterparse(source, events=('start', 'end'))` returns an iterator of `(event, element)` pairs delivered as the parser advances. Take the root from the first `start` event, then on each record's `end` event — at which point that element and its children are complete — process it, call `elem.clear()` to drop its children, text and attributes, and clear the root as well so the emptied shell is no longer referenced by the tree being built. `clear()` alone is not enough: the parser still attaches every finished record to its parent, so without the second step you simply accumulate empty elements. The cost is that you aggregate incrementally, in one pass, with no ability to search the document afterwards.
code
python · 16 linesimport xml.etree.ElementTree as ET
with open("feed.xml", "w", encoding="utf-8") as f: # stand-in for a huge feed
f.write("<feed>")
f.writelines(f"<txn id='{i}'><score>0.5</score></txn>" for i in range(100_000))
f.write("</feed>")
context = ET.iterparse("feed.xml", events=("start", "end"))
_, root = next(context) # the root element, as soon as it opens
total = 0.0
for event, elem in context:
if event == "end" and elem.tag == "txn":
total += float(elem.findtext("score"))
elem.clear() # drop this record's children
root.clear() # and unhook the emptied record itself
print(round(total, 2))go deeper
Know that ET.parse loads the entire document into objects and that a streaming alternative exists in the same module. Being able to name iterparse and say it yields elements as parsing proceeds is enough at this level.
Explain the event model and write the loop: capture the root, act on the end event where the element is complete, clear the element, and say why the root must be cleared too rather than reciting it as a magic line.
Demonstrate that you have run this in production: constant memory versus a fixture too small to expose the problem, incremental aggregation with bounded state, mid-stream ParseError leaving partial side effects, and idempotent or resumable consumption.
Own the ingest architecture. Decide whether a document-shaped feed is the right interface at all, where the pass boundary sits between streaming extraction and queryable storage, and what the operational contract is when a 6 GB nightly file arrives truncated.
## Why the in-memory parse dies `ET.parse(path)` is a batch operation: it reads the whole document and returns only once every element object exists. Each element carries a tag string, an attribute dict, text and tail strings and a child list, so the object graph for a record-heavy feed routinely costs several times the on-disk bytes. Take a fraud-scoring service whose nightly reconciliation reads a 6 GB transaction feed: the process is not slow, it is dead — and the failure mode is nasty, because a 27-minute suite that ingests a scaled-down fixture in CI passes happily and only production sees the resident-set curve climb until the kernel intervenes. The fix is not a bigger box. It is to stop holding the document. ## The iterparse contract `ET.iterparse(source, events=None)` takes a filename or a binary file object and returns an iterator of `(event, element)` pairs, produced incrementally as the parser consumes input. The event names are `'start'`, `'end'`, `'start-ns'`, `'end-ns'`, `'comment'` and `'pi'`; the default is `('end',)`. The semantics of the two main events are what the whole technique rests on. `'start'` fires when the open tag has been read: the element exists and its attributes are set, but its children and text are not yet parsed, so it is not safe to process. `'end'` fires when the close tag has been read: that element and everything beneath it are complete. Record processing therefore belongs on `'end'`, keyed by tag. Crucially, `iterparse` is still *building a tree* as it goes. Each finished element is appended to its parent. If you do nothing, you have simply re-implemented `parse` with extra steps and the same peak memory. ## The two-step release ```python context = ET.iterparse(path, events=("start", "end")) _, root = next(context) # the root element, as soon as it opens for event, elem in context: if event == "end" and elem.tag == "txn": handle(elem) elem.clear() # drop children, text, tail and attributes root.clear() # unhook the emptied record from the tree ``` `Element.clear()` resets one element: children removed, `text`, `tail` and `attrib` emptied. The element object itself survives, and — this is the part candidates miss — it is still in its parent's child list. So `elem.clear()` alone leaves a growing list of empty shells hanging off the root, and peak memory still rises with record count. Clearing the root drops those references so the shells can be collected. In a quick measurement over a synthetic feed, adding `root.clear()` roughly halved peak traced allocation at five thousand records, and the gap widens linearly beyond that, because one path is O(records) and the other is O(1). `root.clear()` is safe here precisely because you have already finished with every previous record. If the root carries attributes you need, read them once at the `start` event before the loop. For nested structures where the record's parent is not the root, keep a reference to that parent and clear it instead, or remove the finished child explicitly. ## What you give up Streaming buys constant memory and pays for it in expressiveness. You get exactly one forward pass, so any aggregate — sums, per-account grouping, deduplication — must be accumulated as you go, and a computation that needs to look back at earlier records needs its own bounded state, not the tree. You cannot run a `findall` over the document afterwards, because the document no longer exists. Anything that genuinely needs random access wants a different design: a first pass that extracts a compact index, or a load into storage that supports queries. Error handling changes shape too. A malformed document raises `ET.ParseError` partway through the loop, after side effects have already happened for the records that parsed. If the consumer is not idempotent, wrap it in a transaction, stage the output, or record a resume position. ## XMLPullParser, and when to prefer it `iterparse` insists on owning the source: you give it a path or a file object and it reads. When the bytes arrive some other way — a socket, a decompressing wrapper, chunks from an HTTP response — use `ET.XMLPullParser(events=('end',))`, feed it with `feed(chunk)` as data arrives, and drain completed pairs with `read_events()` between chunks. It is the same event model with the pump inverted, and it is what lets a streaming consumer stay non-blocking. ## Framing the answer in an interview The strongest version of this answer names the memory model first (objects, not bytes), then the event semantics (`end` means complete), then the two-step release with a reason for the second step, and closes on the tradeoff: one pass, incremental state, no post-hoc search. Mentioning that the CI fixture is too small to catch this is what makes it sound like a lesson learned rather than a recipe memorised.
- Why is calling elem.clear() on each record not enough to keep memory flat?Because `clear()` empties an element but does not detach it. The parser has already appended the finished record to its parent, so the parent accumulates one empty element per record and peak memory still grows linearly with record count. Clearing the parent — usually the root, captured from the first `start` event — drops those references so the shells can be collected.
- Why process on the 'end' event rather than 'start'?`'start'` fires when the open tag has been read: the element exists with its attributes, but its text and children have not been parsed yet, so reading them gives partial or empty results. `'end'` fires after the close tag, when that element's whole subtree is complete. `'start'` is still useful for capturing the root element or for reacting to a container's attributes before its contents arrive.
- When would you use XMLPullParser instead of iterparse?When you own the byte stream rather than handing over a source. `iterparse` opens and reads the file itself, which blocks; `XMLPullParser` is fed with `feed(chunk)` as bytes arrive and drained with `read_events()` between chunks, so it fits a socket, a decompressing wrapper, or a chunked HTTP response, and lets an asynchronous consumer stay non-blocking.
- What does a streaming design cost you compared with parsing the whole tree?One forward pass only. Aggregates must be accumulated as you go with bounded state, cross-record logic cannot look backwards through the tree, and no `findall` is possible afterwards because the document was never retained. Failures also land mid-stream, after earlier records have already produced side effects, so the consumer needs idempotency, staging or a resume position.
iterparse is a conveyor belt rather than a warehouse: each record passes under your hands complete, but if you never take the finished boxes off the belt, the belt fills up anyway.
saying these in an interview costs you the question
- Calls clear but leaves every emptied element on the root
- Processes records on the start event and reads empty text
- Reads the file into a string first, then parses it
- Thinks fromstring uses less memory than parse
- Expects to search the tree after streaming through it
- Proposes a bigger machine instead of changing the parse