skip to content

questions

4

How do Element.find, findall and iter differ in xml.etree.ElementTree?

level: juniorimportance: must knowfreq 45%

answer

  1. Three different ways to reach into a tree
  2. One match, all matches, whole subtree
  3. A missing match is quiet, not loud
  4. Paths are one level deep by default
  5. iter includes the element it started from

basics

~20 s

Element.find returns the first match or None, findall returns a list of every match at that path, and iter walks the whole subtree recursively, including the element itself, yielding each element with the given tag.

solid answer

~40 s

All three search from one `Element`. `find(path)` returns the first matching sub-element or `None` and never raises, so test the result with `is not None` rather than truthiness. `findall(path)` returns a list of every match. Both take the same small XPath subset and are **one level deep by default**: `root.findall('txn')` matches direct children only, while `'txn/score'` or `'.//score'` reaches further down. `iter(tag)` is different in kind: it is a recursive generator over the entire subtree, it starts at the element it was called on, and it ignores path syntax entirely. Around the edges, `iterfind` is the lazy form of `findall`, `findtext` returns the matched element's text with an optional default, and iterating the element itself (`for child in elem`) yields only its direct children.

code

python · 11 lines
python
import xml.etree.ElementTree as ET

doc = """<batch><txn id="1"><score>0.93</score></txn>
         <txn id="2"><score>0.11</score><flag>clock-skew</flag></txn></batch>"""
root = ET.fromstring(doc)

print(root.find("txn").get("id"))                  # 1  -> first match only
print([t.get("id") for t in root.findall("txn")])  # ['1', '2'] direct children
print([e.tag for e in root.iter("score")])         # ['score', 'score'] subtree
print(root.find("score"))                          # None: not a direct child
print(root.find("txn/score").text)                 # 0.93 via a path

go deeper

for a junior

Be ready to parse a small document and pull out a value: fromstring or parse, then find, findall and .text or .get. Know that find returns None when nothing matches and that you must check it explicitly.

for a middle

Explain the mechanics: findall is one level deep unless the path says otherwise, iter recurses over the whole subtree and includes the starting element, and the supported path syntax is a documented XPath subset rather than XPath itself.

for a senior

Show judgement about which search fits the document. Irregular or deeply nested input argues for iter or a .// path; a known record level argues for findall; defensive extraction argues for findtext with a default rather than chained find calls that can hit None.

for a principal

Own the question of whether an element-tree API is the right interface at all. Tree search couples code to document shape, so decide where the mapping from document to domain objects lives, and how that boundary absorbs a schema change without spraying path strings across the codebase.

## Two objects, not one `xml.etree.ElementTree` gives you two distinct things and the interview usually starts by checking that you can tell them apart. `ET.parse(source)` reads a file or file object and returns an **ElementTree** — a wrapper around the whole document that also knows how to write it back. `ET.fromstring(text)` parses a string or bytes and returns the **root Element** directly. If you have a tree, `tree.getroot()` gets you to the root element; almost all searching happens on elements. An `Element` is a small container: a `tag` (a string), an `attrib` dict, a `text` string for the character data directly inside the open tag, a `tail` string for the character data after the close tag, and a sequence of child elements. Because an element behaves like a sequence of its children, `len(elem)` is the child count, `elem[0]` is the first child, and `for child in elem:` iterates **direct children only**. ## find, findall, iterfind `find(path)` and `findall(path)` accept the same path expressions and differ only in how much they return. `find` gives the first matching sub-element or `None`; `findall` gives a list, empty when nothing matches. `iterfind` is `findall` without building the list — useful when the matches are many and you only stream over them. Neither raises when there is no match, which is the single most common source of surprise: a missing element is a quiet `None`, not an exception. The path language is a deliberately small subset of XPath, not the whole thing. What is supported is worth memorising: a plain tag name matches direct children; `a/b` walks a fixed path; `*` matches any single element; `.` is the current element and `..` its parent; `.//b` matches `b` at any depth below; `[@id]` requires an attribute; `[@id='1']` requires a value; `[b]` requires a child element; `[b='x']` requires a child with that text; `[1]` selects by position, counting from one. There are no axes, no functions, no arbitrary expressions. The default depth trips people constantly. Given `<batch><txn><score/></txn></batch>`, `root.findall('score')` is an empty list, because `score` is not a direct child of `batch`. You want `root.findall('txn/score')` or `root.findall('.//score')`. ## iter `iter(tag=None)` is not a path search at all. It is a generator that performs a document-order walk of the element and everything beneath it, yielding every element whose tag matches — or every element, when the tag is omitted. Two consequences follow. First, it **includes the element you called it on**, so `root.iter('batch')` yields the root itself. Second, it takes a tag, not a path: `elem.iter('txn/score')` matches nothing, because no element is literally named that. `iter` is the right tool for "visit every X anywhere in this document"; `findall('.//X')` expresses the same intent as a path and is the better choice when you also need predicates. A related helper, `itertext`, walks the subtree and yields the text fragments rather than the elements. ## The truthiness trap Because an `Element` also implements `__len__`, an element with no children is currently falsy. So: ```python if root.find("txn"): # WRONG: false for a childless <txn/> ... if root.find("txn") is not None: # right ... ``` On CPython 3.14 the first form still returns the old length-based answer but emits a `DeprecationWarning` saying the truth value will always be `True` in future versions. Either way, `is not None` is the only correct test, and interviewers do watch for it. ## Getting data out Once you have an element, its character data is `elem.text` — which is `None`, not `''`, for `<a/>`, and which holds only the text before the first child, with anything after a child living on that child's `tail`. Attributes come from `elem.get('id')`, with an optional default second argument, or from the `attrib` dict; `elem.keys()` and `elem.items()` also work. `findtext(path, default=None)` combines the search and the text read in one call, which keeps simple extraction readable. ## Choosing between them in practice Reach for `find`/`findtext` when the document schema guarantees at most one match and you want a default for the missing case. Reach for `findall` when you are iterating records at a known level. Reach for `iter` when the structure is irregular and you genuinely mean "everywhere below here". And when you only need the immediate children, iterate the element itself rather than searching for `'*'` — it is clearer and does no path parsing.

  • Why is `if root.find('txn'):` an unsafe way to check whether a match was found?
    Because an `Element` implements `__len__`, so a matched but childless element such as `<txn/>` is falsy and the branch silently behaves as though nothing matched. On 3.14 that truth test also raises a `DeprecationWarning` saying it will always return `True` in future versions. Test the result explicitly with `is not None`; use `len(elem)` when you actually mean 'has children'.
  • What path syntax do find and findall actually support?
    A small XPath subset, not the language itself: tag names, `a/b` paths, `*`, `.`, `..`, `.//` for any depth, attribute predicates `[@id]` and `[@id='1']`, child predicates `[b]` and `[b='x']`, and one-based positional `[1]`. There are no axes, functions or computed expressions. Anything richer means pre-selecting in Python or reaching for a full XPath engine outside the standard library.
  • How do you read the text and attributes of an element you have found?
    `elem.text` holds the character data before the first child and is `None` — not an empty string — for `<a/>`; anything after a child element lives on that child's `tail`. Attributes come from `elem.get('id')`, which takes an optional default, or from the `attrib` dict. `elem.findtext('score', '0')` combines the search and the text read, returning the default when nothing matches.

find and findall are like asking for a folder by path; iter is like a recursive search over everything under this folder, and it counts the folder you started in.

saying these in an interview costs you the question

  • Thinks find raises an exception when nothing matches
  • Uses if elem.find(...) truthiness to test for a match
  • Expects findall to search the whole document by default
  • Believes ElementTree supports full XPath expressions
  • Confuses iter over the subtree with iterating direct children
  • Assumes elem.text returns an empty string for an empty element

context

open as a page

Why does ElementTree's findall('txn') match nothing when the XML declares xmlns?

level: middleimportance: must knowfreq 50%

basics

~20 s

The parser expands every namespaced tag into Clark notation, so the element is named {uri}txn and the bare name txn never matches. Search with a prefix-to-URI map, findall('t:txn', {'t': uri}), or write the {uri}txn form yourself.

open as a page

How do you build an XML document with ElementTree's SubElement and write it out?

level: middleimportance: should knowfreq 30%

basics

~20 s

Create the root with ET.Element, add children with ET.SubElement, which attaches them for you, set .text and attributes as strings, optionally pretty-print with ET.indent, then wrap the root in ET.ElementTree and call write with an encoding and xml_declaration=True.

open as a page

ElementTree.parse OOMs on a 6 GB XML feed; how do you stream it with iterparse?

level: seniorimportance: should knowfreq 35%

basics

~20 s

ET.parse materialises the entire document as objects, so peak memory is a multiple of the file size. ET.iterparse yields (event, element) pairs as parsing proceeds: handle each record on its end event, clear it, and unhook it from the root so memory stays flat.

open as a page