Why does ElementTree's findall('txn') match nothing when the XML declares xmlns?
answer
- The bare tag name is not the real name
- Look at what root.tag actually prints
- Braces around a URI in every tag
- The map's prefix is yours, not the document's
- Only the URI carries meaning
basics
~20 sThe parser expands every namespaced tag into Clark notation, so the element is named {uri}txn and the bare name txn never matches. Search with a prefix-to-URI map, findall('t:txn', {'t': uri}), or write the {uri}txn form yourself.
solid answer
~40 sAn `xmlns` declaration puts the element into a namespace, and `xml.etree.ElementTree` records that by rewriting the tag as `{uri}txn` — Clark notation — so `findall('txn')` matches nothing and returns an empty list rather than an error. There are three fixes. Spell the expanded name: `findall('{urn:example:txn}txn')`. Pass a prefix map: `findall('t:txn', {'t': 'urn:example:txn'})`, where the prefix `t` is **yours to choose** and need not match the prefix used in the document — only the URI is meaningful. Or use the namespace wildcard `findall('{*}txn')` when you genuinely do not care which namespace it came from. Two traps: a default `xmlns` applies to elements but **not** to unprefixed attributes, so `elem.get('id')` still works; and stripping `xmlns` from the source to make searches work destroys the document's meaning.
code
python · 14 linesimport xml.etree.ElementTree as ET
doc = """<feed xmlns="urn:example:txn" xmlns:m="urn:example:meta">
<txn id="1"><m:flag>clock-skew</m:flag></txn>
</feed>"""
root = ET.fromstring(doc)
print(root.tag) # {urn:example:txn}feed
print(root.findall("txn")) # [] - the plain name never matches
ns = {"t": "urn:example:txn", "m": "urn:example:meta"}
print(root.findall("t:txn", ns)) # matched via the prefix map
print(root.find("t:txn/m:flag", ns).text) # clock-skew
print(root.findall("{urn:example:txn}txn")) # same, in Clark notation
print(root.find("t:txn", ns).get("id")) # 1 - attributes are unprefixedgo deeper
Recognise the symptom: a search that returns an empty list on a document containing xmlns. Print root.tag, see the URI in braces, and know that the fix is a namespace map rather than editing the document.
Explain the mechanics — parse-time expansion to Clark notation, the prefix map being your own alias rather than the document's, the wildcard forms added in 3.8, and the fact that unprefixed attributes are in no namespace.
Show that you handle producers who change prefixes or version their namespace URI: build the map from parsed bindings, choose wildcards deliberately rather than defensively, and reject namespace-stripping hacks as data corruption in a code review.
Own the contract with the feed's producer. Decide whether a namespace URI change is a breaking version signal your ingest must reject or a compatible one it should tolerate, and make that policy explicit rather than leaving each parser to guess.
## What a namespace declaration actually does `xmlns="urn:example:txn"` on an element says that this element and its unprefixed descendants belong to that namespace URI. `xmlns:m="urn:example:meta"` binds a *prefix* to a second URI, so `<m:flag>` belongs to that one. The prefixes are a serialization convenience: two documents that use `t:txn` and `q:txn` are identical in meaning if both prefixes resolve to the same URI, and a document can rebind prefixes partway through. Because only the URI carries meaning, `ElementTree` throws the prefixes away at parse time and stores the fully expanded name on each element in **Clark notation**: `{urn:example:txn}txn`. Print `root.tag` on a namespaced document and you see it immediately. That is the whole explanation for the empty result: `findall('txn')` asks for elements literally named `txn`, and no element in the tree is named that. ## The three ways to search **Clark notation inline.** `root.findall('{urn:example:txn}txn')` works with no extra machinery. It is verbose and it hard-codes the URI into every path string, but for a one-off script it is fine. **A prefix map.** `find`, `findall`, `iterfind` and `findtext` all take a second argument mapping prefixes to URIs: ```python ns = {"t": "urn:example:txn", "m": "urn:example:meta"} root.findall("t:txn", ns) root.find("t:txn/m:flag", ns).text ``` The critical point, and the one interviewers probe, is that `t` is **your** prefix, not the document's. You could call it `x`; the search would still work as long as it maps to the right URI. Conversely, copying the document's prefix and assuming it is stable is a bug waiting for the day the producer re-serializes with different prefixes. A prefix that appears in the path but not in the map you passed raises `SyntaxError` — a useful, loud failure. Passing no map at all with a prefixed path fails quietly instead, which is one more reason always to pass the map. A default namespace has no prefix in the document, so there is nothing to copy: you invent one for your map. You can also pass `''` as the key to move every unprefixed name in the path into that namespace. **Wildcards.** Since 3.8 the path grammar accepts `{*}txn` (that tag in any namespace, or none), `{urn:example:txn}*` (any tag in that namespace) and `{}txn` (only the unqualified form). `{*}` is genuinely useful when you consume feeds from several producers who version their namespace URI, and it is a trap when you need to distinguish two vocabularies that share a tag name. ## Attributes behave differently This is the detail that separates a middle from a junior answer. A default `xmlns` declaration applies to elements only. An unprefixed attribute is in **no** namespace, ever, so on `<txn id="1" xmlns="urn:example:txn">` the element is `{urn:example:txn}txn` while the attribute key stays plain `id` and `elem.get('id')` works unchanged. Only an explicitly prefixed attribute such as `m:kind` becomes `{urn:example:meta}kind` in `attrib`. ## Writing namespaced XML back out On serialization the process runs in reverse and `ElementTree` has to invent prefixes for the URIs it finds, producing `ns0`, `ns1` and so on. Two controls exist. `ET.register_namespace(prefix, uri)` sets a global preferred prefix before you serialize; the `default_namespace=` argument to `write` and `tostring` promotes one URI to the unprefixed default. Neither changes the meaning of the document, only its readability — which matters mainly when a human or a diff will read the output. ## The anti-patterns Three shortcuts show up in real code and all three are worth rejecting out loud in an interview. Stripping `xmlns` attributes from the text before parsing makes the paths simple and throws away the only thing that distinguishes two vocabularies with the same tag names. Regex-stripping `{...}` from every tag after parsing has the same effect one step later, and breaks the moment two namespaces collide. Matching on `tag.endswith('txn')` looks harmless until a tag named `outertxn` shows up. If you truly do not care about the namespace, say so explicitly with a `{*}` wildcard rather than mutilating the data. ## Debugging recipe When a search mysteriously returns `[]`, print `root.tag` first. If it comes back wrapped in braces, the answer is namespaces; build the map from the URIs you see and the searches start working. Capturing the mapping as it is parsed is also possible by asking `iterparse` for `start-ns` events, which is how generic tooling discovers URIs it was not told about in advance.
- Does the prefix in your namespaces map have to match the prefix used in the document?No. Prefixes are local to the document's serialization and can be rebound at any element; ElementTree discards them and keeps only the URI. Your map is a private alias table, so `{'t': uri}` and `{'anything': uri}` behave identically as long as the URI matches exactly, character for character. A prefix used in a path but missing from the map you pass raises SyntaxError.
- Does a default xmlns declaration also put the element's attributes into that namespace?No — this is a common misreading of the XML rules. A default namespace applies to element names only; an unprefixed attribute is in no namespace at all, so `elem.get('id')` still works on a document with a default `xmlns`. Only an explicitly prefixed attribute is expanded, appearing in `attrib` under its Clark-notation key such as `{urn:example:meta}kind`.
- How do you avoid ns0 and ns1 prefixes when serializing a namespaced tree?Call `ET.register_namespace(prefix, uri)` before serializing to register the prefix you want globally, or pass `default_namespace=uri` to `write` or `tostring` to promote one URI to the unprefixed default. Neither affects the document's meaning — the URIs and therefore the expanded names are identical — so this is purely about output that a human or a diff can read.
- How would you discover the namespace URIs in a document you have never seen?Print `root.tag` and a few descendant tags: the Clark-notation braces show the URIs directly. For a systematic sweep, `ET.iterparse` accepts a `start-ns` event that yields each `(prefix, uri)` binding as the parser encounters it, which is how generic tooling builds a map for documents whose vocabulary is not known in advance.
The prefix in the document is a local nickname; the URI is the legal name. ElementTree files everything under the legal name, so searching by nickname finds nobody.
saying these in an interview costs you the question
- Strips xmlns from the source so plain searches work
- Copies the document's prefix and assumes it is fixed
- Thinks find and findall resolve namespaces automatically
- Regex-removes the brace-wrapped URI from every tag
- Assumes a default xmlns applies to attributes as well
- Matches namespaced tags with endswith on the tag string