How do you build an XML document with ElementTree's SubElement and write it out?
answer
- One call creates and attaches the child
- Everything you store must already be a string
- The failure surfaces at serialization, not assignment
- Pretty printing mutates the tree in place
- Encoding and declaration defaults are not what you expect
basics
~20 sCreate the root with ET.Element, add children with ET.SubElement, which attaches them for you, set .text and attributes as strings, optionally pretty-print with ET.indent, then wrap the root in ET.ElementTree and call write with an encoding and xml_declaration=True.
solid answer
~40 s`ET.Element(tag, attrib)` makes a standalone element; `ET.SubElement(parent, tag, **attrs)` makes one **and attaches it**, which is why hand-rolled `Element` plus `parent.append` is redundant. Text and every attribute value must be `str` — assigning an `int` or a `float` is accepted silently and then raises `TypeError: cannot serialize ...` at write time, which is a favourite gotcha. `ET.indent(tree_or_element)` adds indentation whitespace in place, added in 3.9. Serialisation goes through `ET.ElementTree(root).write(target, encoding='utf-8', xml_declaration=True)` or `ET.tostring(root)`. Two defaults matter: the default encoding is `us-ascii`, so non-ASCII text comes out as numeric character references and `tostring` returns `bytes`; and `xml_declaration` defaults to `None`, which **omits** the declaration for `utf-8` and `us-ascii`, so pass `True` when a consumer expects one. Escaping of `&`, `<` and quotes is handled for you.
code
python · 10 linesimport io, xml.etree.ElementTree as ET
root = ET.Element("batch", {"run": "nightly"})
txn = ET.SubElement(root, "txn", id="1")
ET.SubElement(txn, "score").text = str(0.93) # str(), not the float itself
ET.indent(root, space=" ") # pretty-print in place, 3.9+
out = io.BytesIO()
ET.ElementTree(root).write(out, encoding="utf-8", xml_declaration=True)
print(out.getvalue().decode())go deeper
Be able to build a small document: Element for the root, SubElement for each child, set text and attributes, then write it to a file. Know that the library escapes special characters so you never concatenate XML by hand.
Explain the mechanics and the defaults: SubElement attaches as it creates, values must already be strings or serialization raises TypeError, indent mutates the tree, and encoding plus xml_declaration decide whether you get bytes, escapes and a prolog.
Show judgement about the output contract: precision and timezone conventions for the values you emit, stable prefixes via register_namespace so diffs stay readable, and the point at which a whole-tree build must give way to incremental emission.
Own whether producing a document-shaped payload is right for this integration at all, who owns the vocabulary and its versioning, and how the emitting service and its consumers stay compatible when the shape has to change.
## Constructing the tree Two constructors cover almost everything. `ET.Element(tag, attrib={}, **extra)` creates a free-standing element — typically your root. `ET.SubElement(parent, tag, attrib={}, **extra)` creates one *and appends it to the parent* in a single call, returning the new element. That attachment is the whole point: writing `child = ET.Element('x'); parent.append(child)` produces the same result with more moving parts, and forgetting the `append` is a common bug that yields a silently empty document. Attributes go in either as a dict or as keyword arguments, and the dict form is required for names that are not Python identifiers, such as anything hyphenated or namespaced. Character data is assigned to `.text`; the text that follows an element's closing tag, inside its parent, is `.tail`, which matters mainly when you are editing an existing document and want to preserve or produce spacing. ```python root = ET.Element("batch", {"run": "nightly"}) txn = ET.SubElement(root, "txn", id="1") ET.SubElement(txn, "score").text = str(0.93) ``` Editing is symmetric: `parent.remove(child)` detaches, `parent.insert(i, child)` places at a position, `parent.extend(iterable)` adds several, and slice assignment works because an element behaves like a sequence of children. ## Everything is a string `ElementTree` stores whatever object you assign to `.text` or to an attribute value and validates nothing at assignment time. The serializer, however, only knows how to write `str`. So `elem.text = 0.93` raises `TypeError: cannot serialize 0.93 (type float)` — but only when you serialize, potentially far from the line that caused it. The same applies to `elem.set('n', 3)`. Convert at the point of assignment, and make the formatting decision there too: `f"{value:.4f}"` says something a bare `str(value)` does not. There is no schema and no type system here. An XML document carries text; anything about numbers, dates or booleans is a convention between you and the consumer, and a fraud-scoring service emitting `<score>` values had better agree on precision and on the timezone rules for any timestamp before a clock-skew artefact turns into a downstream reconciliation dispute. ## Pretty-printing `ElementTree` writes no gratuitous whitespace: the default output is one long line. `ET.indent(tree, space=' ', level=0)`, added in 3.9, walks the tree and rewrites `text` and `tail` to produce indentation **in place**. That is a mutation, not a rendering option, so it changes the document's character data — harmless for most consumers, meaningful for anything whitespace-sensitive, and worth doing only when a human or a diff will read the output. ## Serialisation and its defaults Two entry points exist. `ET.tostring(element, encoding=None, ...)` returns the serialized form; `ET.ElementTree(root).write(file_or_filename, ...)` writes it. Their defaults are identical and both surprise people: - **Encoding.** The default is `us-ascii`, so `tostring(elem)` returns `bytes` and any non-ASCII character is escaped as a numeric character reference such as `café`. Pass `encoding='utf-8'` for real UTF-8 bytes, or the special `encoding='unicode'` to get a `str` back — which is also what you need when writing to a file opened in text mode. - **The XML declaration.** `xml_declaration` defaults to `None`, meaning "write one only if the encoding is something other than UTF-8, US-ASCII or unicode". So `write(f, encoding='utf-8')` produces **no** `<?xml ...?>` line. Pass `xml_declaration=True` explicitly when the consumer expects one. - **Empty elements.** `short_empty_elements` defaults to `True`, giving `<x />`; set it `False` when a consumer insists on `<x></x>`. - **Namespaces.** URIs are serialized with generated prefixes (`ns0`, `ns1`) unless you call `ET.register_namespace(prefix, uri)` beforehand or pass `default_namespace=uri`. ## Why not just format strings? The interview subtext of this question is usually escaping. Building XML with f-strings works until a value contains `&`, `<` or a quote character, at which point the output is malformed or, worse, subtly restructured by attacker-controlled input. `ElementTree` escapes text and attribute values on serialization, so correctness is the default rather than a review checklist item. The only real cost is that the whole document is materialised before writing, which is fine for a report and wrong for a multi-gigabyte export — for that, an incremental writer or a hand-managed chunked emission is the right shape.
- Why does assigning a float to an element's text fail only when you serialize?Because `Element` is a plain container that stores the object you give it without checking. Only the serializer requires `str`, and it raises `TypeError: cannot serialize 0.93 (type float)` when it reaches that node — possibly thousands of lines and one function boundary away from the assignment. Convert at assignment time, where you can also pin the formatting the consumer expects.
- Does write(f, encoding='utf-8') emit an XML declaration?No. `xml_declaration` defaults to `None`, which means 'only when the encoding is not UTF-8, US-ASCII or unicode', so that call writes the elements alone. Pass `xml_declaration=True` when a consumer requires the prolog. The related default catches people too: with no encoding at all the output is `us-ascii`, and non-ASCII characters come back as numeric character references.
- How do you get a str rather than bytes out of ET.tostring?Pass the special value `encoding='unicode'`. With any real codec name — or with the default, which is `us-ascii` — `tostring` returns encoded `bytes`. The same value works for `write`, and it is what you must use when the target is a file object opened in text mode, since a text stream cannot accept bytes.
- When is building the whole tree in memory the wrong approach to producing XML?When the output is large enough that materialising every element costs more than the process can afford — a bulk export rather than a report. The tree API has no incremental writer, so the alternatives are emitting records in chunks yourself with careful escaping, or writing well-formed fragments per record into a framing document. Both trade the API's automatic correctness for constant memory.
saying these in an interview costs you the question
- Assigns ints or floats to text and expects a cast
- Expects write with encoding utf-8 to emit a declaration
- Builds XML with f-strings and hand-written escaping
- Creates an Element then forgets to append it
- Assumes tostring returns a str by default
- Believes output is pretty-printed without calling indent