Is xml.etree.ElementTree safe for parsing XML supplied by untrusted users?
answer
- The docs carry a security banner
- One C library under every front-end
- External entities off since a patch release
- Amplification guard lives in libexpat
- Hardened replacement library for untrusted input
basics
~20 sNo. Python's xml package documentation states its parsers are not secure against maliciously constructed data. ElementTree will not fetch external entities, but it still processes a DTD and expands internal entities, so untrusted XML belongs in a hardened parsing library.
solid answer
~40 sNo, and the standard library says so itself: the `xml` package documentation opens with a warning that its modules are not secure against erroneous or maliciously constructed data. `xml.etree.ElementTree`, `xml.dom.minidom`, `xml.dom.pulldom` and `xml.sax` are all front-ends over the same C library, libexpat. Two of the worst cases are blunted on 3.14: external general entities have not been processed by default since Python 3.7.1, so `xml.etree.ElementTree.fromstring` raises `ParseError: undefined entity` instead of reading a local file, and libexpat 2.4.1 and newer refuses runaway entity amplification. But those guards live in C, are not tunable from Python, and depend on which libexpat the build links — read `xml.parsers.expat.EXPAT_VERSION`. For genuinely untrusted documents use a hardened third-party XML library that refuses DTDs and entity declarations outright, and cap the input size before parsing.
code
python · 14 linesimport xml.etree.ElementTree as ET
external = ('<?xml version="1.0"?>'
'<!DOCTYPE d [<!ENTITY xxe SYSTEM "file:///etc/hosts">]>'
'<d>&xxe;</d>')
try:
ET.fromstring(external)
except ET.ParseError as exc:
print("external rejected:", exc)
internal = ('<?xml version="1.0"?>'
'<!DOCTYPE d [<!ENTITY greet "hello">]>'
'<d>&greet;</d>')
print("internal expanded:", ET.fromstring(internal).text)go deeper
Recall that Python's own xml package documentation warns these parsers are not secure against maliciously constructed data, and that the answer for untrusted input is a hardened parsing library plus a size cap — not a hand-written filter over the XML text.
Be ready to explain the mechanics: external general entities off by default since 3.7.1, ElementTree raising ParseError on the reference, internal entities still expanded, and the amplification guard living in libexpat rather than in Python.
An interviewer expects you to know which knobs do not exist. Show that you check xml.parsers.expat.EXPAT_VERSION, cap input bytes before parsing, and isolate hostile parses in a process you can kill rather than relying on a timeout.
Own the policy rather than the patch: decide once that untrusted XML is parsed only through a hardened library at a named boundary, make that the reviewable rule, and treat every direct stdlib parse of external data as a defect the build should catch.
**The warning is in the documentation, not just in folklore.** Python's `xml` package documentation opens with a security banner: the XML modules are not secure against erroneous or maliciously constructed data, followed by a table of which module is exposed to which attack. That banner is what an interviewer is listening for — "no, and here is what the parser actually does" — rather than a vague "XML is dangerous". **One engine, several front-ends.** `xml.etree.ElementTree`, `xml.dom.minidom`, `xml.dom.pulldom` and `xml.sax` are different APIs over the same C library, libexpat, which is reachable directly as `xml.parsers.expat`. The interesting behaviour is therefore mostly libexpat's, and the differences between the front-ends come down to what each does with the events expat emits. **Three attack shapes aimed at those parsers.** 1. *External entity resolution.* The document's internal DTD subset declares `<!ENTITY xxe SYSTEM "file:///etc/passwd">` and the body references `&xxe;`. A parser that resolves it reads that file, or issues a request from inside your network, and splices the bytes into the document your code then reads back out. 2. *Entity-expansion amplification.* Purely internal declarations, each referencing the previous one ten times. Nine levels deep turns a few hundred bytes into a gigabyte of text, with no network access required. 3. *Oversized tokens and deep nesting.* A single enormous token used to cost quadratic time inside libexpat. **What is actually true on CPython 3.14.** External general entities have not been processed by default since Python 3.7.1. `xml.etree.ElementTree` goes one step further and raises `xml.etree.ElementTree.ParseError: undefined entity &xxe;`, because as far as ElementTree is concerned that entity was never defined. `xml.sax` is quieter: it simply produces no character data for the reference unless `xml.sax.handler.feature_external_ges` is switched on. Amplification is caught inside libexpat 2.4.1 and newer, and surfaces in Python as a `ParseError` reading "limit on input amplification factor (from DTD and entities) breached"; the quadratic large-token case was fixed in libexpat 2.6.0. Which of those protections you get depends on the library your interpreter links, so read `xml.parsers.expat.EXPAT_VERSION` at runtime instead of assuming — a distribution build may link an older system library. **Why "not safe" still stands.** Those mitigations are hard-coded in C and are not tunable from Python. There is no supported way to say "reject any document that declares a DOCTYPE", "expand at most N entities" or "stop after ten megabytes" through `xml.etree.ElementTree`; its `xml.etree.ElementTree.XMLParser` exposes only `entity` and `target`. The one lever the standard library gives you is to drive `xml.parsers.expat` yourself and install callbacks — an `EntityDeclHandler` that raises on any entity declaration, an `ExternalEntityRefHandler` that refuses, plus `SetParamEntityParsing` with `xml.parsers.expat.XML_PARAM_ENTITY_PARSING_NEVER`. That genuinely works, but you are then hand-maintaining a security-critical parser configuration and rebuilding the tree-building layer on top of it. **So what the real fix looks like.** For untrusted XML, reach for a hardened third-party XML library — the kind that ships drop-in replacements for the standard parse entry points and refuses DTDs, entity declarations and external references outright, raising a dedicated exception rather than quietly returning a tree. Keep the standard-library parsers for XML your own systems produced. Around either choice, the boring controls still matter: cap the byte size before anything reaches a parser, never let a parse reach out to the network, and run hostile input somewhere you can kill. **One more Python-specific trap.** `xmlrpc.client` decompresses responses before parsing them, so a compressed payload can expand enormously before the XML layer ever sees it — a size cap applied to the compressed bytes is not a cap on what gets parsed. **What a junior is expected to say.** 'No — the standard library documents these parsers as not secure against maliciously constructed data. Modern Python will not fetch an external entity by default and ElementTree raises `ParseError` on one, but the DTD is still processed and internal entities still expand, so for untrusted input I would use a hardened parsing library and limit the input size first.' **Where this shows up in a codebase.** The risky call is rarely dramatic. It is `xml.etree.ElementTree.fromstring` on a request body, `xml.etree.ElementTree.parse` on an uploaded file, or `xml.dom.pulldom.parse` inside a feed reader — ordinary lines that nobody flags in review. A cheap audit is to list every parse entry point and label each with where its bytes came from: your own services, a named partner, or the open internet. Only the last two need the hardened path. Making that distinction explicit is usually worth more than any individual flag, because it turns "is XML safe?" into a question with a per-call-site answer that a reviewer can actually check.
- Which standard-library module would you use to check what protection your build actually has?`xml.parsers.expat` exposes `EXPAT_VERSION` and `version_info`. The amplification guard arrived in libexpat 2.4.1 and the large-token fix in 2.6.0, so a build linking an older system library has neither, however new the Python is. Reading that at startup and refusing to run untrusted-XML paths on an old library is a cheap, honest control.
- If ElementTree already refuses external entities, why is a hardened library still recommended?Because refusing external entities is only one of the three problems. Internal entity declarations are still expanded, the DTD is still processed, and none of the limits are reachable from Python — there is no supported knob for document size, nesting depth or entity count. A hardened library forbids the DOCTYPE and entity declarations outright, which is a policy you can state, test and audit.
- Does parsing untrusted XML in a thread with a timeout protect you?No. The parse runs inside a C call that has no cancellation point, so a timer cannot interrupt it and the thread keeps consuming memory until the process dies. If you need a hard bound, parse in a separate process you can kill, and cap the input bytes before you hand them over.
The stdlib parsers are a mail-sorting machine that will cheerfully follow a forwarding address printed inside the letter; the hardened library is the one that refuses to take instructions from the mail at all.
saying these in an interview costs you the question
- Claims the stdlib XML parsers are safe by default
- Thinks only external entities matter, not internal ones
- Believes ElementTree is safer than xml.sax by design
- Says validating the XML against a schema first prevents it
- Assumes a thread timeout can abort a running parse
- Treats 'no network access' as full protection