skip to content

A chat-transcript archiver using xml.etree.ElementTree on user-supplied exports intermittently blows its 92nd-percentile latency budget — how do you confirm entity expansion is the cause and harden the parse?

level: seniorimportance: should knowfreq 24%

answer

  1. Plot cost against input size first
  2. A tiny payload with an enormous cost
  3. The parser's own message is the proof
  4. The guard is a C library version
  5. Kill a process, not a stuck thread

basics

~20 s

Correlate parse time and memory against input size: a bomb is tiny input with huge cost. Catch xml.etree.ElementTree.ParseError and read its message, check xml.parsers.expat.EXPAT_VERSION, then move untrusted parses behind a hardened library, a byte cap, and a killable worker process.

solid answer

~50 s

First make the signal measurable: record document bytes, parse duration and peak allocation per export. An entity-expansion payload is unmistakable once you plot it — a few hundred input bytes against seconds of CPU and hundreds of megabytes resident, while ordinary slow parses scale with input size. Second, look at what the parser says: on 3.14 the linked libexpat raises `xml.etree.ElementTree.ParseError` reading "limit on input amplification factor (from DTD and entities) breached", which is proof rather than inference — but only if the build links libexpat 2.4.1 or newer, so check `xml.parsers.expat.EXPAT_VERSION`; an older library gives no error, just consumption. Then harden: cap the byte size before parsing, reject exports that declare a DOCTYPE, parse untrusted documents through a hardened third-party XML library, and run the parse in a process you can kill, because a thread stuck in the C parse call has no cancellation point.

code

python · 14 lines
python
import xml.etree.ElementTree as ET
import xml.parsers.expat

print("libexpat:", xml.parsers.expat.EXPAT_VERSION)

decls = ['<!ENTITY a0 "AAAAAAAAAA">']
for i in range(1, 9):
    decls.append(f'<!ENTITY a{i} "{f"&a{i - 1};" * 10}">')
bomb = f'<?xml version="1.0"?><!DOCTYPE d [{"".join(decls)}]><d>&a8;</d>'

try:
    ET.fromstring(bomb)
except ET.ParseError as exc:
    print("refused:", exc)

go deeper

for a junior

Take away the diagnostic shape: a document that is tiny on disk but expensive to parse points at entity expansion, and the parser's own error message names the cause rather than you having to guess.

for a middle

Be ready to describe the instrumentation — bytes, duration and peak allocation per parse — and to explain that external entities are refused by default while internal ones still expand under a C-level limit.

for a senior

An interviewer expects the operational judgement: verify xml.parsers.expat.EXPAT_VERSION on the deployed image, stop swallowing ParseError, cap decompressed input, and isolate untrusted parses in a killable process because a stuck parse cannot be interrupted.

for a principal

Own the standard rather than the incident: one hardened parse entry point for externally supplied documents, build metadata that records the linked C library, and separate counters so an attack signal never hides inside a generic parse-error rate.

## Start by separating "slow" from "amplified" The two look identical on a latency dashboard and are trivial to tell apart once you record the right pair of numbers. For every export, log: - the input size in bytes, - the wall time of the parse, - and peak allocation during it — `tracemalloc` around the call is enough while you are investigating, or a resident-memory sample if you cannot afford the overhead. Normal parses scale roughly linearly with input size. An entity-expansion document is a **vertical outlier**: a payload measured in hundreds of bytes producing seconds of CPU and a memory spike, because the cost lives in expanded text that never appeared in the input at all. If your outliers sit on the size trend line, you have a slow parse — a large export, a cold disk, a contended worker — and entity expansion is not your problem. ## Then get proof rather than a plausible story On CPython 3.14 with a modern linked libexpat, the amplification guard fires and the failure arrives as `xml.etree.ElementTree.ParseError` whose message reads "limit on input amplification factor (from DTD and entities) breached". That string is the confirmation you want, and it means the attack was refused rather than absorbed. Two caveats decide how much comfort to take from it. 1. First, the guard is in the C library, so read `xml.parsers.expat.EXPAT_VERSION` on the exact image you deploy: 2.4.1 introduced the amplification limit and 2.6.0 fixed the quadratic large-token case, and a build linking an older system library has neither — there the same document produces no error, just consumption, which is precisely the intermittent-timeout shape you are chasing. 2. Second, if your code wraps the parse in a broad `except` and returns a default, that message never reaches a log; the first fix is often just to log the exception with the document's size, source and a truncated prefix, and to keep the offending payload for analysis. ## Also check what the archiver does with entity references it does not expand External general entities have not been processed by default since Python 3.7.1, and `xml.etree.ElementTree` raises on the undefined reference — but if any code path reads the same exports through `xml.sax` with default features, that path completes silently and stores empty fields. Two ingestion routes disagreeing about the same document is worth finding while you are in there. ## Now harden, in the order the controls actually bite 1. **Cap the bytes before the parser sees them.** An upload limit is not an XML control, but it removes the largest class of accidental blowups and costs nothing. Apply it to the decompressed size if exports arrive compressed — a size limit on compressed bytes bounds nothing about what gets parsed. 2. **Refuse the DOCTYPE.** An export produced by your own client has no legitimate reason to declare entities. Rejecting any document with a DOCTYPE is a policy you can state in one line, test, and explain in a review. The standard library has no switch for it: `xml.etree.ElementTree.XMLParser` exposes only `entity` and `target`, so you would drive `xml.parsers.expat` yourself, install an `EntityDeclHandler` that raises, and call `SetParamEntityParsing` with `xml.parsers.expat.XML_PARAM_ENTITY_PARSING_NEVER` — then rebuild tree construction on top of that. 3. **Which is why the real fix is a hardened third-party XML library.** Its parse entry points are drop-in replacements that refuse DTDs, entity declarations and external references outright and raise a dedicated exception, giving every ingestion path one stated policy instead of a per-module mix of loud errors, silent skips and version-dependent C guards. 4. **Bound the blast radius.** A parse already running inside libexpat cannot be interrupted: there is no cancellation point, so a timer in a thread will not stop it and the worker keeps allocating. If you need a hard ceiling on the tail, parse untrusted exports in a subprocess you can kill on a deadline, with a memory limit applied to that process. That also converts the failure mode from "the archiver stops serving" into "one export failed", which is what a percentile budget actually cares about. ## Close on the operational loop - Alert on parse failures by class, not just on latency; keep a counter for amplification refusals separately from ordinary parse errors, because a rising count is an attack signal rather than a quality signal. - And record the libexpat version alongside the Python version in your build metadata, so the next person does not have to rediscover that the protection they are relying on came from a C library they never chose.

  • Why will a timeout on a worker thread not rescue the tail latency here?
    The parse runs inside a C call with no cancellation point, so nothing checks a flag or delivers a signal until it returns. The thread keeps expanding text and allocating regardless of your timer. A separate process is the only reliable hard bound: you can kill it on a deadline and apply a memory limit to it without taking the service down.
  • The exports arrive gzipped. What does that change about your size cap?
    A cap on the compressed bytes bounds nothing that matters, because the ratio is attacker-controlled. Enforce the limit on the decompressed stream — decompress incrementally and abort once the output crosses the ceiling — and only then hand the result to a parser. Otherwise a small upload becomes a large document before any XML-level control has an opinion.
  • How would you tell an amplification refusal apart from an ordinary malformed export in your metrics?
    Both arrive as xml.etree.ElementTree.ParseError, so counting the exception class is not enough. Classify on the message — the amplification case reports a breached input amplification factor — and increment separate counters. A rising amplification count with steady traffic is a signal about senders, while malformed-XML counts usually track a client release.

saying these in an interview costs you the question

  • Blames document size without plotting cost against bytes
  • Assumes a thread timeout can abort a running parse
  • Trusts the amplification guard without checking libexpat
  • Swallows ParseError and loses the diagnostic message
  • Caps compressed bytes and calls the input bounded
  • Proposes stripping DOCTYPE text with a regular expression

context