skip to content

Untrusted Input Handling

Hardening every path attacker data takes in: deserialized payloads, XML, uploaded archives, template strings, query parameters and lookalike text. Each has a default that trusts what it is given.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

28

Why is calling pickle.loads() on untrusted bytes the same as running that data?

level: juniorimportance: must knowfreq 70%

answer

  1. Loading is closer to executing than parsing
  2. The stream carries instructions, not only values
  3. One opcode imports a name, another calls it
  4. __reduce__ is the attacker's authoring interface
  5. Checking the result afterwards is already too late

basics

~20 s

Unpickling is not parsing. A pickle stream is a tiny program whose opcodes import names and call them, so pickle.loads on attacker-controlled bytes can invoke any importable callable with attacker-chosen arguments before it returns anything.

solid answer

~50 s

`pickle.loads` runs a small stack machine. Its opcodes can import a module and push any name from it, call that object with arguments the stream supplies, and push state into the result through `__setstate__`. An attacker writes a throwaway class whose `__reduce__` returns `(callable, args)`, pickles one instance, and every process that loads those bytes performs that call — before your code sees a return value, so nothing you check afterwards helps. What is at stake is the identity of the loading process: its credentials, its network access, its filesystem. The rule is that a pickle is exactly as trusted as whoever could write the bytes. For untrusted input use a data-only format such as JSON and construct your objects yourself; for pickles moving between components you own, authenticate the bytes with an HMAC before loading them.

code

python · 11 lines
python
import pickle


class Payload:
    def __reduce__(self):
        return (print, ("this call happened during unpickling",))


data = pickle.dumps(Payload())
result = pickle.loads(data)
print("returned:", result)

go deeper

for a junior

Be ready to state the rule in one sentence: never unpickle bytes you did not produce yourself, and reach for JSON when data crosses a trust boundary. Knowing that loading can run code, and saying so without hedging, is the whole ask here.

for a middle

Explain the mechanism rather than repeating the warning: opcodes in the stream resolve globals and call them, and reduce is how an attacker chooses that callable and its arguments. Be able to say precisely why checks placed after the load are too late.

for a senior

Show that you can find the exposure in a running system — pickles hiding in caches, shelve files, queue messages and inter-process channels — and describe both the migration to a data-only format and the HMAC stopgap for the paths you cannot change yet.

for a principal

Own the framing that a pickle read is a code-deployment path: whoever can write those bytes ships code into your process. Decide where that is ever acceptable, how format choice and key management are enforced across teams, and what holds the line during the migration.

## Loading a pickle is running a program The `pickle` module does not parse a document the way a JSON reader does. A pickle stream is a sequence of opcodes for a small stack machine that lives inside CPython, and the loader executes those opcodes in order. Most of them are dull: push a small integer, push a short string, build a list from the last N items on the stack. Three are not. * `GLOBAL` / `STACK_GLOBAL` carry a module name and a qualified name. The loader **imports that module** and pushes the object the name refers to. * `REDUCE` pops a callable and an argument tuple from the stack and **calls the callable**. * `BUILD` pushes a state dictionary into an object, invoking that object's `__setstate__` if it defines one. Put those together and the conclusion is unavoidable: whoever writes the bytes chooses which modules get imported, which objects get called, and with what arguments. `pickle.loads` is not reading a value out of the stream; it is following instructions that construct one. ## `__reduce__` is the authoring interface You do not need to hand-assemble opcodes to build such a payload, because Python hands you a supported way to emit them. When `pickle.dumps` meets an object whose class defines `__reduce__`, it calls it, and the method returns a callable plus the arguments to call it with. That pair is written out as exactly the `GLOBAL` + `REDUCE` sequence above. So an attacker writes a throwaway class whose `__reduce__` returns `(some_importable_callable, ("argument",))`, pickles one instance, and ships the resulting bytes. The class itself is never named in the stream and does not have to exist on the loading side — only the callable it points at does. A stream is also not limited to a single call. It is a program: it can resolve several globals, perform several `REDUCE` calls, use the result of one as the argument of the next, and drive `__setstate__` on objects it constructs. "It is only one function call, I can eyeball it" is not a property pickle gives you. ## What is actually at risk The attacker gains the identity of the loading process. That is not just "code execution" in the abstract — it is your service's database credentials, its cloud instance role, its outbound network access, its filesystem, its ability to read the secrets mounted next to it. Running the process as an unprivileged user narrows the blast radius but leaves everything the service itself is authorised to do, which is usually the interesting part. ## Defences that do not work * **`try` / `except` around the load.** The call happens while the opcodes are being executed. By the time an exception reaches you, whatever the stream wanted has already run. * **Size limits.** A payload that imports a module and calls one function fits in a few dozen bytes. Caps are a useful availability control against memory bombs; they are not a security control. * **Inspecting or validating the object after loading.** Too late, by definition. * **"It comes from our internal network / our own cache / our own queue."** That is an assumption about who can write those bytes, and it is the assumption that fails: a compromised neighbour, a writable shared volume, a cache-poisoning bug or a misconfigured queue all turn into code execution here, with no memory-safety bug required. * **Choosing a different protocol version, or a different pickle-based library.** The opcodes that import and call are present in every protocol. ## What does work For genuinely untrusted input, do not unpickle it. Use a data-only format — JSON is the usual answer — and build your objects yourself from the parsed primitives, with an explicit constructor that validates fields. The distinction is that a JSON document cannot name a callable; the mapping from data to objects is code you wrote. For pickles moving between components you control and cannot immediately re-format, change the question from *"is this payload safe?"* — which is undecidable — to *"did a holder of the key produce it?"*. Compute an HMAC over the pickled bytes with `hmac.new` and a key from a secret store, and verify it with `hmac.compare_digest` before calling `pickle.loads`. This is authentication, not validation, and it inherits every weakness of your key handling: a key in the source tree buys nothing, and it does not help if the entities holding the key are themselves the untrusted ones. Where you must keep accepting a pickle-shaped format from a semi-trusted producer, a restricted unpickler that allowlists the globals the stream may resolve narrows the surface considerably — but it narrows, it does not sandbox. ## Where pickles hide Interviewers like this question partly because the risky call is often not spelled `pickle.loads` in your code. Values in a `shelve` database are pickles. Objects passed between processes by `multiprocessing` are pickled and unpickled on the far side. Cached objects, session blobs and "just save this object to a file" helpers frequently are too. Finding those call sites is the real work; the rule itself is one sentence.

  • Does wrapping pickle.loads in try/except, or capping the payload size, make it safe?
    No. The call happens while the opcodes are being executed, so by the time an exception surfaces the callable has already run. A size cap only limits how much a payload can do, and importing a module and calling one function fits in a few dozen bytes. Both are reasonable availability controls against memory bombs and malformed input; neither is a security control.
  • Two services you own exchange pickles over a queue. How do you make that defensible?
    Authenticate rather than validate. Compute an HMAC-SHA256 over the pickled bytes with a key from a secret store, ship it alongside, and verify with hmac.compare_digest before calling pickle.loads. That changes the question from “is this payload safe” to “did a key holder produce it”. It buys nothing if the key lives in the repository, and nothing if the key holders are themselves the untrusted parties. A data-only format is still the better answer wherever the schema allows it.
  • Can a single pickle stream do more than one call?
    Yes — the stream is a program. It can resolve several globals, issue several calls, feed the result of one into the arguments of the next, build containers over the results and drive __setstate__ on objects it constructs. Treating a payload as “one call I could eyeball” misreads the format.

A JSON document is a filled-in form: the reader decides what to do with the fields. A pickle is a set of instructions handed to a clerk who follows them without question, including "go fetch this tool and use it".

saying these in an interview costs you the question

  • Says pickle only reads data and cannot execute anything
  • Thinks try/except around the load contains the damage
  • Believes checking the object after loading prevents the attack
  • Assumes an internal network makes untrusted pickles acceptable
  • Calls pickle “just a serialization format like JSON”
  • Claims running as an unprivileged user removes the risk

context

open as a page

Why pass values to sqlite3.Cursor.execute as a parameter tuple instead of formatting them into the SQL text?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Placeholders keep the statement and the data on separate channels. sqlite3 compiles the SQL text first, then binds each parameter into a slot of the compiled statement, so a value can never become syntax. String formatting merges the two.

open as a page

What does tarfile's 'data' extraction filter block, and why is it the 3.14 default?

level: middleimportance: must knowfreq 35%

basics

~20 s

The data filter refuses tar members that would land outside the destination directory, refuses links whose target is absolute or escapes, and rejects device and FIFO members, while clearing ownership and risky permission bits. Python 3.14 makes it the default for extraction.

open as a page

Why is str.format with an attacker-supplied template string a data-leak risk?

level: middleimportance: must knowfreq 42%

basics

~20 s

Python's format mini-language allows attribute access and indexing inside a placeholder, so whoever writes the template chooses what is read out of the objects you pass. A chain through a method's globals reaches module-level secrets.

open as a page

Why does Python's re pattern ^(\d+)+$ hang on a crafted 30-character string?

level: middleimportance: must knowfreq 55%

basics

~20 s

Python's re module uses a backtracking engine. Nesting one quantifier inside another over the same characters gives it exponentially many ways to split the input, so a 30-character string that ultimately fails to match costs about a billion steps.

open as a page

Why can't a table or column name be a sqlite3 query parameter, and what replaces it?

level: middleimportance: must knowfreq 55%

basics

~20 s

Parameters bind after the statement is compiled, but identifiers must be known during compilation so the engine can resolve the table and columns. An identifier therefore belongs in the SQL text, chosen from a fixed allowlist your code owns.

open as a page

What does unicodedata.normalize('NFKC', s) fold that NFC leaves alone, and why can that turn a rejected username into a reserved one?

level: middleimportance: must knowfreq 45%

basics

~20 s

NFKC adds compatibility mappings on top of NFC: fullwidth letters, ligatures, superscripts, circled digits and abbreviation characters all collapse to plain ASCII equivalents. A string that failed a reserved-name check before folding can therefore equal a reserved name after it.

open as a page

What does xml.etree.ElementTree do with an external entity reference like &xxe;?

level: middleimportance: must knowfreq 42%

basics

~20 s

It refuses it. ElementTree does not process external general entities, so nothing defines that reference and the parse fails with xml.etree.ElementTree.ParseError reading 'undefined entity'. Entities declared inline in the document's DTD subset are still expanded normally.

open as a page

Does zipfile.ZipFile.extractall stop archive members from escaping the destination?

level: juniorimportance: should knowfreq 30%

basics

~20 s

Yes, for member names: ZipFile.extract and extractall strip drive letters and leading separators and remove every parent-directory component, so members land under the destination. They do not recreate symlink members, do not apply stored permission bits, and do not limit decompressed size.

open as a page

Why can't an f-string be exploited the way str.format with a user template can?

level: juniorimportance: should knowfreq 32%

basics

~10 s

An f-string is syntax: CPython compiles its interpolations from your source file, so a template that arrives at runtime can never become one. Only str.format, or eval, interprets a runtime string as a template.

open as a page

Why does int() on a 5,000-digit string raise ValueError, and how do you lift the cap?

level: juniorimportance: should knowfreq 28%

basics

~20 s

CPython caps conversion between a decimal string and an int at 4,300 digits, because that conversion is quadratic in the digit count and makes a cheap request expensive. Lift it with sys.set_int_max_str_digits(), PYTHONINTMAXSTRDIGITS, or -X int_max_str_digits.

open as a page

Why do two Python str values that render identically sometimes compare unequal, and how does unicodedata.normalize fix it?

level: juniorimportance: should knowfreq 40%

basics

~20 s

Python compares str values code point by code point, and the same visible text has several spellings: 'e' with an acute accent can be one code point or 'e' plus a combining accent. unicodedata.normalize('NFC', s) rewrites both into one canonical spelling first.

open as a page

Is xml.etree.ElementTree safe for parsing XML supplied by untrusted users?

level: juniorimportance: should knowfreq 30%

basics

~20 s

No. Python's xml package documentation states its parsers are not secure against maliciously constructed data. ElementTree will not fetch external entities, but it still processes a DTD and expands internal entities, so untrusted XML belongs in a hardened parsing library.

open as a page

How does overriding pickle.Unpickler.find_class restrict what a payload can build?

level: middleimportance: should knowfreq 35%

basics

~10 s

find_class is called for every global a pickle stream resolves, receiving the module and the qualified name. Subclass pickle.Unpickler, allowlist the pairs your format legitimately needs, and raise pickle.UnpicklingError for everything else.

open as a page

Which placeholder styles does sqlite3 accept, and what does the DB-API paramstyle global tell you?

level: middleimportance: should knowfreq 42%

basics

~20 s

sqlite3 accepts qmark placeholders (a bare ?, filled from a sequence) and named placeholders (:name, filled from a mapping). sqlite3.paramstyle is the PEP 249 global naming a driver's expected spelling; for sqlite3 it reads 'qmark'.

open as a page

When does str.casefold() differ from str.lower(), and why does that matter for a case-insensitive uniqueness check?

level: middleimportance: should knowfreq 35%

basics

~20 s

str.lower() maps each character to its lowercase form; str.casefold() applies Unicode's aggressive full case folding, which can change length — the German sharp s folds to 'ss'. A uniqueness check using str.lower() can therefore admit two strings that are the same word.

open as a page

How do you cap a decompression bomb when a worker extracts uploaded archives?

level: seniorimportance: should knowfreq 26%

basics

~20 s

Extraction filters guard paths, not volume, so enforce your own limits: a member count cap, a total decompressed-byte budget counted as you stream members out, and a per-member cap. Treat the sizes in archive headers as hints, and unpack inside a disposable, bounded scratch directory.

open as a page

How do you render a user-supplied str.format template safely in a webhook receiver?

level: seniorimportance: should knowfreq 26%

basics

~10 s

Never hand untrusted text to str.format. Prefer string.Template, whose $name placeholders have no attribute or item syntax; if the mini-language is required, subclass string.Formatter, allowlist field names in get_field, and pass only pre-stringified primitives.

open as a page

How do you defend json.loads against a deeply nested untrusted document?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Bound the input before parsing: cap the payload size and reject documents nested deeper than the contract allows. json.loads recurses once per level and raises RecursionError on deep input, and raising sys.setrecursionlimit() only moves the failure, it does not remove it.

open as a page

A shelve database on a shared volume caches a metrics scraper's results — what does the nightly run risk?

level: seniorimportance: should knowfreq 25%

basics

~20 s

shelve stores every value as a pickle, so reading a key unpickles it. Anyone able to write that file — any workload sharing the volume — gets code execution inside the scraper the moment the nightly run reads its cache.

open as a page

When does sqlite3.Cursor.executemany beat a loop of execute calls, and what will it refuse to run?

level: seniorimportance: should knowfreq 38%

basics

~10 s

executemany compiles the statement once and binds each parameter set from any iterable, including a generator, so a large batch streams without being materialized. It refuses statements that return rows, raising sqlite3.ProgrammingError.

open as a page

Which bytes.decode error handler silently loses characters in an inventory sync that ingests product codes at a 1,200-request-per-minute peak, and what should replace it?

level: seniorimportance: should knowfreq 30%

basics

~20 s

errors='ignore' drops every undecodable byte without raising, so a product code arrives shortened and two different codes can collapse into one. Decode with the default errors='strict' at the boundary and reject or quarantine the payload that raises UnicodeDecodeError.

open as a page

A chat-transcript archiver using xml.etree.ElementTree on user-supplied exports intermittently blows its 92nd-percentile latency budget — how do you confirm entity expansion is the cause and harden the parse?

level: seniorimportance: should knowfreq 24%

basics

~20 s

Correlate parse time and memory against input size: a bomb is tiny input with huge cost. Catch xml.etree.ElementTree.ParseError and read its message, check xml.parsers.expat.EXPAT_VERSION, then move untrusted parses behind a hardened library, a byte cap, and a killable worker process.

open as a page

What do Python 3.14 t-strings change about interpolating untrusted values?

level: middleimportance: nice to knowfreq 14%

basics

~20 s

A t-string literal evaluates to a string.templatelib.Template rather than a str: the literal chunks and the interpolated values stay separate until a renderer joins them, so the renderer can escape each value for its destination first.

open as a page

Why does hash('abc') differ between two Python processes, and what does PYTHONHASHSEED=0 cost?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

CPython mixes a random per-process seed into the hash of str and bytes so an attacker cannot precompute keys that all collide in a dict. PYTHONHASHSEED=0 turns that randomization off, and removing it re-opens the collision-flood attack.

open as a page

Is marshal.loads a safer alternative to pickle.loads for untrusted bytes?

level: middleimportance: nice to knowfreq 12%

basics

~20 s

No. marshal is CPython's internal format for compiled bytecode, and its documentation says it is not secure against maliciously constructed data. It carries code objects, and its reader offers no promise of failing cleanly on hostile input.

open as a page

In xml.sax, what do feature_external_ges and feature_external_pes control?

level: middleimportance: nice to knowfreq 14%

basics

~20 s

They are the SAX feature identifiers for external general and external parameter entities. Enabling feature_external_ges makes the reader fetch whatever an entity's system identifier points at; the expat-backed reader refuses any attempt to enable feature_external_pes.

open as a page

How do you write a custom tarfile extraction filter on top of tarfile.data_filter?

level: seniorimportance: nice to knowfreq 12%

basics

~20 s

A filter is any callable taking the TarInfo member and the destination path. Return a member to extract it, return None to skip it, or raise to abort the whole extraction. Call tarfile.data_filter first to keep the standard checks, then add your own policy.

open as a page