skip to content

When should you parse a string with json.loads instead of ast.literal_eval?

level: juniorimportance: should knowfreq 46%

answer

  1. Match the parser to the format
  2. One of them reads Python, not data
  3. Three JSON words it cannot spell
  4. Error messages with a position
  5. Roughly thirty times the cost

basics

~20 s

Whenever the string is JSON — anything produced by another system. json.loads speaks the actual format, including true, false and null, reports errors with a position, and runs roughly thirty times faster than parsing Python source into a tree.

solid answer

~40 s

Parse a format with that format's parser. `ast.literal_eval` reads *Python source*, so it is right only when Python produced the string — a stored `repr`, a literal typed into a field. JSON from an HTTP body or a queue goes to `json.loads`: it accepts `true`/`false`/`null`, which `ast.literal_eval` rejects as names, it enforces double quotes rather than silently tolerating Python spellings, it raises `json.JSONDecodeError` with a position, and on a 2 MB document it is about 30x faster on 3.14 because it is a C scanner rather than a full tokenizer, parser and tree walk. TOML configuration goes to `tomllib` (3.11+), ini files to `configparser`, delimited text to `csv`. And `pickle.loads` is never the answer for anything crossing a trust boundary.

code

python · 9 lines
python
import ast
import json

payload = '{"lang": "fr", "hits": 12, "stale": null}'
print(json.loads(payload))
try:
    ast.literal_eval(payload)
except ValueError as exc:
    print("literal_eval refused:", str(exc)[:44])

go deeper

for a junior

Learn the one-step rule: if the string came from another system it has a format, and that format has a parser. Only reach for the literal evaluator when Python itself wrote the string.

for a middle

Explain the concrete mismatches — true/false/null are names to Python, quoting and trailing-comma rules differ — and know that JSON decoding gives you an error position while an AST walk does not.

for a senior

Bring the cost and failure-mode argument: a full tokenize-parse-walk is roughly an order of magnitude slower, and a lenient parser on foreign input hides producer bugs until a schema change surfaces them.

for a principal

Set the policy that data crossing a trust boundary is read by a format parser with a schema check behind it, and that pickle never appears on that path regardless of convenience.

### The question behind the question Both `json.loads` and `ast.literal_eval` turn a string into a dict. Choosing between them is really choosing **what the string is**, and the answer follows from that in one step: * The string is **JSON** — it came from an HTTP body, a message queue, a log line, another language's serializer → `json.loads`. * The string is **TOML** configuration → `tomllib.loads` (or `tomllib.load` on a file opened in binary mode). Added in **3.11**, read-only by design; there is no writer in the standard library. * The string is **ini-style** configuration → `configparser`. * The string is **delimited tabular text** → the `csv` module, which handles quoting, embedded newlines and delimiters correctly where `str.split(",")` does not. * The string is **Python source that Python itself produced** — a stored `repr`, a literal typed into a form field → `ast.literal_eval`. ### Why JSON in particular is the wrong job for the literal evaluator JSON and Python literals *look* alike, and that resemblance is a trap. They differ exactly where real data lives: * `true`, `false` and `null` are `Name` nodes to Python's parser, and `ast.literal_eval` refuses every `Name`. A payload with a single `null` in it raises `ValueError`. * JSON requires double quotes; Python accepts either, so a Python-literal reader silently tolerates input no JSON producer would emit — the class of leniency that hides bugs until the day a different producer appears. * JSON allows no trailing comma; Python does. Again, `ast.literal_eval` is the lenient one. * `\/` escapes, surrogate pairs and JSON's number grammar are handled by a real JSON decoder and not by Python's tokenizer. So the accidental cases where `ast.literal_eval` happens to read a JSON document are the ones with no booleans and no nulls — which is to say, until the data changes. ### Errors and speed `json.loads` raises `json.JSONDecodeError` (a `ValueError` subclass) carrying `msg`, `pos`, `lineno` and `colno` — you can point at the byte that broke. `ast.literal_eval` gives you either a `ValueError` quoting an AST node dump or a `SyntaxError`, neither of which is a useful error message for a caller who sent JSON. Speed is not close either. Measured on **3.14** with a 2.2 MB JSON document of 50,000 small objects: `json.loads` about **9 ms**, `ast.literal_eval` about **290 ms** — roughly **30x**. `json.loads` runs a C scanner straight to objects; `ast.literal_eval` runs the full Python tokenizer and parser, materializes an entire AST, then walks it in Python. On a hot path that difference is the whole budget. ### The other direction Going the other way, the symmetry holds: `json.dumps` produces JSON for other systems; `repr` produces a debugging string that only sometimes happens to be a valid literal. If you find yourself writing `repr(value)` into a database column so that `ast.literal_eval` can read it back, you have chosen a serialization format by accident, and it is one with no schema, no versioning and no cross-language readers. ### And the one that is never the answer for foreign data `pickle.loads` will read almost any Python object graph, and it will also execute code embedded in the stream. It is a fine choice for data your own process wrote and can vouch for, and never a choice for anything crossing a trust boundary. "It has to be Python objects, so I pickled it" is a remote-code-execution report waiting to be filed; `json` plus an explicit conversion step is the boring correct answer. ### Tabular text deserves its own mention The `csv` module is the case people skip most often, because splitting a line on commas looks like it works. It works until a field contains a comma inside quotes, an embedded newline, a doubled quote character, or a different delimiter in a regional export. `csv.reader` and `csv.DictReader` implement the quoting rules, handle the embedded-newline case when the file is opened with `newline=""`, and let you state the dialect explicitly instead of discovering it from data. Neither `str.split` nor a literal evaluator has any notion of a quoted field, so both are wrong for delimited text in the same way and for the same reason: the format has structure that the shortcut cannot see. ### The rule to state in an interview Parse a format with that format's parser. `ast.literal_eval` is not a JSON reader that is a bit slower — it is a *Python source* reader that will silently accept the wrong things and loudly reject the right ones. Reach for it only when the string genuinely is Python, and reach for `json`, `tomllib`, `configparser` or `csv` the moment the string belongs to someone else's format.

  • A colleague says ast.literal_eval reads their JSON fine, so why change it?
    It works only while the documents contain no `true`, `false` or `null` and no JSON-specific escapes. The first payload with a boolean in it raises ValueError in production, far from the code that chose the parser. It is also lenient in the wrong direction, accepting single quotes and trailing commas that no JSON producer emits, so malformed input passes silently.
  • Which parser reads a TOML configuration file, and what is the call shape?
    `tomllib`, added in 3.11. Use `tomllib.load(f)` with the file opened in **binary** mode (`"rb"`), or `tomllib.loads(text)` for a string. It is read-only by design — the standard library ships no TOML writer — and it raises `tomllib.TOMLDecodeError` on bad input.

saying these in an interview costs you the question

  • Treats JSON and Python literals as the same syntax
  • Uses ast.literal_eval as a general-purpose data parser
  • Splits CSV text on commas instead of using csv
  • Reaches for pickle to read data from another system
  • Assumes the two parsers cost about the same

context