skip to content

JSON with the json Module

load and loads read, dump and dumps write, and hooks cover the objects that are not serializable. Interviewers ask how you serialize a datetime or a Decimal — that is default= and JSONEncoder.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

How do json.load, json.loads, json.dump and json.dumps differ?

level: juniorimportance: must knowfreq 75%

answer

  1. Four names, only two jobs
  2. One extra letter is the whole difference
  3. One pair wants an open file
  4. s is for string, not stream
  5. load reads, dump writes

basics

~20 s

The trailing s means string. json.loads parses JSON text already in memory and json.load reads it from an open file object; json.dumps returns a JSON string, while json.dump writes that text straight into a file object.

solid answer

~40 s

`json.loads` takes JSON text you already hold — a `str`, `bytes` or `bytearray` — and returns Python objects. `json.load` does the same parse but takes any object with a `.read()` method, so you hand it an open file. Going the other way, `json.dumps` serializes a Python object to a `str`, and `json.dump` writes that same text into any object with a `.write()`. The `s` suffix means *string*, not *stream*: the file-object pair is the one **without** it. Because `json.dump` emits `str`, the byte encoding is the file's and is fixed by `open(..., encoding=...)`, while `json.loads` given `bytes` autodetects UTF-8, UTF-16 or UTF-32. Neither file variant is incremental — `json.load` still reads the whole document and builds the entire object graph in memory.

code

pycon · 11 lines
pycon
>>> import io, json
>>> json.loads('{"sent": 4, "ok": true}')
{'sent': 4, 'ok': True}
>>> buf = io.StringIO()
>>> json.dump({"sent": 4, "ok": True}, buf)
>>> buf.getvalue()
'{"sent": 4, "ok": true}'
>>> buf.seek(0)
0
>>> json.load(buf)
{'sent': 4, 'ok': True}

go deeper

for a junior

Recall which two names take a file object and which two take a string, and that the s stands for string. Being able to type a read-modify-write round trip from memory, with encoding='utf-8' on both opens, is the bar here.

for a middle

Explain that json.dump emits str into whatever you pass it, so the encoding is chosen by open() and not by the module, and that json.load builds the entire object graph in memory rather than parsing incrementally.

for a senior

Be ready to say where this four-call API stops being enough: multi-gigabyte documents, newline-delimited records parsed one line at a time, or a memory ceiling that rules out holding the parsed graph at all.

for a principal

Own the interchange-format decision at a service boundary: JSON's size and parse cost, what it cannot express without a convention, and whether the team funnels every encode and decode through one shared helper instead of scattering raw json calls.

## Four names, two jobs The `json` module exposes its work through four top-level callables, and the whole naming scheme rests on one letter. `load` and `loads` **decode**: JSON text in, Python objects out. `dump` and `dumps` **encode**: Python objects in, JSON text out. The trailing `s` on `loads` and `dumps` stands for **string** — those are the in-memory variants. The two names without the `s` are the file-object variants. Candidates routinely misread the `s` as "stream", which inverts the whole API in their head, so it is worth fixing the mnemonic early. ## The decode side `json.loads(text)` accepts a `str` and returns the corresponding Python object graph. It also accepts `bytes` and `bytearray` directly, autodetecting UTF-8, UTF-16 or UTF-32 from the leading bytes, which is why you can hand it a response body without decoding it yourself. `json.load(fp)` accepts a **file object** — anything with a `.read()` method, which includes a file returned by `open()`, an `io.StringIO`, an `io.BytesIO`, or a network stream wrapper. It calls `fp.read()` once and then does exactly what `loads` does. Passing a *filename* is the classic beginner error: `json.load("data.json")` raises `AttributeError: 'str' object has no attribute 'read'`, because a `str` has no `.read()`. Open the file first, or read the text yourself and call `loads`. ## The encode side `json.dumps(obj)` returns a `str` of JSON text — never `bytes`. `json.dump(obj, fp)` serializes and calls `fp.write(...)` with the resulting text; it returns `None`. Note the argument order: the object comes first, the destination second. Writing `json.dump(fp, obj)` is a frequent slip and usually fails with a confusing serialization error about the file object. Because `json.dump` writes `str`, it never chooses an encoding itself. The bytes on disk are decided by the file you opened: ```python import json from pathlib import Path with open("digest.json", "w", encoding="utf-8") as fp: json.dump({"sent": 4}, fp) print(json.loads(Path("digest.json").read_text(encoding="utf-8"))) ``` Relying on the platform default encoding is how a file that reads fine on one machine becomes mojibake on another; pass `encoding="utf-8"` explicitly. If you genuinely need bytes, encode the result of `dumps` yourself: `json.dumps(obj).encode("utf-8")`. ## Formatting knobs are shared All the presentation options belong to the encoder, so `dump` and `dumps` accept the same keywords: `indent` for pretty-printing (an integer or a string), `sort_keys=True` for deterministic key order, `separators` to strip whitespace for the most compact possible output, and `ensure_ascii` (default `True`) which escapes every non-ASCII character as a `\uXXXX` sequence. Setting `ensure_ascii=False` keeps the characters literal, which is smaller and far more readable for non-English text — as long as the file is actually written as UTF-8. ## What the file variants do *not* do The most common wrong assumption at this level is that `json.load` streams. It does not. It reads the entire document and builds the entire Python object graph before returning, so a 2 GB file needs the file's bytes plus the parsed structure in memory at once. `json.dump` is slightly better on the write side — it writes incrementally through the encoder's chunk iterator rather than building one giant string — but the object graph is still fully in memory before you call it. When documents outgrow RAM, the answer is a different shape entirely: newline-delimited records parsed one line at a time with `loads`, or a pull parser, not a different member of this family. ## Round trip in practice The read-modify-write pattern is three lines and worth being able to type from memory: ```python import json with open("config.json", encoding="utf-8") as fp: config = json.load(fp) config["recipients"] = 4 with open("config.json", "w", encoding="utf-8") as fp: json.dump(config, fp, indent=2, sort_keys=True) ``` One anti-pattern to name: `fp.write(json.dumps(obj))` is not wrong, merely redundant — that is precisely what `json.dump` does. The reverse mistake, `json.dump(json.dumps(obj), fp)`, double-encodes and produces a JSON *string* containing JSON text, which the consumer then has to parse twice. ## From the shell The module ships a command line for exactly this: `python -m json --indent 2 file.json` on Python 3.14, and `python -m json.tool`, which has existed far longer and still works. Both parse and re-emit the document, so they double as a fast syntax check on a file you were handed.

  • Can you pass a filename to json.load instead of an open file object?
    No. `json.load` calls `.read()` on whatever you give it, so a `str` path raises `AttributeError: 'str' object has no attribute 'read'`. Open the file first, or read the text yourself with `pathlib.Path.read_text()` and call `json.loads`. The same holds in reverse: `json.dump` calls `.write()`, so it needs a writable file object, not a path.
  • Where is the byte encoding decided when you write JSON to a file?
    In `open()`. `json.dump` produces `str` and hands it to the file object, so the file's `encoding=` argument decides the bytes — pass `encoding='utf-8'` explicitly rather than trusting the platform default. Coming back in, `json.loads` accepts `bytes` and autodetects UTF-8, UTF-16 or UTF-32, but if you read a text file yourself, `open()` has already decoded it.
  • What does ensure_ascii do, and when would you turn it off?
    `ensure_ascii=True` is the default and escapes every non-ASCII character as `\uXXXX`, so the output is pure ASCII and safe through any transport. Setting `ensure_ascii=False` emits the characters literally, which is smaller and much easier to read for non-English text — provided the file or socket is genuinely UTF-8. Both forms decode back to the identical `str`.

Think of a photocopier with two input trays: one takes a sheet you are already holding, the other takes the document feeder. Same copying mechanism, different way of supplying the paper.

saying these in an interview costs you the question

  • Says json.load accepts a filename string
  • Thinks json.dumps returns bytes rather than str
  • Calls json.dump with the file object first
  • Believes json.load streams instead of loading everything
  • Reads the s in loads as stream
  • Writes json.dump(json.dumps(obj), fp) and double-encodes

context

open as a page

How do you make json.dumps serialize a datetime or a Decimal?

level: middleimportance: must knowfreq 66%

basics

~20 s

Give the encoder a fallback: pass default=fn, a callable json.dumps invokes only for objects it cannot handle, or subclass json.JSONEncoder, override default and pass it as cls=. Return a substitute — isoformat() for a datetime, str() for a Decimal.

open as a page

How does json.dumps map Python types, and what does it refuse?

level: middleimportance: should knowfreq 60%

basics

~20 s

dict becomes an object; list and tuple both become arrays; str becomes a string; int and float become numbers; True, False and None become true, false and null. Everything else — set, bytes, datetime, Decimal — raises TypeError.

open as a page

How do you diagnose a json.JSONDecodeError in a production ingest path?

level: seniorimportance: should knowfreq 44%

basics

~20 s

json.JSONDecodeError subclasses ValueError and carries msg, doc, pos, lineno and colno. Catch that class specifically, log the message plus a short slice of doc around pos rather than the whole payload, and keep malformed text separate from valid-but-wrong-shape data.

open as a page