skip to content

Why does tomllib.load require a file opened in binary mode ('rb')?

level: juniorimportance: must knowfreq 50%

answer

  1. TOML has exactly one legal encoding
  2. who decodes the bytes here
  3. text mode returns str, not bytes
  4. the parser calls decode() itself
  5. open(path, 'rb') or TypeError

basics

~20 s

tomllib.load reads bytes and decodes them as UTF-8 itself, because a TOML document is UTF-8 by definition. A file opened in text mode has already been decoded with the platform encoding, so load raises TypeError; open the path with 'rb'.

solid answer

~40 s

The TOML format fixes its encoding: a TOML document is UTF-8. So `tomllib.load`, added in 3.11, takes a **binary** file object, calls `read()`, and decodes those bytes as UTF-8 itself. Hand it a text-mode file and `read()` returns `str`, the internal `decode()` call fails, and you get `TypeError: File must be opened in binary mode`. The rule exists so the platform's locale encoding can never decide how a config file is read — the same file parses identically on every machine. If you already hold a string, call `tomllib.loads` instead; if you hold a path, `Path.read_bytes()` feeds `loads` after a decode, or just `open(path, 'rb')`. `tomllib` is parse-only: through 3.14 the standard library has no TOML writer. Malformed input raises `tomllib.TOMLDecodeError`, a subclass of `ValueError`.

code

python · 11 lines
python
import io
import tomllib

document = b"[extract]\nworkers = 4\nprobe_streams = true\n"
print(tomllib.load(io.BytesIO(document)))
print(tomllib.loads(document.decode()))

try:
    tomllib.load(io.StringIO("workers = 4"))
except TypeError as exc:
    print("TypeError:", exc)

go deeper

for a junior

Recall the one-liner: TOML is UTF-8, so the parser decodes the bytes itself and needs open(path, 'rb'). Be able to say what goes wrong with text mode — a TypeError, not a silent misparse.

for a middle

Explain the mechanism: load calls read(), expects bytes, decodes UTF-8, and a str has no decode. Know loads for in-memory text, the TOML-to-Python type mapping, and that tomllib cannot write.

for a senior

Show where this bites in production: locale-dependent decoding making a config file machine-specific, validating the returned dict rather than trusting it, and designing around the missing writer when a feature wants to persist settings.

for a principal

Own the format choice. Argue when typed, read-only TOML beats INI or JSON for operator-edited configuration, what the absent writer implies for any save-settings feature, and how you keep configuration parsing identical across every environment you deploy to.

## Where `tomllib` came from Python 3.11 added `tomllib` to the standard library (PEP 680), after TOML had already become the default shape of Python tool configuration. It is deliberately small: `tomllib.loads(s)` parses a `str`, `tomllib.load(fp)` parses a **binary** file object, and that is the whole public surface apart from `tomllib.TOMLDecodeError`. Before 3.11 you needed a third-party parser for the same job. ## The encoding is a property of the format, not of your machine TOML's specification states that a TOML document is a valid UTF-8 encoded Unicode document. The encoding is not a per-file choice and not a platform detail — it is fixed by the format. `tomllib` therefore takes responsibility for decoding rather than letting the caller do it. `load` calls `fp.read()`, expects `bytes` back, and decodes them as UTF-8. Passing a text-mode file breaks that contract in two ways. The quiet failure is that Python has *already* decoded the file for you using the locale encoding (`locale.getencoding()`), which is UTF-8 on most Linux systems but historically something else on Windows. A config file containing a non-ASCII character then parses on your laptop and produces mojibake — or a `UnicodeDecodeError` — on a colleague's. The loud failure is what actually happens today: `load` tries to call `.decode()` on what it read, `str` has no such method, and the `AttributeError` is turned into `TypeError: File must be opened in binary mode, e.g. use `open('foo.toml', 'rb')``. The error message *is* the documentation. Note also that you cannot compromise by passing an encoding to a binary open: `open(path, 'rb', encoding='utf-8')` raises `ValueError`, because binary mode takes no encoding argument. Binary really means binary; the parser owns the decode. ```python import tomllib with open("settings.toml", "rb") as handle: # 'rb', never 'r' settings = tomllib.load(handle) ``` If you already have the text — from a network response, a test fixture, or a heredoc — use `loads`. If you want a `load`-shaped call in a test without touching the filesystem, wrap the bytes in `io.BytesIO`. ## What comes back A plain `dict`. Keys are always `str`; values map from TOML types to Python ones: string to `str`, integer to `int`, float to `float`, boolean to `bool`, array to `list`, table to `dict`, and the date-time types to `datetime.datetime`, `datetime.date` and `datetime.time`. An offset date-time carries a `tzinfo`; a local one does not. Nested tables are nested dicts, so a `[tool.section]` header becomes `data["tool"]["section"]`. Two hooks matter in practice. `parse_float` lets you swap the float constructor — passing `decimal.Decimal` keeps exact decimal values instead of binary floats. And because the result is an ordinary dict, nothing validates it: `tomllib` guarantees the document is well-formed TOML, not that it contains the keys your program needs. That validation is yours to write. ## Read-only, on purpose There is no `dump` or `dumps` in `tomllib`, and none is coming from PEP 680, which scoped the addition to reading. The reason is that writing TOML *well* is a round-tripping problem: preserving comments, key ordering, table style and quoting is a design space with more than one right answer, and the standard library declined to pick one. The practical consequence is architectural: with the standard library alone, your program reads configuration and never edits it. A "save my settings" feature needs a different format — or a third-party writer. This is the sharpest contrast with `configparser`, which does have a `write` method and can round-trip an INI file (losing comments). ## Errors `tomllib.TOMLDecodeError` subclasses `ValueError`, so a broad `except ValueError` around configuration loading catches both a malformed document and a downstream conversion failure. In 3.14 that exception carries structured detail — the message, the document, and the position of the problem — instead of only free-form text, which makes it far easier to point a user at the offending line. ## What good practice looks like Open with `open(path, "rb")` inside a `with` block; do not pass `encoding=`; catch `tomllib.TOMLDecodeError` at the boundary and report the file name alongside it; and treat the returned dict as untyped input to be validated, not as your settings object.

  • You already have the TOML text in a string. What do you call?
    `tomllib.loads(text)`. `load` exists only for the file case and its whole extra job is the UTF-8 decode. If you want a `load`-shaped call in a test, wrap the bytes in `io.BytesIO` rather than writing a temporary file.
  • Why is there no tomllib.dump in the standard library?
    PEP 680 scoped the addition to reading. Writing TOML well means deciding how to preserve comments, key order, table style and quoting, and there is no single right answer, so the standard library left serialisation to third-party writers. Practically: with the standard library your program reads config and never rewrites it.
  • What does tomllib raise on a malformed document, and what can you do with it?
    `tomllib.TOMLDecodeError`, a subclass of `ValueError`, so a broad `except ValueError` at the config boundary catches it. On 3.14 it carries the message, the document and the failing position, so you can report the exact line back to whoever wrote the file.

Handing a text-mode file to tomllib is like giving a translator a page someone already translated badly: the parser wants the original bytes so it can do the one decoding the format allows.

saying these in an interview costs you the question

  • Says open(path) then tomllib.load works fine
  • Thinks tomllib can also write TOML files
  • Passes encoding='utf-8' and expects text mode to work
  • Believes the platform locale encoding decides how TOML is read
  • Confuses tomllib.load with tomllib.loads for a str
  • Expects a TOMLDecodeError not to be a ValueError

context