skip to content

What does the compression.zstd module added in Python 3.14 give you?

level: middleimportance: nice to knowfreq 15%

answer

  1. New in 3.14, from a PEP
  2. A package namespace, not just a module
  3. Its API mirrors an existing codec's
  4. Default level is far below the maximum
  5. Trained dictionaries for many small records

basics

~10 s

Python 3.14 adds Zstandard to the standard library as compression.zstd, with the familiar compress/decompress/open shape plus ZstdFile, incremental compressor and decompressor objects, trained dictionaries, and a tuning-parameter enum.

solid answer

~40 s

Before 3.14 Zstandard meant a third-party binding; PEP 784 landed it in the stdlib as `compression.zstd`, inside a new `compression` package that also re-exports the existing codecs as `compression.gzip`, `compression.bz2`, `compression.lzma` and `compression.zlib`. The API deliberately mirrors `lzma`: `compress`/`decompress` for one-shot work, `open` and `ZstdFile` for streaming files, `ZstdCompressor`/`ZstdDecompressor` for incremental work, and `ZstdError` for failures. Beyond that it exposes what zstd is actually known for — `ZstdDict` with `train_dict` and `finalize_dict` for trained dictionaries that make many small payloads compress well, and `CompressionParameter`/`DecompressionParameter` enums covering `nb_workers`, `window_log`, `checksum_flag` and friends. The default level is 3 (`COMPRESSION_LEVEL_DEFAULT`), and the valid range runs up to 22. `zipfile.ZIP_ZSTANDARD` lets archives use it too.

code

python · 11 lines
python
from compression import zstd

blob = b"sensor=7 celsius=21.5\n" * 1000

with zstd.open("batch.zst", "wb", level=3) as f:
    f.write(blob)

with zstd.open("batch.zst", "rb") as f:
    print(len(f.read()))

print(zstd.get_frame_info(zstd.compress(blob)).decompressed_size)

go deeper

for a junior

Recall that Python 3.14 ships Zstandard in the standard library as compression.zstd, and that it opens files the same way gzip.open does. Knowing it exists and needs no dependency is enough here.

for a middle

Explain the API surface — one-shot functions, open/ZstdFile, the incremental compressor and decompressor — and the level model: default 3, range up to 22, and why reflexively maxing it is the deflate habit rather than a zstd one.

for a senior

Show what you would do with it in production: dictionary training for small repeated records, window_log_max and get_frame_info as guards on data you did not produce, and a fallback plan for interpreters older than 3.14.

for a principal

Own the rollout decision: whether adopting a 3.14-only codec is worth pinning your minimum interpreter, who owns the trained dictionaries and their versioning, and what the migration story is for archives already written in another format.

### What actually shipped Python 3.14 added Zstandard support to the standard library through PEP 784. The module is `compression.zstd`, and it arrives inside a new top-level `compression` package that also carries alias modules for the codecs that already existed: `compression.gzip`, `compression.bz2`, `compression.lzma` and `compression.zlib`. The old top-level names are unchanged and are not going anywhere — the package is a namespace for the family, not a migration. ### The API is deliberately familiar Nothing about the shape is new; that is the point. It follows `lzma` almost exactly: * `compression.zstd.compress(data, level=None, options=None, zstd_dict=None)` and the matching `decompress` for one-shot use. * `compression.zstd.open(file, mode='rb', *, level=None, options=None, zstd_dict=None, encoding=None, errors=None, newline=None)` — binary by default, text with `'rt'`/`'wt'`, exactly like `gzip.open`. * `ZstdFile`, the binary file object `open` wraps. * `ZstdCompressor` and `ZstdDecompressor` for incremental streaming over data too large to hold. * `ZstdError`, raised for malformed frames and bad parameters. So the migration cost from `lzma` or `gzip` for a service that already streams is close to nothing. ### Levels, and the default that is not 9 `COMPRESSION_LEVEL_DEFAULT` is `3`, and the accepted range runs up to `22` (a `level` outside the valid range raises `ValueError`, which names the bounds). This is a different mental model from the deflate codecs, where reflexive `compresslevel=9` is common: zstd's low levels are already fast *and* competitive on ratio, and the high end is a distinct, much slower regime you opt into for archival data. Measure on your own payloads rather than assuming higher is better. ### Frame metadata `get_frame_info(data)` reads a frame header and tells you the `decompressed_size` the producer recorded and the `dictionary_id` it used, without decompressing anything. That is genuinely useful before you allocate: you can decide whether to accept a frame, or route it to a streaming path, from its header alone. ### Trained dictionaries The feature that most distinguishes zstd from the deflate family is dictionary compression. Ordinary compression finds redundancy *within* one payload, so a stream of small, similar records — a few hundred bytes each of the same telemetry shape — compresses badly, because each one is too short to build a model from. `train_dict(samples, dict_size)` builds a `ZstdDict` from a corpus of representative samples; `finalize_dict` refines a dictionary you already have. Pass it as `zstd_dict=` to `compress` and `decompress` and each small record is compressed against shared context it did not have to carry. The cost is coordination: producer and consumer need the same dictionary, and the frame records only its id, not its contents. ### Tuning parameters `CompressionParameter` and `DecompressionParameter` are `IntEnum`s naming the knobs the C library exposes, passed as an `options` mapping. The ones that come up in practice are `nb_workers` (multi-threaded compression inside a single call), `window_log` (how far back the matcher looks — larger finds more redundancy at more memory), `checksum_flag` (write a content checksum into the frame), and `enable_long_distance_matching`. On the decompression side `window_log_max` caps how much memory a frame may make you allocate, which is the knob you set when decompressing data you did not produce. `Strategy` selects the matcher's algorithm. ### Beyond the module The archive formats learned zstd too: `zipfile.ZIP_ZSTANDARD` is a compression method you can pass when writing entries, and `shutil`'s archive-format list gained `zstdtar`. ### The caveat worth stating `compression.zstd` binds to the system's Zstandard library, so a CPython built without it will raise `ImportError` on the import — a possibility on stripped or unusual builds, and a reason to keep the import at the top where it fails loudly rather than deep inside a request path. And it is 3.14: code that must also run on 3.13 or earlier needs a fallback, whether that is a third-party binding or simply using `gzip` there. ### What an interviewer is checking Not encyclopedic knowledge of the parameter enum. They want to hear that you know zstd is in the stdlib as of 3.14, that its API is the same shape as the codecs you already use, that the default level is low on purpose, and ideally that you can say what a trained dictionary is for — because that is the part that changes what is possible rather than merely what is available.

  • What problem does a trained ZstdDict solve that a higher level cannot?
    Short payloads. Compression finds redundancy inside one payload, so a few hundred bytes of a repeating record shape has almost nothing to exploit and raising the level barely helps. `train_dict(samples, dict_size)` builds shared context from a corpus of representative records, and compressing each record against it removes the redundancy that lives *between* records. The price is that producer and consumer must both hold the same dictionary.
  • Which parameter would you set when decompressing zstd frames from an untrusted producer?
    `DecompressionParameter.window_log_max`, which caps the window — and therefore the memory — a frame is allowed to make you allocate. A frame header can request a very large window; without a cap you honour it. Pairing that with `get_frame_info`, which reports the recorded decompressed size before you decompress anything, lets you reject a frame on its header instead of on your memory limit.
  • Does importing compression.zstd always succeed on 3.14?
    No. The module binds to the Zstandard library, so a CPython built without it raises `ImportError` at import. That is unusual on mainstream builds but real on stripped ones, which is an argument for importing at module top level where the failure is immediate and obvious, rather than lazily inside a code path that only some requests reach.

saying these in an interview costs you the question

  • Claims zstd was always in the standard library
  • Assumes the default level is 9 like gzip's
  • Thinks the compression package replaces the gzip module
  • Confuses a trained dictionary with a Python dict
  • Expects a higher level to fix tiny-payload ratios
  • Names it zstd rather than compression.zstd

context