How does gzip.open in text mode differ from its default binary mode?
answer
- One call, two independent layers
- Its default mode is not the builtin's
- Bytes come out unless you ask otherwise
- Three arguments only text mode accepts
- 'rt' plus encoding='utf-8'
basics
~20 sgzip.open defaults to mode 'rb' and yields bytes. Adding a 't' — 'rt' or 'wt' — wraps the decompressed stream in a text layer that decodes to str, and only then may you pass encoding, errors or newline.
solid answer
~40 s`gzip.open(filename, mode='rb', ...)` differs from the builtin `open`, whose default mode is text. In binary modes you read and write `bytes` exactly as they sit inside the member; in text modes (`'rt'`, `'wt'`, `'at'`) gzip wraps the decompressed byte stream in a text layer, so you read and write `str`. Passing `encoding`, `errors` or `newline` in a binary mode raises `ValueError`, and omitting `encoding` in a text mode falls back to the locale encoding — which is why you should write `gzip.open(path, 'rt', encoding='utf-8')` explicitly for any file that travels between machines. `compresslevel` is meaningful only for write modes. `bz2.open`, `lzma.open` and, on 3.14, `compression.zstd.open` follow the same mode grammar, while the underlying `gzip.GzipFile`, `bz2.BZ2File` and `lzma.LZMAFile` objects are binary-only.
code
python · 10 linesimport gzip
with gzip.open("readings.jsonl.gz", "wt", encoding="utf-8") as f:
f.write('{"sensor": 7, "celsius": 21.5}\n')
with gzip.open("readings.jsonl.gz", "rb") as f:
print(f.read()) # b'{"sensor": 7, "celsius": 21.5}\n'
with gzip.open("readings.jsonl.gz", "rt", encoding="utf-8") as f:
print(f.readline().rstrip()) # {"sensor": 7, "celsius": 21.5}go deeper
Recall the default: gzip.open opens in binary and gives you bytes. Be ready to say what you change to get str back, and to name the argument you would add alongside it.
Explain the two-layer model — decompression produces bytes, a separate text layer decodes them — and why encoding, errors and newline raise ValueError in binary mode while compresslevel matters only on write.
Demonstrate the operational habit: pin encoding='utf-8' on every text-mode open so archives read back identically on another host, and distinguish a BadGzipFile container failure from a UnicodeDecodeError decode failure when triaging a bad file.
Own the convention across services: decide whether your archived payloads are byte streams with a documented codec or text files, write it down once, and avoid mixed conventions that make a shared reader guess between locale and UTF-8.
### Two layers, not one A gzip file is a byte container. Nothing inside it says "this is UTF-8 text" — it holds a compressed byte sequence plus a header and a CRC32 trailer. So decompression can only ever produce `bytes`. Text is a second, independent decision: which codec turns those bytes into `str`. `gzip.open` lets you make both decisions in one call. Its signature is `gzip.open(filename, mode='rb', compresslevel=9, encoding=None, errors=None, newline=None)`. In a binary mode it returns a `gzip.GzipFile` (a binary file object). In a text mode it builds that same object and then wraps it in a text layer, exactly as `io.TextIOWrapper` does for a raw binary stream. What you get back reads and writes `str`, and `readline()` splits on universal newlines rather than on a literal `b'\n'`. ### The default that surprises people The builtin `open` defaults to `'r'`, meaning **text**. `gzip.open` defaults to `'rb'`, meaning **binary**. The same is true of `bz2.open`, `lzma.open` and, on Python 3.14, `compression.zstd.open`. Code that was written against the builtin and then had `gzip.` pasted in front of it starts failing with `TypeError: a bytes-like object is required, not 'str'` on the first `write`, or produces `b'...'` reprs on the first `print`. The fix is not `str(chunk)` — that renders the `b'...'` repr into your output — it is either decoding deliberately with `chunk.decode('utf-8')` or asking for text mode in the first place. ### encoding, errors and newline are text-only These three parameters describe the text layer, so they are rejected outright in binary mode: ```pycon >>> import gzip >>> gzip.open('readings.gz', 'rb', encoding='utf-8') Traceback (most recent call last): ... ValueError: Argument 'encoding' not supported in binary mode ``` That error is a feature: it stops you believing a decode happened when none did. In text mode, leaving `encoding=None` means the interpreter's locale encoding is used, which is not the same on every machine your file lands on. For anything durable — a telemetry archive read back on another host — name the codec: `encoding='utf-8'`. ### compresslevel is write-only `compresslevel` (0–9, default 9) is passed straight to the deflate compressor and is simply ignored on read: the level used to *write* a member is not recorded as a tuning knob, and the decompressor infers everything it needs from the stream. Setting it on a read open is harmless but meaningless, and reaching for level 9 by reflex is often the wrong tradeoff — see the codec-and-level question on this topic. ### GzipFile itself is binary `gzip.GzipFile` has no text mode at all. If you need `str` over an already-open binary object, wrap it yourself: ```python import gzip, io with gzip.GzipFile('readings.gz', 'rb') as raw: with io.TextIOWrapper(raw, encoding='utf-8') as text: first = text.readline() ``` That is exactly what `gzip.open` does for you in `'rt'`. ### Filenames and file objects The first argument may be a path (`str`, `bytes` or a `os.PathLike`) *or* an already-open binary file object — a socket-backed stream, an `io.BytesIO`, or the body of a network response. That is how you decompress without ever touching the filesystem: ```python import gzip, io blob = gzip.compress(b'{"sensor": 7}\n') with gzip.open(io.BytesIO(blob), 'rt', encoding='utf-8') as f: line = f.readline() ``` ### Errors you will actually see Opening a file that is not gzip raises `gzip.BadGzipFile` on the first read, not at open time, because the header is only inspected when bytes are pulled. A truncated file raises `EOFError`. A file whose *bytes* decompress fine but whose *text* is not valid in the codec you named raises `UnicodeDecodeError` from the text layer — and the distinction matters when you are diagnosing: a `BadGzipFile` means the container is wrong, a `UnicodeDecodeError` means the container was fine and your encoding guess was not. ### Appending, and files with several members A gzip file may hold several compressed members back to back, and readers concatenate them transparently. That is what `'ab'` gives you: `gzip.open(path, 'ab')` starts a *new* member rather than reopening the previous one, and a later `gzip.open(path, 'rb').read()` returns all members joined, as if the file had been written in one pass. It is a genuinely useful shape for an append-only capture — each batch is a self-contained member — and it explains why the compressed file grows by a fresh header and trailer every time you append. ### The habit worth forming Decide, per call, whether you are moving bytes or moving text. If bytes, stay in `'rb'`/`'wb'` and never call `str()` on what comes out. If text, say `'rt'`/`'wt'` and always name `encoding`. Almost every gzip bug that reaches a code review is one of those two rules skipped.
- What exception tells you the file you opened with gzip.open is not actually gzip data?`gzip.BadGzipFile`, raised on the first read rather than at open time, because the header is only parsed when bytes are pulled. A file that starts as valid gzip but ends early raises `EOFError` instead. Both are container-level failures; a `UnicodeDecodeError` from the same call means the container was fine and the text encoding you named was wrong.
- Can you hand gzip.open something other than a filename?Yes — the first argument accepts a path or an already-open binary file object, so you can decompress an `io.BytesIO` holding a response body, or a pipe, without writing to disk. In text mode gzip still wraps the result in a text layer, so `gzip.open(io.BytesIO(blob), 'rt', encoding='utf-8')` gives you `str` lines straight from memory.
- Does compresslevel do anything when you open for reading?No. `compresslevel` configures the deflate compressor, so it only affects write and append modes. A decompressor reads whatever parameters the stream itself encodes, so the level a file was written with is not something you re-specify on read. Passing it on a read open is accepted and ignored.
saying these in an interview costs you the question
- Thinks gzip.open defaults to text mode like the builtin open
- Calls str() on decompressed bytes instead of decoding them
- Passes encoding in binary mode and expects it to apply
- Believes the gzip file itself records a text encoding
- Omits encoding in text mode and assumes UTF-8 everywhere
- Expects gzip.GzipFile to support a 'rt' mode