skip to content

Binary Data and Buffers

When the payload is bytes, not text: struct frames records, base64 and hex make them printable, gzip or zstd shrink them, uuid gives a byte-shaped id. Interviewers use these for protocol work.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

16

Why does base64.b64encode reject a str, and what do you write instead?

level: juniorimportance: must knowfreq 60%

answer

  1. Two layers: characters and bytes
  2. The function works below the text layer
  3. A conversion brackets the call on each side
  4. str.encode in, .decode('ascii') out
  5. Decoding is the lenient direction

basics

~10 s

base64.b64encode is a bytes-to-bytes function: it needs a bytes-like object and returns bytes. Encode text first with str.encode('utf-8'), then call .decode('ascii') on the result to get a str for JSON.

solid answer

~40 s

Base64 is a transformation over **bytes**, so `base64.b64encode` takes a bytes-like object and returns bytes; handing it a `str` raises `TypeError: a bytes-like object is required, not 'str'`. The idiom is a sandwich: `base64.b64encode(text.encode('utf-8')).decode('ascii')` — the inner `str.encode` chooses how the text becomes bytes, and the outer `bytes.decode` turns the Base64 output, which is ASCII by construction, into a `str` you can put in JSON. Decoding is deliberately asymmetric: `base64.b64decode` **does** accept a `str`, because Base64 text is ASCII by definition, and it raises `ValueError` if the string contains non-ASCII characters. What comes back from `b64decode` is bytes, and turning those back into text requires knowing the original codec — Base64 does not carry one.

code

python · 14 lines
python
import base64
import json

token = base64.b64encode("héllo".encode("utf-8")).decode("ascii")
print(token)                                     # aMOpbGxv
print(base64.b64decode(token).decode("utf-8"))   # héllo
print(json.dumps({"payload": token}))            # {"payload": "aMOpbGxv"}

for bad in (lambda: base64.b64encode("héllo"),
            lambda: json.dumps({"payload": base64.b64encode(b"hi")})):
    try:
        bad()
    except TypeError as exc:
        print("TypeError:", exc)

go deeper

for a junior

Memorise the sandwich: base64.b64encode(text.encode('utf-8')).decode('ascii'). Be ready to say that the function works on bytes and returns bytes, and that Base64 hides nothing from anyone.

for a middle

Explain the mechanics: three input bytes become four ASCII characters from a 64-symbol alphabet, and the decode side accepts a str because Base64 text is ASCII by definition. Name the TypeError and the ValueError you would actually see.

for a senior

Show where this bites in production — double encoding through a helper, a str() repr leaking into a payload, and a codec mismatch that Base64 round-trips happily while the text still arrives wrong. Put the text codec in the schema.

for a principal

Own the contract question: if binary is crossing a service boundary, decide whether it belongs inline at all, and if it does, pin the alphabet, the padding convention and the text codec in the interface definition rather than letting each client rediscover them.

### Two different conversions wearing the same word Python uses the verb "encode" for two unrelated jobs, and this question exists to check that you keep them apart. * **Text encoding** — `str.encode(codec)` turns characters into bytes according to a codec such as UTF-8, and `bytes.decode(codec)` reverses it. This is where meaning lives: the same characters produce different bytes under UTF-8, UTF-16 or Latin-1. * **Binary-to-text encoding** — `base64.b64encode` takes arbitrary bytes and re-expresses them in a 64-character ASCII alphabet so they survive a transport that only tolerates text. No meaning is added or removed; it is a reversible re-spelling. Base64 sits entirely on the byte layer. It reads the input three bytes (24 bits) at a time and re-cuts those 24 bits into four 6-bit groups, mapping each group through the alphabet `A-Z a-z 0-9 + /`. Nothing in that procedure has any notion of characters, so accepting a `str` would force the module to guess a text codec on your behalf, and the output would silently depend on that guess. CPython refuses instead: `base64.b64encode("héllo")` raises `TypeError: a bytes-like object is required, not 'str'`. ### The sandwich ```python token = base64.b64encode("héllo".encode("utf-8")).decode("ascii") ``` Read it inside out. `"héllo".encode("utf-8")` is the decision you own — it produces `b'h\xc3\xa9llo'`, six bytes for five characters. `b64encode` turns those six bytes into eight ASCII bytes, `b'aMOpbGxv'`. The trailing `.decode("ascii")` exists purely because most places you want to *put* the token — a JSON body, a header value, a log line, an f-string — want a `str`. Any codec that is an ASCII superset gives the same result, but naming `ascii` documents the invariant and fails loudly if a non-Base64 byte ever reaches it. Going the other way you unwrap in the mirror order: `base64.b64decode(token)` returns `b'h\xc3\xa9llo'`, and `.decode("utf-8")` turns that back into `"héllo"`. Note what is **not** stored anywhere: the fact that UTF-8 was used. Base64 round-trips bytes, not text. If the producer used UTF-16 and the consumer assumes UTF-8, the Base64 layer round-trips perfectly and the text still comes out wrong — which is why the codec belongs in your protocol or your schema, not in a comment. ### The asymmetry on decode `b64decode` accepts either bytes or a `str`. That is not sloppiness: the Base64 alphabet is a subset of ASCII, so a Base64 `str` has exactly one sensible byte reading and the module can do the conversion itself. It calls the string's ASCII encoding internally and raises `ValueError: string argument should contain only ASCII characters` if anything else shows up. Encoding has no such unique answer, so the API stays strict. Remembering the rule as "decode is lenient about the wrapper, encode is not" covers `binascii.unhexlify` and `base64.urlsafe_b64decode` too. ### The failure modes an interviewer is listening for The first is `str(b'aGk=')`, which produces the six-character string `"b'aGk='"` — the *repr* of the bytes object, quotes and prefix included. It looks almost right, it serialises without complaint, and it corrupts every consumer. Always use `.decode()`, never `str()`, on bytes. The second is skipping the outer decode and handing raw bytes to `json.dumps`, which raises `TypeError: Object of type bytes is not JSON serializable`. JSON has no binary type at all; that is precisely why Base64 is in the pipeline. The third is double encoding — Base64-encoding an already-Base64 token, usually because a helper does it and the caller does it again. It round-trips, so tests pass, and the payload grows by a third each time. ### Where the bytes actually come from In real code the input to `b64encode` is usually already bytes and the `.encode()` step is unnecessary — a file opened in binary mode, a `hashlib` digest, a `struct`-packed record, an image read from disk. The text sandwich shows up specifically when the *thing being carried* is text that must survive a transport which mangles it: a header value with newlines in it, a filename with non-ASCII characters, a JSON string that also has to hold a raw byte sequence. Recognising which case you are in stops the second common error — reaching for Base64 when the real problem was that a codec was never chosen, and a plain `str` would have travelled fine. It also decides where the failure surfaces. If the payload is genuinely binary, the only decision on the wire is the Base64 alphabet. If it is text, there are two independent decisions — the text codec and the Base64 alphabet — and only one of them is visible in the encoded string. Writing the codec into the schema, rather than assuming UTF-8 on both sides, is what keeps the second one from becoming a bug that only one client ever sees. ### Encoding is not encryption Say this before the interviewer has to ask. Base64 is a public, keyless, reversible transformation; anyone with the token has the bytes. It is for transport safety — getting arbitrary bytes through a text-only channel — and never for confidentiality. A credential that is "Base64 encoded" in a config file is a credential in plain text with an extra step.

  • base64.b64decode accepts a str while b64encode refuses one — why the asymmetry?
    Base64 output is ASCII by construction, so a Base64 `str` has exactly one byte reading and the module can do that conversion itself; it raises `ValueError` if the string holds non-ASCII characters. Encoding has no unique answer — the module would have to guess a text codec, and the output would silently depend on the guess. So the strict side is the one where a guess would change the result.
  • After base64.b64decode you get bytes back. What do you need to turn them into text?
    The codec the producer used. Base64 round-trips bytes and stores nothing about how those bytes represented characters, so `.decode('utf-8')` is an assumption you are making, not information you recovered. If the producer used UTF-16 the Base64 layer still round-trips perfectly and the text comes out wrong, which is why the codec belongs in the protocol or schema.
  • Why can't you put a bytes object into a JSON body directly?
    JSON has no binary type — its values are strings, numbers, booleans, null, arrays and objects — so `json.dumps` raises `TypeError: Object of type bytes is not JSON serializable`. Base64 exists in this pipeline exactly to bridge that gap: it re-spells arbitrary bytes in an ASCII alphabet that is safe inside a JSON string, at the cost of about a third more size.

Base64 is a way of spelling a phone number out loud so it survives a bad line — it changes nothing about the number and hides nothing from a listener.

saying these in an interview costs you the question

  • Claims Base64 encrypts or protects the data
  • Calls str() on the encoded bytes, producing a b'...' repr
  • Assumes b64encode takes a str like most text helpers
  • Thinks b64decode returns text rather than bytes
  • Confuses str.encode with base64.b64encode entirely
  • Base64-encodes an already-encoded token twice

context

open as a page

How do struct.pack and struct.unpack convert Python values to and from bytes?

level: juniorimportance: must knowfreq 40%

basics

~20 s

struct.pack takes a format string plus values and returns a bytes object laid out exactly as that format describes. struct.unpack reverses it and always returns a tuple, even when the format has a single field.

open as a page

How do Python's uuid.uuid4, uuid.uuid1 and uuid.uuid5 differ in where their bits come from?

level: juniorimportance: must knowfreq 62%

basics

~20 s

uuid.uuid4 fills 122 bits from the operating system's random source. uuid.uuid1 encodes a timestamp plus the host's 48-bit node address. uuid.uuid5 hashes a namespace UUID together with a name, so the same name always produces the same UUID.

open as a page

Why can zlib.compress exhaust memory on a multi-gigabyte file, and what replaces it?

level: middleimportance: must knowfreq 50%

basics

~10 s

zlib.compress takes one bytes object and returns another, so both the whole input and the whole output sit in memory at once. Stream instead: zlib.compressobj(), feed it chunks with compress(), and finish with flush().

open as a page

Why does struct.calcsize('@ci') exceed struct.calcsize('<ci')?

level: middleimportance: must knowfreq 34%

basics

~20 s

The '@' prefix means native byte order with native alignment, so struct inserts pad bytes before the int — typically eight bytes in total. The '<' prefix selects little-endian standard sizes with no alignment padding, giving exactly five.

open as a page

How does gzip.open in text mode differ from its default binary mode?

level: juniorimportance: should knowfreq 40%

basics

~20 s

gzip.open defaults to mode 'rb' and yields bytes. Adding a 't' — 'rt' or 'wt' — wraps the decompressed stream in a text layer that decodes to str, and only then may you pass encoding, errors or newline.

open as a page

How do you decode a base64.urlsafe_b64encode token whose '=' padding was stripped?

level: middleimportance: should knowfreq 45%

basics

~10 s

Re-add the padding before decoding: base64.urlsafe_b64decode(s + '=' * (-len(s) % 4)). The decoder needs a length that is a multiple of four and otherwise raises binascii.Error('Incorrect padding'); the '=' carries no data.

open as a page

When would you use binascii.hexlify instead of bytes.hex() in Python?

level: middleimportance: should knowfreq 35%

basics

~10 s

Only when you want bytes out: they compute the same hex digits, but binascii.hexlify returns bytes and bytes.hex() returns a str. Prefer bytes.hex(), which also takes a separator argument for readable dumps.

open as a page

Why use struct.unpack_from and struct.pack_into over struct.unpack and pack?

level: middleimportance: should knowfreq 22%

basics

~20 s

struct.unpack_from reads fields at an offset inside a larger buffer without slicing a copy out of it, and ignores trailing bytes. struct.pack_into writes fields in place into a pre-allocated writable buffer instead of allocating a new bytes object.

open as a page

How do you store a uuid.UUID as sixteen bytes and rebuild the object later?

level: middleimportance: should knowfreq 47%

basics

~20 s

Write u.bytes, which is the raw sixteen-byte value, and rebuild it with uuid.UUID(bytes=...). The hex form is 32 characters, int is a 128-bit integer, and str(u) is the 36-character hyphenated text — all four round-trip to the same object.

open as a page

A nightly report job Base64-encodes a 2.4 GB artifact into one JSON field. What goes wrong, and how would you fix it?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Base64 turns three bytes into four characters, so 2.4 GB becomes 3.2 GB of text, and json.dumps then builds further whole copies in memory. Stream the bytes out of band instead, or encode in chunks that are multiples of three bytes.

open as a page

Compressing sensor telemetry with gzip level 9 saturates the collector's CPU — how do you pick a codec and level?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Measure ratio and CPU on your own data. Level 9 typically buys under a percent over level 6 for several times the CPU, and codecs differ sharply: lzma smallest, zstd fast at a good ratio, gzip most portable.

open as a page

struct.unpack per record dominates a 45-second index cold start — how do you cut it?

level: seniorimportance: should knowfreq 20%

basics

~20 s

Compile the layout once as a module-level struct.Struct and stream the buffer through Struct.iter_unpack instead of slicing and calling struct.unpack per record. The C-level iterator removes the per-record slice, the lookup and most of the Python loop overhead.

open as a page

How does uuid.uuid5 with a namespace make a retried 6,800-row billing batch idempotent?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Derive each charge's identifier with uuid.uuid5(namespace, name) from stable inputs such as the run key and the subscription id. The retry recomputes the same UUID, so a unique constraint or an upsert rejects the second write instead of charging twice.

open as a page

What does the compression.zstd module added in Python 3.14 give you?

level: middleimportance: nice to knowfreq 15%

basics

~10 s

Python 3.14 adds Zstandard to the standard library as compression.zstd, with the familiar compress/decompress/open shape plus ZstdFile, incremental compressor and decompressor objects, trained dictionaries, and a tuning-parameter enum.

open as a page

Do uuid.uuid1 values sort in creation order, and what does uuid.uuid7 change?

level: middleimportance: nice to knowfreq 20%

basics

~20 s

No. A version 1 UUID puts the low 32 bits of its timestamp first, so text and byte order do not track time. Sort by UUID.time instead, or use uuid.uuid7(), added in Python 3.14, whose leading bytes are a millisecond timestamp.

open as a page