Why does base64.b64encode reject a str, and what do you write instead?
answer
- Two layers: characters and bytes
- The function works below the text layer
- A conversion brackets the call on each side
- str.encode in, .decode('ascii') out
- Decoding is the lenient direction
basics
~10 sbase64.b64encode is a bytes-to-bytes function: it needs a bytes-like object and returns bytes. Encode text first with str.encode('utf-8'), then call .decode('ascii') on the result to get a str for JSON.
solid answer
~40 sBase64 is a transformation over **bytes**, so `base64.b64encode` takes a bytes-like object and returns bytes; handing it a `str` raises `TypeError: a bytes-like object is required, not 'str'`. The idiom is a sandwich: `base64.b64encode(text.encode('utf-8')).decode('ascii')` — the inner `str.encode` chooses how the text becomes bytes, and the outer `bytes.decode` turns the Base64 output, which is ASCII by construction, into a `str` you can put in JSON. Decoding is deliberately asymmetric: `base64.b64decode` **does** accept a `str`, because Base64 text is ASCII by definition, and it raises `ValueError` if the string contains non-ASCII characters. What comes back from `b64decode` is bytes, and turning those back into text requires knowing the original codec — Base64 does not carry one.
code
python · 14 linesimport base64
import json
token = base64.b64encode("héllo".encode("utf-8")).decode("ascii")
print(token) # aMOpbGxv
print(base64.b64decode(token).decode("utf-8")) # héllo
print(json.dumps({"payload": token})) # {"payload": "aMOpbGxv"}
for bad in (lambda: base64.b64encode("héllo"),
lambda: json.dumps({"payload": base64.b64encode(b"hi")})):
try:
bad()
except TypeError as exc:
print("TypeError:", exc)go deeper
Memorise the sandwich: base64.b64encode(text.encode('utf-8')).decode('ascii'). Be ready to say that the function works on bytes and returns bytes, and that Base64 hides nothing from anyone.
Explain the mechanics: three input bytes become four ASCII characters from a 64-symbol alphabet, and the decode side accepts a str because Base64 text is ASCII by definition. Name the TypeError and the ValueError you would actually see.
Show where this bites in production — double encoding through a helper, a str() repr leaking into a payload, and a codec mismatch that Base64 round-trips happily while the text still arrives wrong. Put the text codec in the schema.
Own the contract question: if binary is crossing a service boundary, decide whether it belongs inline at all, and if it does, pin the alphabet, the padding convention and the text codec in the interface definition rather than letting each client rediscover them.
### Two different conversions wearing the same word Python uses the verb "encode" for two unrelated jobs, and this question exists to check that you keep them apart. * **Text encoding** — `str.encode(codec)` turns characters into bytes according to a codec such as UTF-8, and `bytes.decode(codec)` reverses it. This is where meaning lives: the same characters produce different bytes under UTF-8, UTF-16 or Latin-1. * **Binary-to-text encoding** — `base64.b64encode` takes arbitrary bytes and re-expresses them in a 64-character ASCII alphabet so they survive a transport that only tolerates text. No meaning is added or removed; it is a reversible re-spelling. Base64 sits entirely on the byte layer. It reads the input three bytes (24 bits) at a time and re-cuts those 24 bits into four 6-bit groups, mapping each group through the alphabet `A-Z a-z 0-9 + /`. Nothing in that procedure has any notion of characters, so accepting a `str` would force the module to guess a text codec on your behalf, and the output would silently depend on that guess. CPython refuses instead: `base64.b64encode("héllo")` raises `TypeError: a bytes-like object is required, not 'str'`. ### The sandwich ```python token = base64.b64encode("héllo".encode("utf-8")).decode("ascii") ``` Read it inside out. `"héllo".encode("utf-8")` is the decision you own — it produces `b'h\xc3\xa9llo'`, six bytes for five characters. `b64encode` turns those six bytes into eight ASCII bytes, `b'aMOpbGxv'`. The trailing `.decode("ascii")` exists purely because most places you want to *put* the token — a JSON body, a header value, a log line, an f-string — want a `str`. Any codec that is an ASCII superset gives the same result, but naming `ascii` documents the invariant and fails loudly if a non-Base64 byte ever reaches it. Going the other way you unwrap in the mirror order: `base64.b64decode(token)` returns `b'h\xc3\xa9llo'`, and `.decode("utf-8")` turns that back into `"héllo"`. Note what is **not** stored anywhere: the fact that UTF-8 was used. Base64 round-trips bytes, not text. If the producer used UTF-16 and the consumer assumes UTF-8, the Base64 layer round-trips perfectly and the text still comes out wrong — which is why the codec belongs in your protocol or your schema, not in a comment. ### The asymmetry on decode `b64decode` accepts either bytes or a `str`. That is not sloppiness: the Base64 alphabet is a subset of ASCII, so a Base64 `str` has exactly one sensible byte reading and the module can do the conversion itself. It calls the string's ASCII encoding internally and raises `ValueError: string argument should contain only ASCII characters` if anything else shows up. Encoding has no such unique answer, so the API stays strict. Remembering the rule as "decode is lenient about the wrapper, encode is not" covers `binascii.unhexlify` and `base64.urlsafe_b64decode` too. ### The failure modes an interviewer is listening for The first is `str(b'aGk=')`, which produces the six-character string `"b'aGk='"` — the *repr* of the bytes object, quotes and prefix included. It looks almost right, it serialises without complaint, and it corrupts every consumer. Always use `.decode()`, never `str()`, on bytes. The second is skipping the outer decode and handing raw bytes to `json.dumps`, which raises `TypeError: Object of type bytes is not JSON serializable`. JSON has no binary type at all; that is precisely why Base64 is in the pipeline. The third is double encoding — Base64-encoding an already-Base64 token, usually because a helper does it and the caller does it again. It round-trips, so tests pass, and the payload grows by a third each time. ### Where the bytes actually come from In real code the input to `b64encode` is usually already bytes and the `.encode()` step is unnecessary — a file opened in binary mode, a `hashlib` digest, a `struct`-packed record, an image read from disk. The text sandwich shows up specifically when the *thing being carried* is text that must survive a transport which mangles it: a header value with newlines in it, a filename with non-ASCII characters, a JSON string that also has to hold a raw byte sequence. Recognising which case you are in stops the second common error — reaching for Base64 when the real problem was that a codec was never chosen, and a plain `str` would have travelled fine. It also decides where the failure surfaces. If the payload is genuinely binary, the only decision on the wire is the Base64 alphabet. If it is text, there are two independent decisions — the text codec and the Base64 alphabet — and only one of them is visible in the encoded string. Writing the codec into the schema, rather than assuming UTF-8 on both sides, is what keeps the second one from becoming a bug that only one client ever sees. ### Encoding is not encryption Say this before the interviewer has to ask. Base64 is a public, keyless, reversible transformation; anyone with the token has the bytes. It is for transport safety — getting arbitrary bytes through a text-only channel — and never for confidentiality. A credential that is "Base64 encoded" in a config file is a credential in plain text with an extra step.
- base64.b64decode accepts a str while b64encode refuses one — why the asymmetry?Base64 output is ASCII by construction, so a Base64 `str` has exactly one byte reading and the module can do that conversion itself; it raises `ValueError` if the string holds non-ASCII characters. Encoding has no unique answer — the module would have to guess a text codec, and the output would silently depend on the guess. So the strict side is the one where a guess would change the result.
- After base64.b64decode you get bytes back. What do you need to turn them into text?The codec the producer used. Base64 round-trips bytes and stores nothing about how those bytes represented characters, so `.decode('utf-8')` is an assumption you are making, not information you recovered. If the producer used UTF-16 the Base64 layer still round-trips perfectly and the text comes out wrong, which is why the codec belongs in the protocol or schema.
- Why can't you put a bytes object into a JSON body directly?JSON has no binary type — its values are strings, numbers, booleans, null, arrays and objects — so `json.dumps` raises `TypeError: Object of type bytes is not JSON serializable`. Base64 exists in this pipeline exactly to bridge that gap: it re-spells arbitrary bytes in an ASCII alphabet that is safe inside a JSON string, at the cost of about a third more size.
Base64 is a way of spelling a phone number out loud so it survives a bad line — it changes nothing about the number and hides nothing from a listener.
saying these in an interview costs you the question
- Claims Base64 encrypts or protects the data
- Calls str() on the encoded bytes, producing a b'...' repr
- Assumes b64encode takes a str like most text helpers
- Thinks b64decode returns text rather than bytes
- Confuses str.encode with base64.b64encode entirely
- Base64-encodes an already-encoded token twice