A nightly report job Base64-encodes a 2.4 GB artifact into one JSON field. What goes wrong, and how would you fix it?
answer
- Start by sizing the field honestly
- Four characters for every three bytes
- The wire cost is not the peak cost
- Chunk boundaries must respect the grouping
- The truncation arrives without an exception
basics
~20 sBase64 turns three bytes into four characters, so 2.4 GB becomes 3.2 GB of text, and json.dumps then builds further whole copies in memory. Stream the bytes out of band instead, or encode in chunks that are multiples of three bytes.
solid answer
~50 sThe encoded length is `4 * ceil(n / 3)`, so a 2.4 GB artifact is exactly 3.2 GB of ASCII — a third more, not a rounding error. Worse is the memory profile: `base64.b64encode` materialises that 3.2 GB `bytes` object while the input is still live, `json.dumps` builds another whole `str`, and encoding the body to UTF-8 builds another, so peak RSS is several times the artifact and the receiver must buffer the entire field before it can decode anything. The right fix is not to embed it: write the artifact to object storage or a separate transfer and put a reference in the JSON. If it must be inline, encode incrementally — and split the input on **multiples of three bytes**. A 4096-byte chunk is one byte past a multiple of three, so each chunk gets its own `=` padding and the concatenation silently decodes to only the first chunk.
code
python · 19 linesimport base64
import binascii
import math
n = 2_400_000_000
print(4 * math.ceil(n / 3)) # 3200000000 characters
data = bytes(range(256)) * 40 # 10240 bytes
def joined(chunk):
return b"".join(base64.b64encode(data[i:i + chunk])
for i in range(0, len(data), chunk))
print(len(base64.b64decode(joined(4096)))) # 4096 - silently truncated
print(len(base64.b64decode(joined(4095)))) # 10240 - 4095 is a multiple of 3
try:
base64.b64decode(joined(4096), validate=True)
except binascii.Error as exc:
print("validated:", exc) # Excess data after paddinggo deeper
Know the size rule: three bytes become four characters, so Base64 grows a payload by about a third, and JSON cannot hold raw bytes at all. That is enough to question a gigabyte-scale blob in a JSON field.
Compute 4 * ceil(n / 3) and explain why encoding, serialising and UTF-8 encoding each allocate another full copy, so peak memory is a multiple of the payload rather than 1.33 times it.
Diagnose the chunk-boundary bug from a truncated artifact and no exception, name the multiple-of-three rule and validate=True, and argue for moving the blob out of the JSON body entirely.
Own the boundary contract: decide when payloads travel inline versus by reference, set the threshold, and require a digest and length beside every reference so integrity does not depend on a transport that truncates quietly.
### The arithmetic first Base64 reads three input bytes and emits four characters, padding the final group. The encoded size is therefore `4 * ceil(n / 3)` — a growth factor of 4/3, about 33%, plus at most two padding characters. For a 2.4 GB artifact that is 3,200,000,000 characters. Since the Base64 alphabet contains no JSON metacharacters, `json.dumps` adds only the two quotes and escapes nothing, so the field really is 3.2 GB on the wire. A candidate who says "a bit of overhead" has not sized it. A candidate who says "it doubles" is describing hex. The one-third figure is the number to have. ### The memory profile is worse than the wire size Wire size is the part people estimate; peak resident memory is the part that pages the job out. `base64.b64encode(data)` is not a generator. It allocates one contiguous 3.2 GB `bytes` object while `data` is still referenced by the caller, so you are already at 5.6 GB. Then `.decode('ascii')` or `json.dumps` builds a `str` — another 3.2 GB, since a pure-ASCII `str` is one byte per character in CPython's compact representation. Then serialising the response body back to UTF-8 makes another. A nightly job that reads its artifact into memory and hands it to `json.dumps` can touch four times the artifact size at peak, and the allocations are large contiguous blocks, which is the shape most likely to fail even when the total looks affordable. The receiver has the mirror problem: JSON has no framing inside a string value, so a parser cannot hand you the field until it has read the closing quote. The whole 3.2 GB is buffered before a single byte is decoded. ### The fix that actually belongs in the design Do not put a 2.4 GB blob in a JSON field. Write the artifact to a blob store or a file endpoint and put a URL, a size and a digest in the JSON. That removes the encoding overhead entirely (binary goes as binary), makes the transfer resumable and cacheable, and lets the report metadata stay a small document that anything can parse. Base64 in JSON is a good answer for kilobytes — a thumbnail, a signature, an encrypted token — and a bad one for gigabytes. ### If it must be inline: the three-byte boundary Sometimes the pipeline is fixed and you must stream the encoding. The rule is that **every chunk you encode independently must be a multiple of three bytes**, except the last. The reason is padding. Three bytes map to four characters with no `=`; any other length produces padding, and padding means end-of-message. Encode in 4096-byte chunks — 4096 is one more than 4095, which is a multiple of three — and every chunk ends in `=`, so the concatenated text is not one Base64 message but many jammed together. Here is what CPython 3.14 does with that, and it is the reason this question is a senior one: ```python data = bytes(range(256)) * 40 # 10240 bytes blob = b"".join(base64.b64encode(data[i:i + 4096]) for i in range(0, len(data), 4096)) len(base64.b64decode(blob)) # 4096, not 10240 — no exception ``` The decoder stops at the first padding run and returns only the first chunk. No error, no warning, a truncated artifact. Change the chunk to 4095 — or any multiple of three — and the round trip is exact. Passing `validate=True` converts the silent truncation into `binascii.Error: Excess data after padding`, and since Python 3.11 `binascii.a2b_base64(..., strict_mode=True)` rejects trailing data at the lower level too. Both are worth having on any decode path that consumes something a producer assembled in pieces. The practical shape is a chunk size that is a comfortable multiple of three — `3 * 1024 * 1024` is a common choice — written straight to the output stream, so neither the encoded text nor the input is ever fully resident. `base64.encodebytes` exists as a convenience but wraps its output in newlines every 76 characters, which is right for MIME and wrong for a JSON string, where each newline must then be escaped. ### What to say in the interview Lead with the number (4/3, so 3.2 GB), then the memory multiplication that the wire size hides, then the design fix (a reference, not an embed), and keep the chunking rule as the answer to "suppose you cannot change the format". Mentioning that the chunk-boundary bug fails silently rather than loudly is the part that shows you have debugged it rather than read about it.
- How many characters exactly does Base64 produce for n bytes, and does JSON add escaping on top?`4 * ceil(n / 3)` characters, padding included. JSON adds only the surrounding quotes: the standard alphabet is letters, digits, `+` and `/`, and none of those require escaping inside a JSON string — CPython's `json.dumps` does not escape `/`. So the field size is the Base64 size plus two, which makes the 4/3 factor the whole story on the wire, though not in memory.
- Why must a streamed Base64 encoder split its input on multiples of three bytes?Because three bytes map to exactly four characters with no padding, while any other length emits `=`, which means end-of-message. Concatenating independently encoded chunks of, say, 4096 bytes therefore produces many padded messages glued together, and the decoder returns only the first one — silently, unless you pass `validate=True`. Any multiple of three, such as 3 MiB, keeps the stream a single valid message.
- Would using hex instead of Base64 avoid the boundary problem?Yes — hex is stateless, one byte to two characters with no padding or alignment, so it can be split and concatenated anywhere. But it costs 100% overhead instead of 33%, turning 2.4 GB into 4.8 GB, so it trades a bug you can prevent with a chunk-size constant for size you pay on every transfer. Fix the chunking; do not switch encodings to dodge it.
- What would you put in the JSON instead of the artifact?A reference: a URL or object key, the byte length, a content type and a digest for integrity. The binary then moves as binary over a transfer that supports ranges, resumption and caching, the metadata document stays small enough for any client to parse, and neither side needs gigabytes of contiguous memory. Inline Base64 is right for kilobyte-scale payloads, not gigabyte-scale ones.
saying these in an interview costs you the question
- Calls Base64 overhead negligible without sizing it
- Estimates the growth as double, or as a saving
- Concatenates independently encoded chunks of any size
- Assumes json.dumps streams rather than building one string
- Expects corruption to surface as a decode exception
- Ignores that encode, dumps and UTF-8 each copy the payload