skip to content

Why does embedding raw bytes inside a text-encoded message inflate the payload by about a third?

level: middleimportance: nice to knowfreq 30%

answer

  1. characters only, not arbitrary bytes
  2. six bits per character
  3. three bytes become four characters
  4. output is ceil(n/3)*4
  5. about 33 percent, plus padding

basics

~20 s

A text message can only carry characters, so raw bytes must be re-encoded into a safe alphabet. Base64 carries six bits per character, turning every three input bytes into four characters: roughly a third larger, before padding.

solid answer

~50 s

A text grammar admits characters, not arbitrary byte values, so a blob cannot travel through it untouched. The standard fix is **Base64**, which takes the input twenty-four bits at a time and emits four characters drawn from a sixty-four symbol alphabet — six bits each. Since each of those characters occupies one byte in an ASCII-compatible text stream, three bytes in become four bytes out: a factor of `4/3`, about **33 percent**. The output is padded up to a multiple of four characters, so a small blob pays slightly more, and some pipelines add line breaks on top. The receiver also decodes twice: once out of the text grammar, then once out of the alphabet. A compact binary encoding avoids the whole thing because it has a native bytes type — length followed by the bytes, unchanged.

code

pseudocode · 11 lines
pseudocode
output = empty
for each group of 3 bytes in input:          // 24 bits
    emit 4 characters from the 64-symbol alphabet   // 6 bits each

if 1 byte remains:
    emit 2 characters, then 2 padding characters
else if 2 bytes remain:
    emit 3 characters, then 1 padding character

// output length = ceil(n / 3) * 4 characters
// 1048576 bytes -> 349526 groups -> 1398104 characters

go deeper

for a junior

Know that a text message carries characters, so raw bytes must be re-encoded first, and that this makes the payload noticeably larger rather than leaving it unchanged.

for a middle

Be able to derive the ratio: twenty-four bits become four six-bit symbols, each occupying a byte, so the output is ceil(n/3)*4 characters, about a third larger before padding.

for a senior

Judge whether the blob belongs in the message at all — reference plus separate channel, a native bytes type on that hop, or acceptance of the overhead for a small field on a cold path.

for a principal

Set the boundary rule for the platform: what maximum blob size may ride inside a message, and what the alternative path is, so that individual teams are not each rediscovering this arithmetic.

## Why a text channel cannot carry arbitrary bytes A text encoding is defined over characters. Byte values that do not form valid characters, or that collide with the grammar's own structural characters, cannot simply be dropped into a field: the reader would either reject them or mis-parse the surrounding structure. So a blob — an image, a compressed archive, a signature, a key — has to be **re-expressed** using only characters the grammar accepts. ## The arithmetic Base64 is the usual answer, and its cost is arithmetic rather than opinion: 1. The input is taken **three bytes at a time**, which is twenty-four bits. 2. Those twenty-four bits are split into **four groups of six bits**. 3. Each six-bit group indexes a **sixty-four symbol alphabet**, producing one character. 4. Each such character occupies **one byte** in an ASCII-compatible text stream. So three bytes become four bytes: `4/3 = 1.3333`, an expansion of about **33 percent**. The output length is `ceil(n / 3) * 4` characters. A leftover of one input byte emits two characters plus two padding characters; a leftover of two emits three characters plus one. A worked example makes it concrete. Take a blob of one mebibyte, 1,048,576 bytes: - `1048576 / 3 = 349525.33`, so `ceil` gives **349,526** groups. - `349526 * 4 = 1,398,104` characters, hence 1,398,104 bytes of text. - The overhead is **349,528 bytes**, a ratio of 1.3333. ## What else rides on top - **Padding** rounds the output up to a multiple of four characters — negligible on a large blob, proportionally noticeable on a tiny one. - **Line breaks**, added by some pipelines for historical reasons, cost a further couple of percent. - **A second decode pass** on the receiver: the characters come out of the text grammar first, then out of the alphabet. The blob is walked twice before anything can use it. - **Memory during decode**, because the encoded and decoded forms of a large blob tend to exist at the same time. ## What compression does and does not recover This is the part that is usually stated too strongly in both directions. Base64 output uses only sixty-four distinct symbols, so each byte of it carries six bits of information rather than eight. A compressor with an entropy-coding stage notices that restricted alphabet and **largely removes the expansion**, bringing the encoded blob back towards the size of the original bytes. What it cannot do is shrink the blob itself when that blob was already compressed, already encrypted, or otherwise high in entropy — there is no redundancy left to find. So the honest summary is: compression usually undoes the alphabet overhead and nothing beyond it, at the price of compressing and decompressing a large field on both ends. ## The alternatives - **Send the bytes on a separate channel** and carry only a reference in the message. The text message stays small and readable; the blob travels as bytes. - **Use an encoding with a native bytes type** for that hop, so the blob is written as a length followed by its bytes with no expansion at all. - **Split the message**, keeping readable fields readable and moving only the blob out. This preserves the inspectability the text form was chosen for, which a message dominated by a wall of alphabet characters has already lost. That last point is the one worth saying aloud: a large embedded blob defeats the reason the text encoding was chosen. Nobody reads several hundred thousand characters of alphabet in a log line, so the payload is paying the readability tax without collecting the benefit. ## How this is asked Interviewers use it to check that a candidate can reason about an encoding numerically rather than by feel. Give the mapping (three bytes to four characters), give the ratio (`4/3`, about a third), mention padding, and then make the design point: if the blob is large, the right move is usually to stop carrying it inside a text message at all.

  • If the whole message is compressed afterwards, does the expansion disappear?
    Mostly, but not usefully. Base64 output draws on sixty-four symbols, so an entropy stage recognises the restricted alphabet and recovers close to the original blob size. It cannot go further if the blob was already compressed or encrypted, since no redundancy remains. You have then spent compression work on both ends to undo an expansion you could have avoided by not embedding the blob.
  • When is embedding a blob in a text message still the right call?
    When the blob is small and a single round trip matters more than bytes — a thumbnail, a short signature, a certificate. The expansion on a few kilobytes is irrelevant, and keeping one self-contained message is simpler than coordinating a second fetch. It stops being the right call once the blob dominates the message, because the payload then pays the readability cost without anyone being able to read it.

saying these in an interview costs you the question

  • Thinks embedding bytes as characters is free because the bytes are unchanged
  • States the expansion as fifty percent or as a doubling of the payload
  • Assumes a text grammar can carry arbitrary byte values untouched
  • Believes compression can shrink an already-compressed blob further
  • Never questions whether a large blob belongs inside the message at all