How do you decode a base64.urlsafe_b64encode token whose '=' padding was stripped?
answer
- Only two alphabet symbols change
- Padding survives the URL-safe variant
- The decoder wants a multiple of four
- Negate the length before the modulo
- The wrong decoder discards, it does not raise
basics
~10 sRe-add the padding before decoding: base64.urlsafe_b64decode(s + '=' * (-len(s) % 4)). The decoder needs a length that is a multiple of four and otherwise raises binascii.Error('Incorrect padding'); the '=' carries no data.
solid answer
~50 s`base64.urlsafe_b64encode` changes only the last two alphabet symbols — `+` and `/` become `-` and `_` — so the token survives a URL path or query. It still **appends `=` padding**, which is why tokens are usually stored stripped, and `base64.urlsafe_b64decode` still **requires** it: a length that is not a multiple of four raises `binascii.Error: Incorrect padding`. The fix is `s + "=" * (-len(s) % 4)`, which adds zero, one or two `=` and never three, since a length of one more than a multiple of four is not valid Base64 at all. The padding is pure redundancy — the length already tells you whether the final group encodes one, two or three bytes — which is what makes stripping it lossless. Never decode a URL-safe token with plain `base64.b64decode`: by default it silently discards `-` and `_` as non-alphabet characters and hands you wrong bytes.
code
python · 20 linesimport base64
import binascii
tok = base64.urlsafe_b64encode(b"any carnal pleas").rstrip(b"=")
print(tok) # b'YW55IGNhcm5hbCBwbGVhcw'
padded = tok + b"=" * (-len(tok) % 4)
print(base64.urlsafe_b64decode(padded)) # b'any carnal pleas'
try:
base64.urlsafe_b64decode(tok)
except binascii.Error as exc:
print("unpadded:", exc) # Incorrect padding
raw = b"\xfb\xef\xbe"
print(base64.b64encode(raw), base64.urlsafe_b64encode(raw)) # b'++++' b'----'
print("silently wrong:", base64.b64decode("----")) # b''
try:
base64.b64decode("----", validate=True)
except binascii.Error as exc:
print("validated:", exc) # Only base64 data is allowedgo deeper
Know that the URL-safe variant swaps two characters, + to - and / to _, and that '=' padding must be present when you decode. Recognise binascii.Error: Incorrect padding as a missing-padding message.
Derive the formula rather than memorising it: four characters per three bytes means the padded length is a multiple of four, so -len(s) % 4 is the count of '=' to restore. Explain why a remainder of one is impossible.
Lead with the silent failure: the default validate=False discards non-alphabet characters, so decoding a URL-safe token with the standard function returns wrong bytes rather than raising. Say when you turn validation on and where you enforce it.
Own the convention across services — one alphabet, one padding policy, validated at the boundary and written into the interface definition, so clients do not each rediscover that two nearly identical decoders disagree on two characters.
### What "URL-safe" actually changes Standard Base64 uses `A-Z a-z 0-9 + /`. Two of those symbols are hostile in a URL: `/` is the path separator, and `+` means a literal space when a query string is form-decoded. `base64.urlsafe_b64encode` substitutes `-` for `+` and `_` for `/` and changes nothing else — it is exactly `base64.b64encode(data, altchars=b"-_")`, and `urlsafe_b64decode` translates the pair back before decoding. What it does **not** do is drop the padding. `base64.urlsafe_b64encode(b"any carnal pleas")` returns `b'YW55IGNhcm5hbCBwbGVhcw=='`, trailing `=` and all, and `=` is itself awkward in a URL — it is the separator inside a query string and gets percent-encoded as `%3D` in many contexts. That is why the token formats you meet in the wild strip it, and why the interview question is about putting it back. ### The padding rule, derived Base64 consumes three input bytes at a time and emits four characters. When the input length is not a multiple of three there is a leftover group of one or two bytes: one byte becomes two characters plus `==`, two bytes become three characters plus `=`. So the encoded length is always `4 * ceil(n / 3)`, and `len % 4` of a stripped token can only be 0, 2 or 3 — never 1, because one leftover Base64 character carries six bits and cannot encode a whole byte. That gives the restoration formula directly: ```python padded = s + "=" * (-len(s) % 4) ``` Python's `%` returns a non-negative result for a positive modulus, so `-len(s) % 4` is exactly the number of characters missing to reach the next multiple of four: 0, 1 or 2 (and 3 only for the impossible `len % 4 == 1` case, which will fail the decode anyway, as it should). Writing `len(s) % 4` instead is the common off-by-one — it computes how many characters are *past* the boundary, not how many are missing. The deeper point is that the padding is redundant. The decoder can derive the leftover-byte count from the length alone; `=` exists so that concatenated Base64 streams can be split at message boundaries by a decoder that reads forward without knowing lengths in advance. In a JSON field or a URL segment you already know where the token ends, so stripping is lossless — and if your protocol strips, it must document that, because `binascii` will not guess. ### The silent-corruption trap The worst failure here is not the exception, it is the absence of one. `base64.b64decode` defaults to `validate=False`, which **discards every character outside the standard alphabet before decoding**. Feed it a URL-safe token and the `-` and `_` characters simply vanish: ```pycon >>> base64.b64decode("----") b'' >>> base64.b64decode("----", validate=True) binascii.Error: Only base64 data is allowed ``` Four valid URL-safe characters decode to nothing at all, with no error. In a real token the surviving characters usually still form a length that is not a multiple of four, so you often get `Incorrect padding` instead — but not always, and "sometimes an exception, sometimes wrong bytes" is the worst possible failure mode. Passing `validate=True` turns the whole class into a loud `binascii.Error`, and it costs one keyword argument. Use the matching function for the alphabet, and validate when the input is untrusted. ### Which exception, and why it matters All of these raise `binascii.Error`, which is a subclass of `ValueError` — so `except ValueError` catches padding and alphabet problems together if you would rather not import `binascii` to write the handler. The messages you will actually see are `Incorrect padding` (length not a multiple of four), `Only base64 data is allowed` (a non-alphabet character under `validate=True`), and `Invalid base64-encoded string` (a final group of exactly one character). Since Python 3.11 the lower-level `binascii.a2b_base64` also takes `strict_mode=True`, which additionally rejects data appearing after the padding — useful when you are parsing a concatenated stream rather than a single field. ### The rule to carry away Pick one convention per protocol and enforce it at the edge: URL-safe alphabet, padding stripped, restored with `-len(s) % 4` on the way in, decoded with `urlsafe_b64decode` and `validate=True`. Two functions that differ in two characters, one of which fails silently against the other's output, is exactly the kind of pairing that produces a bug report months later from one client on one code path.
- What does the validate argument to base64.b64decode change?By default (`validate=False`) any character outside the standard Base64 alphabet is discarded before decoding, so newline-wrapped text decodes fine — and so does a URL-safe token, into the wrong bytes. With `validate=True` any such character raises `binascii.Error: Only base64 data is allowed`. Turn it on for untrusted or cross-system input, where a silently wrong result is far more expensive than an exception.
- If the '=' padding carries no data, why does the format have it at all?It marks the end of a message in a stream a decoder reads forward without knowing the length in advance: `==` means the final group held one byte, `=` means two. When the token sits in a JSON field or a URL segment you already know where it ends, so the count is derivable from the length and stripping is lossless — which is why so many token formats strip it.
- Can a stripped token's length ever be one more than a multiple of four?No. Every three input bytes produce four characters, and a leftover of one or two bytes produces two or three characters respectively — never one, because a single Base64 character carries only six bits. So `len % 4` of a valid stripped token is 0, 2 or 3; a remainder of 1 means the token was truncated or corrupted, and the decoder rejects it with `binascii.Error`.
saying these in an interview costs you the question
- Thinks urlsafe_b64encode omits the '=' padding
- Strips padding on write and never restores it on read
- Pads with len(s) % 4 instead of -len(s) % 4
- Decodes a URL-safe token with plain b64decode
- Assumes a wrong alphabet always raises an exception
- Pads the encoded text to a multiple of three