skip to content

Why does a JWT use Base64url encoding rather than standard Base64, and why is padding dropped?

level: middleimportance: should knowfreq 45%

answer

  1. Two characters change, one disappears
  2. Plus and slash are hostile in URLs
  3. Equals is a query-string delimiter
  4. Padding is recomputable, so it is removed
  5. Result contains no dot

basics

~20 s

Standard Base64 uses '+', '/' and '=', which need escaping in URLs, query strings and some headers. Base64url substitutes '-' and '_' and JOSE omits the '=' padding, so a token is safe to place anywhere unescaped.

solid answer

~40 s

Base64url is the URL- and filename-safe Base64 alphabet: it keeps `A–Z`, `a–z` and `0–9` but replaces `+` with `-` and `/` with `_`. Those two substitutions matter because `+` is decoded as a space in form-encoded data and `/` is a path separator, so a standard Base64 token would have to be percent-encoded to survive a URL or a redirect. JOSE additionally requires the trailing `=` padding to be **omitted**, since `=` is a reserved delimiter in query strings and cookies and the padding carries no information — the decoder can recompute it from the segment length. The combined effect is that a JWT is a string of `A–Za–z0–9-_` plus dots, which can be dropped into a URL, an `Authorization` header, a cookie or a filename verbatim, and split on the dot without ambiguity.

code

text · 4 lines
text
standard Base64 : A-Z a-z 0-9 + / with '=' padding
Base64url (JOSE): A-Z a-z 0-9 - _ with no padding

conversion for a standard decoder:  '-' -> '+' , '_' -> '/' , re-add '=' to a multiple of 4

go deeper

for a junior

Recall that a JWT is Base64url, a URL-safe variant, and that it is not encryption. Knowing the two substituted characters is a nice extra at this level.

for a middle

Explain the '+'→'-' and '/'→'_' substitutions, why JOSE removes the '=' padding, and why the dot can therefore serve as an unambiguous separator.

for a senior

Be able to diagnose cross-language decode failures caused by padded-versus-unpadded input, and to reject malformed segment lengths before handing bytes to a parser.

for a principal

Own the interoperability angle: mandate URL-safe unpadded handling in shared libraries so tokens are byte-identical everywhere and never get re-encoded on the way through your platform.

## The problem Base64url solves Base64 turns arbitrary bytes into text using 64 characters plus a padding character. The classic alphabet ends with `+` and `/`, and pads with `=`. All three are hostile in the places tokens travel: - `+` means a literal space in `application/x-www-form-urlencoded` data, so a token pasted into a query string silently corrupts. - `/` is a path separator; a token in a URL path segment or a filename would be split. - `=` is the name/value delimiter in query strings and cookie strings, and is reserved in URI syntax. Every one of those would force percent-encoding, meaning the same token has two textual forms depending on where it appears — an invitation for bugs at exactly the layer where security depends on comparing bytes. ## The substitution Base64url is the URL- and filename-safe alphabet defined alongside standard Base64: identical for the first 62 characters, with index 62 encoded as `-` instead of `+` and index 63 as `_` instead of `/`. Nothing else changes; the bit-packing is the same, so converting between the two forms is a two-character search-and-replace, which is exactly what shell one-liners do before calling a standard Base64 decoder. ## Why padding is dropped Base64 encodes three bytes into four characters. When the input length is not a multiple of three, the final group is short and standard Base64 appends one or two `=` characters so the output length is always a multiple of four. That padding is pure redundancy: the decoder can tell how many bytes the final group carries from how many characters are left. JOSE therefore **requires** that the padding be removed, and encoders that leave it on produce tokens some verifiers reject. The consequences are practical. A segment length of `4n+1` is invalid and signals a malformed token. And many standard library decoders — written for padded input — reject an unpadded segment, which is why JWT libraries ship a dedicated Base64url decoder and why hand-rolled decoding is a common source of "works in one language, fails in another" bugs. ## Why the dot works as a separator Once the alphabet is `A–Za–z0–9-_` with no padding, no segment can contain a `.`. That makes the dot an unambiguous delimiter: a parser can split first and decode later, and a malformed token is detected by counting segments rather than by decoding. It also means a JWT survives being logged, copied out of a terminal, put in a `Location` header of a redirect, or used as a cache key without any escaping step. ## Consequences for size Base64 of any flavour costs about 33% expansion: three bytes become four characters. Dropping padding saves at most two characters per segment — negligible, and not the reason it is dropped. If a token is too large for a header or a cookie, the fix is fewer or smaller claims, never a different encoding. It is worth being explicit about that in an interview, because "drop padding to save space" is a tempting but wrong rationale. ## What interviewers are checking This is a details question, and it separates people who have debugged tokens from people who have only consumed a library. The strong answer names the two substituted characters, states that JOSE mandates unpadded output, and connects both to where tokens actually travel. A weak answer says "it is a variant of Base64" and stops.

  • What does it mean if a JWT segment's length modulo 4 equals 1?
    The token is malformed. Base64 emits 2, 3 or 4 characters for a final group, never 1, so a segment of length 4n+1 cannot be valid unpadded Base64url. Rejecting it early is cheaper and safer than letting a decoder produce partial bytes.
  • Why do some standard library decoders fail on a JWT segment?
    Because they expect the padded, standard alphabet: they choke on '-' and '_' or on a length that is not a multiple of four. Use the library's URL-safe, no-padding decoder, or convert the alphabet and re-add padding before calling a classic decoder.
  • Does dropping the padding meaningfully shorten a token?
    No — at most two characters per segment. Base64 of any flavour still costs about a third in expansion. Padding is dropped because '=' is a reserved delimiter in URLs and cookies and carries no information, not to save bytes; token size is controlled by trimming claims.

saying these in an interview costs you the question

  • Says Base64url and Base64 use the same alphabet
  • Claims padding is dropped to save space
  • Thinks Base64url is a different, stronger encoding
  • Cannot name the two substituted characters
  • Assumes any Base64 decoder handles a JWT segment

context