When would you pick base64.RawURLEncoding over base64.StdEncoding in Go?
answer
- only two of the 64 characters differ
- '/' breaks a path, '+' becomes a space
- '=' still needs percent-escaping
- Raw means padding omitted
- the decoder must match the encoder
basics
~20 sPick base64.RawURLEncoding when the value goes into a URL path, query or filename: it swaps '+' and '/' for '-' and '_' and drops the '=' padding, so nothing needs percent-escaping. Use base64.StdEncoding for ordinary opaque payloads.
solid answer
~40 s`encoding/base64` ships four ready-made encodings. `StdEncoding` is RFC 4648's standard alphabet with `=` padding; `URLEncoding` is the same but with `-` and `_` in place of `+` and `/`; the `Raw` twins, `RawStdEncoding` and `RawURLEncoding`, are those two with padding omitted. For a link shortener handing out opaque identifiers I would take `RawURLEncoding`: `/` would break a path segment, `+` is decoded as a space by form decoding, and `=` needs percent-escaping — none of which happen with `-`, `_` and no padding. Dropping the padding loses nothing, because the decoder recovers the byte count from the character count. The one hard rule is that decoding must use the same variant: `base64.StdEncoding.DecodeString` on a URL-alphabet value fails with a corrupt-input error, and a `Raw` decoder rejects a trailing `=`.
code
go · 5 linesid := []byte{0xde, 0xad, 0xbe, 0xef}
fmt.Println(base64.StdEncoding.EncodeToString(id)) // 3q2+7w==
fmt.Println(base64.URLEncoding.EncodeToString(id)) // 3q2-7w==
fmt.Println(base64.RawURLEncoding.EncodeToString(id)) // 3q2-7wgo deeper
Remember there are two alphabets and a padded and unpadded form of each, and that anything destined for a URL should use the URL-safe, unpadded one.
Explain exactly which characters differ, what padding is for, and why a decoder built on one variant rejects a value produced by another.
Demonstrate the triage: given a rejected identifier, read the alphabet and padding off the string itself, then fix the contract so both sides name the variant rather than guessing.
Own the choice as a published format: identifiers are pasted into URLs, logs and spreadsheets by people you will never meet, so the variant and its length are effectively frozen once external consumers exist.
## Four encodings, two independent choices RFC 4648 defines two base64 alphabets, and Go exposes each of them twice — padded and unpadded — as four package-level `*base64.Encoding` values: | value | characters 62 and 63 | padding | |---|---|---| | `base64.StdEncoding` | `+` and `/` | `=` | | `base64.URLEncoding` | `-` and `_` | `=` | | `base64.RawStdEncoding` | `+` and `/` | none | | `base64.RawURLEncoding` | `-` and `_` | none | The first 62 characters (`A–Z`, `a–z`, `0–9`) are identical in both alphabets, which is why a value encoded with the wrong variant often *looks* fine until a byte lands on index 62 or 63. ## Why the URL alphabet exists The three characters the standard alphabet can emit that a URL dislikes are `+`, `/` and `=`. - `/` terminates a path segment, so `GET /r/3q2+7w/` and `GET /r/ab/cd` are structurally different requests. - `+` is decoded as a space by `application/x-www-form-urlencoded` parsing, so an identifier that survives a path can still be corrupted in a query string. - `=` is legal in a query value but must be percent-escaped in other positions, and it makes an ugly identifier. The same three characters are awkward in filenames and in shell arguments. So for a link shortener issuing opaque identifiers — 16 random bytes, say — `base64.RawURLEncoding.EncodeToString(id)` gives 22 characters made only of `A–Z a–z 0–9 - _`, safe in a path, a query, a filename and a log line without any escaping layer. ## What dropping the padding does and does not cost Padding exists so that a stream of concatenated base64 blocks can be split back apart on a multiple-of-four boundary. It carries no information about the data. A decoder that is told there is no padding recovers the length from the character count: 4 characters decode to 3 bytes, 3 characters to 2 bytes, 2 characters to 1 byte, and a remainder of exactly 1 character is invalid input. The length arithmetic is exposed: - `base64.StdEncoding.EncodedLen(16)` is 24 — `(16+2)/3*4`, including two `=`. - `base64.RawURLEncoding.EncodedLen(16)` is 22 — `(16*8+5)/6`. Two characters saved per identifier is not the point; not needing an escaping story is. ## The decoders are strict about the variant Each `*Encoding` decodes only its own form. - `base64.StdEncoding.DecodeString("3q2-7w==")` fails: `-` is not in the standard alphabet. - `base64.RawStdEncoding.DecodeString("3q2+7w==")` fails: with padding disabled, `=` is just an illegal character. - `base64.StdEncoding.DecodeString("3q2+7w")` fails too: the padded encoding wants a complete four-character quantum. The error is a `base64.CorruptInputError`, whose message names the offending input byte position. That is exactly the report you get from a partner who encoded with one variant and decoded with another, and the fastest triage is to look at the value itself: a `-` or `_` means URL alphabet, a trailing `=` means padded, and a length that is not a multiple of four means a `Raw` variant. If you must accept both forms — an unavoidable position for a public endpoint that has already shipped two encoders — normalise the string before decoding, or try one decoder and fall back to the other. Do not invent a third alphabet. ## Building your own encoding `base64.NewEncoding(alphabet)` takes a 64-character alphabet and returns an `*Encoding`. Two chainable modifiers refine it: - `WithPadding(base64.NoPadding)` turns off padding — this is how the `Raw` variants relate to the padded ones — or `WithPadding('.')` picks a different pad byte. - `Strict()` returns an encoding that additionally rejects a final quantum whose unused trailing bits are non-zero, which matters when you need canonical encodings that cannot be mutated without detection. A custom alphabet is rarely the right call for a new system: it makes your identifiers undecodable by every off-the-shelf tool a partner might reach for. ## What an interviewer is listening for That you reach for `RawURLEncoding` for anything that goes in a URL rather than percent-escaping standard base64 after the fact; that you know padding is a framing device, not data; and that you understand the decoder must match the encoder, because that mismatch is the bug report this API generates most often.
- Does dropping the padding lose information?No. Padding only rounds the output up to a multiple of four so concatenated blocks can be split apart. The decoder derives the byte count from the character count — three characters mean two bytes, two characters mean one byte — and a remainder of exactly one character is invalid input either way.
- A partner sends an identifier your base64.StdEncoding decoder rejects. How do you triage it in one look?Read the string. A `-` or `_` means they used the URL alphabet; a trailing `=` means padded; a length that is not a multiple of four means a Raw variant. Match the variant to the value, then fix the contract so both sides name the same encoding explicitly.
- How would you build an unpadded encoding over a custom alphabet?`base64.NewEncoding(alphabet).WithPadding(base64.NoPadding)` returns an `*Encoding` with your 64 characters and no pad byte. It is rarely worth it: a non-standard alphabet means no other tool or language can decode your values without your code.
saying these in an interview costs you the question
- Thinks URLEncoding also drops the padding
- Percent-escapes standard base64 instead of using the URL alphabet
- Believes standard and URL values decode interchangeably
- Claims removing padding loses the length
- Invents a custom alphabet for a public identifier