In Dart, how do the utf8, base64 and latin1 codecs differ, and what happens when utf8.decode receives bytes that are not valid UTF-8?
answer
- text to bytes: utf8, latin1
- bytes to text-safe string: base64
- utf8.encode returns Uint8List
- malformed bytes: FormatException or U+FFFD
- latin1 only covers code points 0 to 255
basics
~10 sutf8 and latin1 convert between String and bytes; base64 converts bytes to ASCII text and back. utf8.decode throws a FormatException on malformed bytes unless allowMalformed is true, which substitutes U+FFFD.
solid answer
~40 s`utf8` is a `Codec<String, List<int>>`: `utf8.encode` returns a `Uint8List` of UTF-8 bytes, covering all of Unicode, and `utf8.decode` turns bytes back into a `String`, throwing a `FormatException` on invalid or truncated sequences unless you pass `allowMalformed: true`, which replaces them with U+FFFD, and dropping a leading byte-order mark. `latin1` maps each character to one byte and only covers code points 0 to 255: encoding a character outside that range throws an `ArgumentError`, and decoding a value above 255 throws a `FormatException` unless `allowInvalid` is set. `base64` works on bytes, not strings: `base64.encode(bytes)` gives ASCII text safe to embed in JSON or headers, `base64.decode` reverses it, and `base64Url` uses the URL-safe alphabet. So encoding a string as base64 is two steps, `base64.encode(utf8.encode(s))`, or one fused codec.
code
dart · 18 linesimport 'dart:convert';
void main() {
const note = 'Café ☕';
final bytes = utf8.encode(note); // Uint8List
print(note.length); // 6 UTF-16 code units
print(bytes.length); // 9 UTF-8 bytes
final token = base64Url.encode(bytes); // bytes -> URL-safe text
print(token); // Q2Fmw6kg4piV
print(utf8.decode(base64Url.decode(token))); // Café ☕
print(utf8.decode([0x43, 0xFF], allowMalformed: true)); // C followed by U+FFFD
// utf8.decode([0xFF]); // throws FormatException
print(latin1.encode('Café')); // [67, 97, 102, 233]
// latin1.encode('☕'); // throws ArgumentError: outside 0..255
}go deeper
Know that utf8 converts text to bytes and back, and base64 converts bytes to text-safe strings; they solve different problems.
Explain FormatException versus allowMalformed, the Uint8List return type, latin1's 0 to 255 limit, and base64Url.
Choose the failure policy for untrusted bytes, avoid codeUnits as a wire format, and keep base64 for small binary values.
Standardise encodings across the app's formats and protocols so every boundary uses UTF-8 and documented base64 alphabets.
## Three codecs, two jobs `dart:convert` exposes ready-made `Codec` constants. They solve two different problems: | Codec | Converts | Typical use | |---|---|---| | `utf8` | `String` to UTF-8 bytes and back | files, sockets, HTTP bodies | | `latin1` | `String` to one byte per character and back | legacy protocols, ISO-8859-1 data | | `base64`, `base64Url` | bytes to ASCII text and back | binary data inside JSON, URLs or headers | Mixing up the two jobs is the most common mistake: base64 is not a text encoding, it is a way to carry **bytes** through channels that only accept text. ## `utf8` - `utf8.encode(string)` returns a **`Uint8List`**. Since Dart 3.2 the declared return type is `Uint8List`, not `List<int>`. - A Dart `String` is a sequence of UTF-16 code units, so `string.length` and the UTF-8 byte count differ: `'Café ☕'` has 6 code units but 9 UTF-8 bytes. - `utf8.decode(bytes)` throws a **`FormatException`** for invalid or unterminated sequences. - `utf8.decode(bytes, allowMalformed: true)`, or a codec created with `Utf8Codec(allowMalformed: true)`, replaces bad sequences with the replacement character **U+FFFD** instead of throwing. - A leading byte-order mark is discarded when decoding. - `utf8.encoder` and `utf8.decoder` are `Converter`s, so they also work as stream transformers for chunked input. ## `latin1` - Each character becomes exactly one byte, which only works for code points **0 to 255**. - `latin1.encode('☕')` throws an **`ArgumentError`** because the string contains characters outside that range. - `latin1.decode` throws a `FormatException` for values outside 0 to 255 unless `allowInvalid: true`, which substitutes U+FFFD. - Use it only when a protocol really specifies ISO-8859-1; for new formats use UTF-8. ## `base64` 1. `base64.encode(bytes)` returns a `String` of the standard alphabet with `=` padding; `base64Encode` is a top-level shorthand. 2. `base64.decode(text)` returns a `Uint8List` and throws a `FormatException` on invalid input. 3. `base64Url` uses `-` and `_` instead of `+` and `/`, safe in URLs and file names. 4. `base64.normalize(text)` validates input and rewrites it into the standard alphabet with correct padding, useful before comparing tokens. Base64 output is about a third larger than its input, so it suits small binary values, such as an avatar thumbnail inside a settings document, better than large files. ## Composing them Because each is a `Codec`, they compose with `fuse`: - `utf8.fuse(base64)` is a `Codec<String, String>` that turns text into base64 and back in one call. - `json.fuse(utf8)` turns JSON-encodable objects into UTF-8 bytes directly. ## Diagnosing garbled text Text that shows up as `Café` instead of `Café` is the classic symptom of decoding UTF-8 bytes with the wrong codec. The two UTF-8 bytes of `é`, `0xC3 0xA9`, read as Latin-1 become two characters, `Ã` and `©`. The reverse mistake, decoding Latin-1 bytes as UTF-8, usually fails with a `FormatException` instead, because a lone byte above 127 is not valid UTF-8. So: - mojibake with `Ã` sequences means UTF-8 data was decoded as Latin-1 somewhere; - a `FormatException` from `utf8.decode` often means the data was never UTF-8; - fix the declared encoding at the source rather than adding `allowMalformed` to hide it. ## Choosing correctly - Receiving bytes you did not produce: decode with `utf8`, and decide deliberately between failing and `allowMalformed`, depending on whether corrupt text is acceptable. - Embedding binary data in JSON: base64 it, and record which alphabet you used. - Never use `string.codeUnits` as a byte encoding for the wire; those are UTF-16 units, and anything above 255 does not fit in a byte.
- How do you store a string as base64 in a settings file?Encode it to bytes first, then to base64: `base64.encode(utf8.encode(text))`, and reverse with `utf8.decode(base64.decode(stored))`. `utf8.fuse(base64)` wraps both steps into one `Codec<String, String>`. Calling base64 on the string itself is impossible, because base64 works on bytes.
- When would you set allowMalformed: true?When displaying text you cannot control and a replacement character is better than failing, for example log lines from a device. For data you must process correctly, such as a settings file, keep the default so corrupt input raises a `FormatException` you can report.
saying these in an interview costs you the question
- base64 is a text encoding like UTF-8
- String.length equals the number of UTF-8 bytes
- utf8.decode silently skips invalid bytes by default
- latin1 can encode any Unicode character
- String.codeUnits gives the UTF-8 bytes of a string