Concretely, what goes wrong when you read a UTF-8 text file through a byte stream and treat each byte as a character?
answer
- UTF-8 is variable-length: 1-4 bytes per char
- byte-as-char splits multi-byte chars -> mojibake
- café bytes 63 61 66 C3 A9; é = C3 A9
- CharsetDecoder reassembles continuation bytes
- fix: InputStreamReader/Files.* with explicit charset
basics
~10 sIn UTF-8 some characters take more than one byte. If you treat each byte as a character, those multi-byte characters get split into several wrong characters, so accented and non-Latin text comes out garbled.
solid answer
~50 sUTF-8 is a variable-length encoding: ASCII characters are one byte, but accented Latin letters take two bytes, most other scripts take three, and emoji take four. A byte stream gives you those raw bytes with no decoding. If you map each byte directly to a char, every multi-byte character is broken into two-to-four separate, meaningless char values — the classic 'mojibake' you see when 'café' shows up as 'café'. The reverse error also happens on write: if your text contains non-ASCII chars and you push them out as single bytes, you lose information. Character streams avoid this by running bytes through a CharsetDecoder that recognizes UTF-8's byte patterns and assembles complete characters. The fix is to never interpret bytes as text yourself: wrap the byte stream in an InputStreamReader with the correct charset, or use Files.readString/newBufferedReader with an explicit charset.
code
java · 15 linesimport java.nio.charset.StandardCharsets;
// The bytes for the string "café" encoded as UTF-8.
byte[] utf8 = "café".getBytes(StandardCharsets.UTF_8);
// -> [0x63, 0x61, 0x66, 0xC3, 0xA9] (é is the two bytes C3 A9)
// WRONG: interpret each byte as a Latin-1 character (1 byte = 1 char).
String wrong = new String(utf8, StandardCharsets.ISO_8859_1);
System.out.println(wrong); // prints: café (mojibake, length 5)
System.out.println(wrong.length()); // 5
// RIGHT: decode the bytes as UTF-8.
String right = new String(utf8, StandardCharsets.UTF_8);
System.out.println(right); // prints: café (length 4)
System.out.println(right.length()); // 4go deeper
Knows that using a byte stream on text can garble accented or non-English characters and that you should use a Reader instead.
Explains UTF-8 is variable-length so multi-byte characters get split when each byte is treated as a char.
Gives the concrete byte-level mechanism, the mojibake example, the write-side mirror error, and the explicit-charset fix.
Adds decoder internals: continuation bytes, buffer-boundary splitting, CodingErrorAction policies, BOM handling, and why centralizing the boundary in a CharsetDecoder is the robust design.
## Background: what a charset really is Text must be stored as bytes, and a **charset** (character encoding) is the rule that maps characters to byte sequences. **UTF-8** is the dominant one. Its defining property is that it is **variable-length**: | Character range | Bytes in UTF-8 | |---|---| | ASCII (`A`, `0`, space) | 1 byte | | Latin accents (`é`, `ñ`), Greek, Cyrillic | 2 bytes | | Most CJK (`中`, `日`) | 3 bytes | | Emoji, rare scripts | 4 bytes | UTF-8 encodes the length in the high bits of the first byte; **continuation bytes** all start with the bit pattern `10xxxxxx`. A correct decoder reads the lead byte, learns how many bytes follow, and assembles them into one character. ## The bug: bytes interpreted as characters A **byte stream** (`InputStream`) hands you raw bytes and does **no decoding**. If you take each byte and widen it to a `char` (e.g. by using ISO-8859-1 implicitly, or by casting), you are asserting "1 byte = 1 character" — which is only true for ASCII. Concrete example: the string `café` in UTF-8 is the bytes `63 61 66 C3 A9`. The `é` is the **two** bytes `C3 A9`. - **Correct (UTF-8 decode):** `c`, `a`, `f`, `é` — 4 characters. - **Wrong (byte-as-char):** `c`, `a`, `f`, `Ã`, `©` — 5 characters. The two bytes of `é` became two separate Latin-1 characters `Ã` and `©`. This garbled output is called **mojibake**. The more non-ASCII content, the worse the corruption; for a 3-byte CJK character you get three junk characters each. ## The reverse: the write side The symmetric error occurs on output. If your in-memory `String` holds `中` (one char) and you write it through a byte stream as a single byte, you truncate the character — the high bits are lost and the file no longer contains valid text. Encoding through a `Writer` with the right charset produces the correct multi-byte sequence. ## Why character streams fix it A **character stream** (`Reader`) — or more precisely the `InputStreamReader` bridge underneath it — runs incoming bytes through a `CharsetDecoder`. The decoder understands UTF-8's lead/continuation byte structure, so `C3 A9` is recognized as the single character `é`. On write, a `CharsetEncoder` turns each char back into the correct byte sequence. The byte/char boundary is handled correctly in exactly one place. ## Related subtleties a senior should mention - **Buffer-boundary splitting:** even with a correct charset, naive code that decodes fixed-size byte chunks can split a multi-byte character across two chunks. `InputStreamReader`/`CharsetDecoder` handle this by holding back the partial bytes; hand-rolled byte-to-string conversion (`new String(bytes, off, len)` per chunk) can corrupt at the seam. This is a strong argument for using the provided character streams. - **Malformed input:** real decoders let you choose what to do with invalid byte sequences via `CodingErrorAction` (REPLACE with `�`, IGNORE, or REPORT/throw). Byte-as-char never reports an error — it silently corrupts. - **BOM:** some files start with a byte-order mark; UTF-8 BOM (`EF BB BF`) can leak in as a stray character if not handled. ## The fix Never interpret bytes as text by hand. Either: 1. Wrap the byte stream: `new InputStreamReader(in, StandardCharsets.UTF_8)` (then a `BufferedReader`), or 2. Use the high-level helpers: `Files.readString(path, UTF_8)`, `Files.newBufferedReader(path, UTF_8)`. Always state the charset explicitly so the same code behaves identically on every machine.
- Why does the bug often pass tests but fail in production?Tests frequently use plain ASCII, which is one byte per character in UTF-8, so byte-as-char happens to be correct. The corruption only surfaces with accented or non-Latin input, which appears with real user data.
- What happens to a multi-byte character split across two read buffers?A naive per-buffer byte-to-String conversion corrupts the seam. A proper CharsetDecoder (used by InputStreamReader) retains the leftover partial bytes and completes the character on the next read.
saying these in an interview costs you the question
- Claiming UTF-8 is fixed 2 bytes per char (that is UTF-16-ish, and still not fixed)
- Saying it 'just works' because ASCII tests pass — ASCII is the one safe case
- Decoding fixed byte chunks with new String per chunk, ignoring boundary splitting
- Assuming a wrong charset throws an error rather than silently corrupting