Why is specifying a charset explicitly important when reading and writing text files in Java, and how does this differ across Java versions?
answer
- Charset = bytes<->chars; wrong one = mojibake
- Pre-18 default charset = OS/locale (Win-1252 vs UTF-8)
- Files.* helpers always defaulted to UTF-8
- JEP 400 / Java 18: UTF-8 is the JVM default
- Always pass StandardCharsets.UTF_8 explicitly
basics
~20 sA charset decides how characters become bytes and back. If you read a file with a different charset than it was written in, you get garbled text. Always pass StandardCharsets.UTF_8 explicitly so behavior is the same on every machine.
solid answer
~40 sText on disk is bytes; a charset (like UTF-8) maps between bytes and characters. Read with the wrong charset and non-ASCII characters get mojibake or decoding errors. The trap is the JVM 'default charset': before Java 18 it depended on the OS/locale (Windows-1252 on Windows, often UTF-8 on Linux), so APIs that used the default — old FileReader/FileWriter, InputStreamReader without a charset — behaved differently per environment, a classic 'works on my machine' bug. The newer Files helpers and Files.newBufferedReader/Writer always defaulted to UTF-8. From Java 18 (JEP 400) UTF-8 became the JVM-wide default, so even the legacy APIs now default to UTF-8 unless file.encoding is overridden. Best practice regardless of version: be explicit — pass StandardCharsets.UTF_8 — and handle malformed input deliberately rather than relying on defaults.
code
java · 9 linesimport java.nio.charset.StandardCharsets;
import java.nio.file.*;
// Explicit charset -> identical behavior on every OS and Java version
String text = Files.readString(Path.of("in.txt"), StandardCharsets.UTF_8);
Files.writeString(Path.of("out.txt"), text, StandardCharsets.UTF_8);
// Legacy, charset-dependent (pre-18 used the platform default):
// new FileReader("in.txt") // avoid unless you really want the defaultgo deeper
Knows files store bytes and you should use UTF-8 so text isn't garbled.
Can explain mojibake and that passing StandardCharsets.UTF_8 makes reads/writes consistent.
Explains the pre-Java-18 platform-default trap, which APIs used it, and that the Files helpers always defaulted to UTF-8.
Knows JEP 400's impact, plans upgrades (file.encoding=COMPAT, audits of getBytes/FileReader), and sets malformed-input handling policy.
## Bytes vs characters A file on disk is a sequence of **bytes** (numbers 0–255). Characters like 'A', 'é', or '€' are an abstraction. A **charset (character encoding)** is the rulebook mapping characters to byte sequences and back. **UTF-8** is the dominant one: ASCII characters take 1 byte, others take 2–4 bytes. Other charsets exist: **ISO-8859-1 (Latin-1)** and **Windows-1252** are single-byte Western encodings; **UTF-16** uses 2+ bytes per character. ## What goes wrong If a file was written as UTF-8 but you read it as Windows-1252, multi-byte characters are misinterpreted: 'é' (UTF-8 bytes C3 A9) becomes 'é'. This garble is called **mojibake**. The reverse (writing with the wrong charset) corrupts the file. With strict decoding, an invalid byte sequence can instead throw a `MalformedInputException`. ## The 'default charset' trap Java has a JVM-wide *default charset* (`Charset.defaultCharset()`, controlled by the `file.encoding` system property). **Before Java 18** this was derived from the operating system locale: typically **Windows-1252 on Windows**, **UTF-8 on most Linux**, sometimes others. APIs that *don't take a charset* used this default: - `new FileReader(file)` / `new FileWriter(file)` (the older constructors) - `new InputStreamReader(in)` / `new OutputStreamWriter(out)` with no charset - `String.getBytes()` with no argument Result: the same code read/wrote files differently on a developer's Mac vs a Windows CI box vs a Linux server — the textbook 'works on my machine' encoding bug. ## What already defaulted to UTF-8 The NIO additions were designed to avoid this: `Files.readString`, `Files.writeString`, `Files.readAllLines`, `Files.lines`, `Files.newBufferedReader`, `Files.newBufferedWriter` all **default to UTF-8** (and offer charset overloads). So preferring the `Files` helpers sidestepped the trap even before Java 18. ## JEP 400 — Java 18 makes UTF-8 the default **Java 18 (JEP 400, 2022)** changed `Charset.defaultCharset()` to **UTF-8** regardless of OS/locale. Now even `FileReader`/`FileWriter`/`String.getBytes()` default to UTF-8. You can still override with `-Dfile.encoding=...` (and `file.encoding=COMPAT` restores the old locale-derived behavior) for legacy compatibility. This removed a huge class of cross-platform bugs but means apps that *relied* on the old platform default may behave differently after upgrading. ## Best practice (every version) 1. **Be explicit:** pass `StandardCharsets.UTF_8` to constructors and `Files` overloads. Then version and OS no longer matter. 2. **Avoid the no-charset legacy constructors** unless you genuinely want the platform default. 3. **Decide on malformed input:** the defaults *replace* bad bytes with the replacement char U+FFFD; if you need to detect corruption, configure a `CharsetDecoder` with `CodingErrorAction.REPORT`. 4. **Round-trip with the same charset** you wrote with. ## One-liner to remember Text I/O without an explicit charset is a latent cross-platform bug; pass `StandardCharsets.UTF_8` and the problem disappears.
- Which JEP and Java version made UTF-8 the JVM-wide default charset?JEP 400, delivered in Java 18 (2022). Before that the default came from the OS locale.
- How can you restore the old locale-based default after upgrading to Java 18+?Set the system property file.encoding=COMPAT (or specify a concrete charset), though being explicit per-call is preferred.
- What is mojibake?Garbled text from decoding bytes with the wrong charset, e.g. UTF-8 'é' shown as 'é' when read as Windows-1252.
saying these in an interview costs you the question
- Believing the default charset was always UTF-8 (false before Java 18)
- Using new FileReader/FileWriter and assuming consistent encoding across OSes
- Reading with a different charset than the file was written with
- Ignoring malformed-input handling (silent replacement hides corruption)