skip to content

What charset do Kotlin's File text helpers use by default, and what concrete bugs arise if a file's actual encoding differs from what you read it with?

level: seniorimportance: should knowfreq 40%

answer

  1. All text helpers default to UTF-8 (not platform default)
  2. Wrong charset -> \uFFFD replacement char or mojibake
  3. Write/read must use the same charset
  4. Helpers don't strip a UTF-8 BOM (\uFEFF)
  5. Pass Charset explicitly for foreign files

basics

~20 s

Kotlin's readText, writeText, readLines, etc. default to UTF-8. If a file is actually in another encoding (like Latin-1 or UTF-16), reading it as UTF-8 corrupts non-ASCII characters or throws, because the bytes are decoded with the wrong rules.

solid answer

~50 s

All the text helpers — `readText`, `writeText`, `appendText`, `readLines`, `useLines`, `forEachLine`, `bufferedReader` — default to `Charsets.UTF_8`, not the JVM platform default. That's deliberate: it makes behavior portable across OSes (unlike Java's old `FileReader`, which used the platform default). If the file's real encoding differs, you decode bytes with the wrong table: a Latin-1 'é' (one byte 0xE9) read as UTF-8 is an invalid sequence and surfaces as the replacement char '\uFFFD' (or throws with a strict CharsetDecoder); UTF-16 read as UTF-8 produces garbage and stray nulls. Round-tripping is also a trap: writing UTF-8 and reading back as UTF-16 mangles everything. The fix is to pass the correct `Charset` explicitly, e.g. `file.readText(Charsets.ISO_8859_1)`, and to standardize on UTF-8 end-to-end. Also beware a UTF-8 BOM: the helpers don't strip it, so the first character may be an invisible '\uFEFF'.

code

kotlin · 11 lines
kotlin
import java.io.File

val f = File("data.txt")
f.writeText("naïve façade", Charsets.ISO_8859_1)

println(f.readText())                    // naïve fa\uFFFDade  (wrong: UTF-8 default)
println(f.readText(Charsets.ISO_8859_1)) // naïve façade        (correct)

// BOM trap
val withBom = "\uFEFFhello"
println(withBom == "hello")              // false — invisible BOM prefix

go deeper

for a junior

Knows the default is UTF-8 and that you can pass a charset.

for a middle

Explains that a mismatch corrupts non-ASCII text and that write/read should share a charset.

for a senior

Details replacement char vs strict decoder, BOM handling, why UTF-8 default beats platform default, and treats it as a silent data-corruption risk.

for a principal

Establishes org-wide UTF-8 policy, defines strict-decoding validation at boundaries, and reasons about i18n correctness for external inputs at scale.

## The default is UTF-8 — by design Every Kotlin stdlib text helper has a `charset: Charset = Charsets.UTF_8` parameter: `readText`, `writeText`, `appendText`, `readLines`, `useLines`, `forEachLine`, `bufferedReader`, `bufferedWriter`. A *charset* is the mapping between raw bytes on disk and characters in memory. Kotlin defaulting to **UTF-8** (rather than the JVM's locale-dependent platform default, as Java's legacy `FileReader`/`FileWriter` did) makes file I/O deterministic across machines. (Note: modern JDK 18+ also defaults file.encoding to UTF-8 via JEP 400, but Kotlin's explicit default predates and doesn't depend on that.) ## What goes wrong on a mismatch Reading is *decoding* (bytes -> chars); writing is *encoding* (chars -> bytes). If your assumed charset differs from the file's real one: - **Latin-1 (ISO-8859-1) read as UTF-8:** 'é' is the single byte `0xE9`. UTF-8 expects a multi-byte sequence there; the lone byte is invalid and the default decoder substitutes the **replacement character** `\uFFFD` (''). - **UTF-16 read as UTF-8:** every other byte is `0x00`; you get interspersed null characters and mojibake. - **Write-then-read mismatch:** `writeText(s)` (UTF-8) then `readText(Charsets.UTF_16)` returns corrupted text — silent data loss, not an exception. - **Strict decoders throw:** a `CharsetDecoder` configured with `CodingErrorAction.REPORT` raises `MalformedInputException` instead of substituting — the helpers use the lenient default (replace). ```kotlin import java.io.File val f = File("latin.txt") f.writeText("café", Charsets.ISO_8859_1) // 4 bytes: c a f 0xE9 f.readText() // "caf\uFFFD" ← wrong (UTF-8 default) f.readText(Charsets.ISO_8859_1) // "café" ← correct ``` ## The BOM trap A *Byte Order Mark* is an optional prefix (`0xEF 0xBB 0xBF` for UTF-8) some editors add. Kotlin's helpers **don't strip it**, so `readText()` may start with an invisible `\uFEFF`, breaking equality checks, JSON/XML parsing, or key matching. Strip it explicitly: `text.removePrefix("\uFEFF")`. ## Practical guidance - **Standardize on UTF-8** for everything you control; only override `charset` when consuming foreign files. - **Always pass the charset explicitly** when you know the source encoding: `readText(Charsets.ISO_8859_1)`. - **Round-trip with the same charset** for write/read pairs. - **Detect or document** the encoding of external inputs; charset auto-detection is heuristic and unreliable. - For strict validation, build a `CharsetDecoder` with `CodingErrorAction.REPORT` and stream through it instead of relying on the lenient helpers. ## Why this is a senior-level concern The bugs are silent and data-dependent: ASCII-only test files pass, then real user data with accents, emoji, or CJK characters corrupts in production. Pinning charset explicitly is a cheap insurance against an entire class of internationalization defects.

  • Why did Kotlin default to UTF-8 instead of the platform default like old Java FileReader?
    For deterministic, portable behavior. The platform default varied by OS/locale, causing files to read differently on different machines; UTF-8 removes that ambiguity.
  • How would you make decoding fail loudly on bad bytes instead of substituting \uFFFD?
    Use a CharsetDecoder with onMalformedInput(CodingErrorAction.REPORT) and stream bytes through it; the lenient stdlib helpers replace by default.

Reading a file with the wrong charset is like reading Morse code as semaphore — same signal, wrong decoding rules, garbled message.

saying these in an interview costs you the question

  • Believing the helpers use the platform/locale default charset
  • Assuming wrong-charset reads always throw (they usually substitute silently)
  • Writing with one charset and reading back with another
  • Ignoring the UTF-8 BOM and then failing equality/parse checks
  • Trusting heuristic charset auto-detection for correctness

context