What does utf8.DecodeRuneInString return for invalid bytes, and how do you spot a genuine U+FFFD?
answer
- two results, and the second matters
- it never returns an error value
- bad input advances exactly one byte
- U+FFFD is also a real character
- size 1 versus size 3 decides it
basics
~10 sIt returns utf8.RuneError with size 1 for invalid bytes, and size 0 for an empty string. A genuine U+FFFD decodes with size 3, so the size, not the rune, tells corruption apart.
solid answer
~40 s`utf8.DecodeRuneInString(s)` returns `(r rune, size int)`: the first rune and how many bytes it consumed. On ill-formed input it returns `(utf8.RuneError, 1)` — it reports the replacement character U+FFFD and advances exactly one byte so a caller can keep going. On an empty string it returns `(utf8.RuneError, 0)`. Both of those size values are impossible for correct non-empty UTF-8, which is precisely how you detect trouble: U+FFFD is itself an ordinary character and encodes to three bytes, so a *real* U+FFFD in the input comes back as `(utf8.RuneError, 3)`. Comparing the rune to `utf8.RuneError` alone therefore cannot distinguish "the input was broken" from "the input contained a replacement character"; you must look at the size. `utf8.DecodeLastRuneInString` is the mirror image for walking backwards.
code
go · 5 linesr1, n1 := utf8.DecodeRuneInString("\ufffd") // a genuine U+FFFD
fmt.Println(r1 == utf8.RuneError, n1) // true 3
r2, n2 := utf8.DecodeRuneInString("\xff") // an impossible byte
fmt.Println(r2 == utf8.RuneError, n2) // true 1go deeper
Recall that this decoder hands back two things, a rune and a byte count, and that Go substitutes the replacement character rather than failing when bytes are malformed.
Explain the two impossible results, size 0 for empty and size 1 for ill-formed, and show a loop that uses the size to advance a byte cursor correctly.
Demonstrate that you would never gate validation on the rune value alone, because U+FFFD is legitimate input; show the size check and say where in a pipeline you would place it.
Decide the platform-wide policy: whether ingress rejects ill-formed text outright, repairs it, or passes it through, and make sure every service applies the same rule so a repair is not undone downstream.
## The signature and its contract ``` func DecodeRuneInString(s string) (r rune, size int) func DecodeRune(p []byte) (r rune, size int) func DecodeLastRuneInString(s string) (r rune, size int) ``` Each decodes **one** rune — the first (or, for the `Last` variants, the final one) — and reports how many **bytes** that rune occupied. The size is what lets you advance a manual cursor: ``` for i := 0; i < len(s); { r, size := utf8.DecodeRuneInString(s[i:]) // use r i += size } ``` The size result is what makes manual decoding possible at all: UTF-8 is variable width, so the decoder is the only thing that knows where the next rune starts. ## The two impossible results The package documents two return shapes that correct, non-empty UTF-8 can never produce: - `(utf8.RuneError, 0)` — the input was **empty**. There was nothing to decode. - `(utf8.RuneError, 1)` — the input was **ill-formed**. The bytes at the front are not a valid UTF-8 sequence. That second case is deliberately non-fatal. Rather than returning an error or panicking, the decoder yields the Unicode replacement character and steps forward exactly one byte, so a loop over corrupt input makes progress instead of spinning. This is the same policy the rest of Go applies to bad UTF-8: substitute U+FFFD, advance one byte, keep going. `utf8.RuneError` is a constant equal to `'\uFFFD'`, the REPLACEMENT CHARACTER — the black-diamond question mark, or an empty box, depending on the font. `unicode.ReplacementChar` is the same value under a different name. ## Why the size, not the rune, is the error signal U+FFFD is a normal Unicode character. Nothing stops a user from typing it, and nothing stops an upstream system from having already substituted it for its own corrupt data. When that character appears legitimately in the input it is encoded as three bytes (EF BF BD) and decodes to `(utf8.RuneError, 3)`. So: | input | result | meaning | |---|---|---| | `""` | `(RuneError, 0)` | empty | | a byte like 0xFF | `(RuneError, 1)` | ill-formed | | the encoded U+FFFD | `(RuneError, 3)` | a real replacement character | Code that writes `if r == utf8.RuneError { return errBadUTF8 }` rejects perfectly valid input containing U+FFFD, and that bug tends to survive review because the comparison *looks* right. The correct test is `r == utf8.RuneError && size <= 1`. ## When you would decode by hand at all Most code should not. `for i, r := range s` decodes for you, `utf8.RuneCountInString` counts for you, and `[]rune(s)` converts for you. Manual decoding earns its place when: - You need to **distinguish corrupt input from a real U+FFFD**, which only the size result can do. - You are walking **backwards** (`DecodeLastRuneInString`), for instance to trim or inspect a trailing rune. - You are implementing a **streaming** parser over `[]byte` where a rune may straddle a chunk boundary and you must know whether the tail is a partial sequence. - You want to stop at the **first** rune without scanning the whole string. ## Companion checks - `utf8.ValidString(s) bool` answers "is all of this well formed" in one call, and is what you want for a validation gate. - `utf8.RuneStart(b byte) bool` reports whether a byte could begin a sequence — it is the cheap way to find a rune boundary without decoding. - `utf8.ValidRune(r rune) bool` asks a different question: whether a **rune value** can legally be encoded at all (surrogate halves and values above U+10FFFF cannot). - `strings.ToValidUTF8(s, replacement)` rewrites ill-formed runs rather than reporting them. ## What an interviewer is checking That you can read a two-result decoder contract precisely, that you know Go's substitute-and-advance policy for bad UTF-8, and that you spot the trap where U+FFFD is both the error sentinel and a legitimate character. The last point is the one that separates a candidate who has read the documentation from one who has debugged text pipelines.
- Why does the decoder advance one byte on bad input instead of skipping the whole malformed sequence?Because there is no reliable way to know how long a malformed sequence was meant to be. Advancing exactly one byte guarantees forward progress and makes the behaviour deterministic: a run of n invalid bytes yields n replacement characters. It also matches the substitution policy the rest of Go uses, so a manual loop and a range loop agree on the same output.
- How does utf8.ValidRune differ from utf8.ValidString?`utf8.ValidString` asks whether a string's bytes are a well-formed UTF-8 encoding. `utf8.ValidRune` asks whether a single rune value can be encoded at all: it is false for negative values, for values above U+10FFFF, and for the surrogate range U+D800..U+DFFF, which UTF-8 forbids. One validates bytes, the other validates a code point.
- When would you use utf8.DecodeLastRuneInString?When you are working from the end: inspecting or removing a trailing rune, checking whether a string ends in a particular character class, or backing a byte cursor off the end of a buffer. It has the same contract as the forward version, including returning size 1 for an ill-formed tail, so the same corruption test applies.
saying these in an interview costs you the question
- Treats r == utf8.RuneError alone as proof of invalid input
- Expects an error result from the decoder
- Thinks the decoder panics on malformed bytes
- Assumes size is always 1
- Confuses ValidRune with ValidString