A Go log pager truncates each line with s[:80] and users see mangled characters — what is happening and how do you cut safely?
answer
- The index unit is not the character unit
- Outside ASCII a code point is two to four bytes
- The cut landed inside an encoded sequence
- The tail is invalid UTF-8, drawn as U+FFFD
- Cut where a rune starts, not at byte 80
basics
~20 ss[:80] cuts at byte 80, which can land inside a multi-byte UTF-8 sequence and leave a half-encoded rune that the terminal draws as a replacement character. Cut at a rune boundary instead, then worry about display width separately.
solid answer
~50 sIndexes into a Go string are byte offsets, and UTF-8 encodes anything outside ASCII in two to four bytes. Slicing at byte 80 can therefore split a rune in half, and the trailing fragment is invalid UTF-8 that a terminal renders as U+FFFD — the classic mangled tail. The fix is to cut where a rune starts: range over the string, which yields the byte offset of each rune's first byte, and slice at the offset of the rune you want to stop before; or walk back with `utf8.DecodeLastRuneInString` until the last rune decodes cleanly. `utf8.RuneCountInString` tells you how many runes a line holds, and `utf8.ValidString` tells you whether a fragment is well-formed. Note that even a correct rune cut is not a correct *column* cut: combining marks add no width and wide CJK glyphs take two cells, so a pager that pads columns needs a width calculation on top.
code
go · 10 linesfunc truncateRunes(s string, n int) string {
count := 0
for i := range s { // i is the byte offset where each rune starts
if count == n {
return s[:i]
}
count++
}
return s
}go deeper
Be ready to say that string indexes and len are in bytes, and that non-ASCII text uses several bytes per character, so slicing at a fixed index can split one.
Explain the UTF-8 encoding widths, that ranging over a string yields the byte offset where each rune starts, and write the rune-safe truncation loop on the spot.
Diagnose the symptom from the report alone, name the downstream systems that will reject invalid UTF-8 rather than display it, choose the cut strategy against the pager's redraw cost, and volunteer the grapheme and display-width caveat.
Own where text is validated and normalised in the system. Deciding that lines are checked once at ingest, rather than defensively repaired at every render site, is the call that keeps this class of bug from recurring.
## What the code actually asked for `s[:80]` is a byte-offset slice expression. It builds a new string header pointing at the same array with length 80. It does not know or care what those 80 bytes encode. Go source is UTF-8 and Go strings conventionally hold UTF-8, in which a code point occupies one to four bytes: ASCII is one, most Latin-1 accented letters and Cyrillic are two, most CJK is three, emoji are four. So byte 80 lands inside a rune whenever the first 80 bytes do not end on a rune boundary — which, for any non-ASCII log line, is most of the time. ## Why the output is mangled The truncated string ends with a partial UTF-8 sequence: a lead byte with some or none of its continuation bytes. A terminal decoding that stream cannot form a code point and substitutes U+FFFD, the replacement character. Anything downstream that validates UTF-8 — a JSON encoder, a database column with a UTF-8 constraint, a log shipper — may reject the line instead, so the same bug shows up as "invalid UTF-8" errors in systems that never displayed anything. A related symptom on the read side: `s[i]` yields a `byte` (a `uint8`), not a `rune`. Printing it with `%c` prints whatever code point that single byte number happens to be, so mid-rune bytes print as garbage Latin-1-looking characters. Anyone doing `for i := 0; i < len(s); i++ { fmt.Printf("%c", s[i]) }` has written the same bug in a different shape. ## Cutting on a rune boundary Ranging over a string decodes UTF-8 and gives you `i` = the byte offset at which each rune begins, and `r` = the decoded rune. That makes the safe truncation a short loop: count runes, and when you reach the limit, slice at the current offset — which is by construction a rune boundary. The `unicode/utf8` package gives you the same thing in pieces: - `utf8.RuneCountInString(s)` — how many runes the string decodes to (as opposed to `len(s)`, which is bytes). - `utf8.DecodeRuneInString(s)` — the first rune and its size in bytes; returns `utf8.RuneError` with size 1 on a malformed byte. - `utf8.DecodeLastRuneInString(s)` — the same from the end, which is how you trim a fragment back off a cut you already made. - `utf8.ValidString(s)` — whether the whole string is well-formed UTF-8, useful as a test assertion or a guard at an ingest boundary. - `utf8.RuneLen(r)` — how many bytes a given rune will need. Converting with `[]rune(s)` and slicing that is also correct, and it is the obvious move for someone coming from a language with fixed-width characters. It costs an allocation and a full decode of the line, which matters in a pager redrawing hundreds of lines per keystroke; the range-based cut touches only the first 80-odd bytes of each line. ## Runes are still not what the user sees Cutting on a rune boundary makes the output *valid*; it does not make it the right *width*. Three separate notions are in play: - **Bytes** — what `len` and slicing count. - **Runes / code points** — what `range` and `[]rune` yield. - **User-perceived characters (grapheme clusters)** — `e` plus a combining acute is two runes and one visible character; a flag emoji or an emoji with a skin-tone modifier is several runes and one glyph. And for a terminal there is a fourth: **cells**. A wide CJK glyph occupies two columns; a combining mark occupies zero. A pager that aligns columns has to measure display width, not runes, and the stdlib does not provide that measurement — it is a table lookup that lives outside the standard library. Cutting mid-cluster also produces visible nonsense (a lone combining accent landing on whatever follows), even though every byte is valid UTF-8. ## What to say in an interview Diagnose it as "indexes are byte offsets, UTF-8 is variable width, so the cut split a rune", show the range-offset fix, mention `utf8.DecodeLastRuneInString` for repairing a cut you inherited, and then volunteer the honest caveat that rune-correct is not the same as column-correct. The last part is what separates someone who has shipped a terminal UI from someone who has read the spec.
- What type does `s[3]` give you for a Go string, and what does printing it with %c show?A `byte` (`uint8`), not a `rune`. Printing it with `%c` prints the code point whose number equals that single byte, so for a mid-sequence byte of a multi-byte rune you get an unrelated character. Use `range` or `utf8.DecodeRuneInString` when you want a code point.
- Why not just do `[]rune(s)[:80]` and convert back?It is correct, and fine for occasional use. It allocates a rune array and decodes the whole line even though you only need the first 80 runes, which matters in a pager redrawing many lines per keystroke. The range-based cut touches only the bytes it must.
- After you cut on rune boundaries, why can the columns still not line up?Runes are not cells. A wide CJK glyph takes two terminal columns and a combining mark takes none, so equal rune counts can render at different widths. Aligning columns needs a display-width calculation, which the standard library does not provide; cutting mid-grapheme-cluster also leaves stray combining marks.
- How would you assert in a test that a truncation function never emits broken text?Assert `utf8.ValidString(out)` for a corpus of lines containing multi-byte runes at every offset around the limit, and assert `utf8.RuneCountInString(out) <= n`. Table-driven cases with a two-byte, a three-byte and a four-byte rune straddling the cut catch every off-by-one.
Slicing at byte 80 is like cutting a sentence off at the eightieth pen stroke rather than the eightieth letter: you can end mid-letter, and what is left is not a letter at all.
saying these in an interview costs you the question
- Says Go strings are broken or should store characters
- Believes len(s) and the index unit are characters
- Thinks s[i] yields a rune rather than a byte
- Converts to []rune on every line without noticing the cost
- Assumes a rune-correct cut is also a column-correct cut