skip to content

Why does truncating a Go display name with name[:32] leave a replacement box, and how do you fix it?

level: seniorimportance: should knowfreq 46%

answer

  1. the index is not what you think
  2. slice indices are byte offsets
  3. a rune can be four bytes wide
  4. back up off continuation bytes
  5. assert validity before you store

basics

~20 s

Slicing a Go string cuts at a byte offset, so a multi-byte rune gets split in half and the stored bytes are no longer valid UTF-8. Back the cut up to a rune boundary and check utf8.ValidString before storing.

solid answer

~40 s

`name[:32]` slices **bytes**, and a rune can be two, three or four bytes wide, so the cut lands mid-sequence whenever byte 32 is inside a rune. The tail left behind is an orphaned lead byte or continuation byte: no longer well-formed UTF-8, and every renderer downstream shows U+FFFD. The fix is to move the cut backwards to a rune boundary — walk down while `utf8.RuneStart(name[max])` is false, or decode forward and stop before you exceed the budget. Then assert the invariant at the storage boundary with `utf8.ValidString`, and repair anything that arrives already broken with `strings.ToValidUTF8`. To confirm the diagnosis on live data, hexdump the stored bytes: you will see a lone 0xF0 or a dangling 0x80-series continuation byte at the end where a complete sequence should be.

code

go · 10 lines
go
func truncateBytes(s string, max int) string {
	if len(s) <= max {
		return s
	}
	// s[max] is the first excluded byte; back up off any continuation byte.
	for max > 0 && !utf8.RuneStart(s[max]) {
		max--
	}
	return s[:max]
}

go deeper

for a junior

Remember that slicing a Go string uses byte offsets, so a cut can land inside a multi-byte character and produce bytes that no longer decode.

for a middle

Explain the one-to-four byte encoding, show how RuneStart identifies a boundary, and name the check that proves a stored value is well formed.

for a senior

Walk the diagnosis end to end: confirm with a validity assertion and a hexdump, fix the cut, repair existing rows, then place the invariant at ingress so the bug cannot return through another handler.

for a principal

Own the contract across services: define once whether length limits are counted in bytes or runes, who validates, and whether downstream systems may assume well-formed text, then make that assumption testable rather than folklore.

## What actually happened A Go string is a byte sequence, and slice expressions on it use **byte** indices. `name[:32]` therefore means "the first 32 bytes", not "the first 32 characters". UTF-8 encodes a rune in one to four bytes: | range | bytes | examples | |---|---|---| | U+0000..U+007F | 1 | ASCII | | U+0080..U+07FF | 2 | accented Latin, Greek, Cyrillic | | U+0800..U+FFFF | 3 | most CJK, U+FFFD itself | | U+10000..U+10FFFF | 4 | emoji, rare scripts | If byte 32 falls inside a four-byte emoji, the stored value ends with one, two or three bytes of that emoji and nothing else. Those bytes are not a valid sequence. Every decoder that later reads the value substitutes U+FFFD, which fonts draw as a black diamond or an empty box — the "box at the end" support is reporting. The bug is quiet in tests because ASCII test fixtures never trigger it. It appears the moment a real user types an emoji or a non-Latin script near the limit. ## Diagnosing it Two cheap steps confirm it without guessing: 1. **Assert the invariant.** Run `utf8.ValidString(stored)` over the affected rows. If the broken ones come back false, the corruption is in the bytes you stored, not in the renderer or the font. 2. **Look at the bytes.** Hexdump one offending value. A truncated four-byte emoji leaves a trailing `f0` or `f0 9f`; a truncated three-byte CJK character leaves `e6` or `e6 97`. Continuation bytes are always in the range `80..bf`, and a value ending in one of those with no lead byte to match is the signature. It is worth checking *where* the corruption entered. If the API rejects invalid UTF-8 at ingress and the database still holds broken rows, the truncation is yours; if ingress accepts anything, the client may have sent it. ## Fixing the truncation The direct fix is to move the cut backwards to a rune boundary. `utf8.RuneStart(b byte)` reports whether a byte can begin an encoded rune (it is false exactly for continuation bytes), so: ``` func truncateBytes(s string, max int) string { if len(s) <= max { return s } for max > 0 && !utf8.RuneStart(s[max]) { max-- } return s[:max] } ``` At most three decrements run, because no rune is wider than `utf8.UTFMax` (4) bytes. An equivalent approach decodes forward with `utf8.DecodeRuneInString`, accumulating sizes and stopping before the budget is exceeded; that one is easier to extend when you also want to stop at a whole word. Note what this **does not** fix: it still cuts a grapheme cluster. A truncated name can lose the skin-tone modifier from an emoji or the accent from a letter, leaving something valid but ugly. That is a product decision, not a correctness bug — the correctness bug is invalid UTF-8. ## Repairing what already exists For data already stored, or for input you did not produce, `strings.ToValidUTF8(s, replacement)` returns a copy with each **run** of invalid bytes replaced by the replacement string — pass `""` to delete them or `"\uFFFD"` to make the damage explicit. `bytes.ToValidUTF8` is the `[]byte` twin. Deciding between deleting and marking is worth a moment: deleting produces cleaner output but hides that data was lost; marking keeps the evidence but propagates a character your search index probably cannot match. ## Getting the invariant to stick One-off repairs do not hold. What holds is an invariant enforced at a boundary: - **Validate at ingress.** Reject or repair invalid UTF-8 the moment a display name enters the system, so every layer behind it can assume well-formedness. - **Truncate once, at a known place.** Scattered `[:n]` slices across handlers are how this bug returns. One helper, tested with a four-byte rune straddling the limit, is the fix that survives. - **Decide what the limit means.** If the column is `VARCHAR(32)` counting characters and your code counts bytes, the two limits disagree and one of them will surprise you. Write down which one is authoritative. - **Test the boundary.** A table test whose input places a four-byte rune at exactly the cut point, one byte before it, and one byte after it, catches every off-by-one in the backing-up loop. ## What an interviewer is checking Whether you connect a user-visible symptom to a byte-level cause, whether you reach for `utf8.ValidString` and a hexdump instead of speculating, and whether your fix is a boundary invariant rather than a patch at the one call site that reported the bug.

  • How many times can the backing-up loop decrement before it finds a rune boundary?
    At most three. UTF-8 encodes a rune in at most `utf8.UTFMax` bytes, which is four, so a cut point can be at most three continuation bytes deep inside a sequence. The loop is bounded and cheap; the `max > 0` guard only matters for pathological input that is entirely continuation bytes, which is already invalid UTF-8.
  • Should the repair delete the invalid bytes or replace them with U+FFFD?
    Deleting with `strings.ToValidUTF8(s, "")` gives clean output but silently loses evidence that data was damaged. Replacing with U+FFFD keeps the evidence visible but propagates a character that search and comparison will not match, and that the next system may itself flag as corrupt. I would delete in user-facing values and log the event, so the signal lives in the logs rather than in the data.
  • Would truncating on a rune boundary ever still look wrong to a user?
    Yes. A rune boundary is not a grapheme boundary: you can strip a combining accent off its base letter, or a skin-tone modifier or zero-width joiner off an emoji, leaving something well-formed but visually altered. That is a presentation decision rather than a corruption bug, and fixing it properly needs grapheme segmentation, which the standard library does not provide.

It is like cutting a sentence at the 32nd letter of a language written in ligatures: you can end up with half a symbol that no reader can interpret.

saying these in an interview costs you the question

  • Thinks string slice indices count characters
  • Blames the font or the database encoding first
  • Converts to []rune only to reintroduce a byte limit later
  • Repairs one row instead of enforcing an ingress invariant
  • Assumes every non-ASCII rune is exactly two bytes