A signup form truncates display names with name[:20] and some stored names now end in a replacement character. What went wrong?
answer
- the slice bound is in the wrong unit
- nothing validates the result
- half a character was kept
- the renderer only reports the damage
- cut where a character starts
basics
~20 sSlicing a Go string cuts bytes, not characters, so name[:20] can land inside a multi-byte UTF-8 sequence. The partial bytes left behind are still a legal Go string, and only the layer that decodes them shows a replacement character.
solid answer
~50 s`name[:20]` is a **byte** slice expression, not a character one. Any name whose twentieth byte falls inside a multi-byte UTF-8 sequence gets cut in half, leaving a leading byte with no continuation. Go does not validate strings, so nothing fails at the cut: the value is stored happily and only surfaces when something decodes it, at which point the orphaned bytes render as U+FFFD. Confirm it by dumping the stored bytes in hex — a trailing `c3` or `e6` with nothing after it is the signature. The fix is to back up to the last character boundary at or before the limit, which you can find by ranging the string, since every offset it yields starts a character. First, though, settle what the limit means: twenty bytes is a storage constraint, twenty characters is a product one, and they are not interchangeable.
code
go · 3 linesname := "José" // 5 bytes: 4a 6f 73 c3 a9
short := name[:4] // a byte bound, not a character bound
fmt.Printf("%x\n", short) // prints: 4a6f73c3go deeper
Remember that slicing a string uses byte positions, so cutting at a fixed number can split a multi-byte character. Know that the replacement character in the output is the symptom, not the cause.
Explain why nothing failed at the cut: strings are unvalidated bytes, so the partial sequence stores fine and only a decoder notices. Be able to describe a boundary-aware truncation using the offsets a range loop yields.
Diagnose from the stored bytes rather than the rendered output, identify the affected rows for a backfill, and argue which unit the limit should be in. Be ready to say that silent truncation of an identity field is itself the wrong behaviour.
Own the rule across the system: one definition of the limit and its unit, enforced at the edge, with storage sized to match. An interviewer expects you to weigh reject-versus-truncate as a product decision, not just patch the slice expression.
### What the code does `name[:20]` takes the first twenty **bytes**. A Go string is a byte sequence, and every slice expression on it is positioned in bytes. UTF-8 encodes a character in one to four bytes, so for any name that is not pure ASCII, byte 20 has a good chance of being in the middle of a character rather than at its start. Cut there and you keep the first byte or two of a multi-byte sequence and throw away the rest. The result is a string containing an incomplete encoding. ### Why nothing failed at the cut Go never validates a string's contents. There is no encoding tag, no check on assignment, and no check on the slice expression. A string is whatever bytes you put in it, valid UTF-8 or not. So the truncation succeeds, the value goes into the database, and the system stays quiet. The failure surfaces later and elsewhere: at whatever layer decodes the bytes for display. A decoder that meets a leading byte with no continuation emits U+FFFD, the Unicode replacement character, which is what the user sees as a small diamond or box. The bug therefore appears to be in the renderer, several layers away from the line that caused it. ### Confirming it Dump the stored bytes rather than printing the string — printing it will just show you the replacement character again. `fmt.Printf("%x", s)` gives a plain hex dump and `%q` shows a quoted form with invalid bytes escaped as `\\xc3`-style values. A trailing byte in the `c2`-`f4` range with nothing following it is the leading byte of a sequence whose tail was cut off, and that is the signature you are looking for. Comparing the stored byte length against the length of what the user typed usually confirms the truncation happened at exactly the limit. ### Fixing it The correct byte-limited truncation cuts at the last **character boundary** at or before the limit. Ranging the string gives you those boundaries directly: every offset it yields is the first byte of a character, so the largest such offset that fits the budget is where to cut. That drops at most one whole character and always leaves valid text. The other repair is to decide the limit is really about characters, convert with `[]rune(s)` and keep the first N code points. That is valid text too, but be clear about what it costs: a rune limit is not a byte limit. Twenty code points can occupy up to eighty bytes, so if the underlying column is sized in bytes you have simply moved the failure from corruption to a length error at insert. ### Deciding which limit you actually have This is the part that matters more than the code: - **A storage limit** is in bytes (or in the database's own notion of character, which may differ again). Enforce it with a boundary-aware byte truncation. - **A product limit** — "names may be at most 20 characters" — is in characters, and the honest implementation counts decoded code points. - **A display limit** is neither, because glyph width has nothing to do with either count: CJK characters occupy two columns, and combining marks occupy none. Pick one, state it in the form's validation message, and enforce the same one everywhere. Half of these bugs come from validation counting one unit and storage counting another. ### Truncate or reject Silently truncating user-entered identity data is itself a questionable behaviour: the user gets an account under a name they did not choose, and no error tells them why. For a signup form, rejecting with a clear message is usually better than trimming. Truncation belongs to display and to log lines, where losing the tail is acceptable. ### One more caveat Even a character-boundary cut can look wrong to a reader. A code point is not a user-perceived glyph: a letter followed by a combining accent is two code points, and cutting between them leaves a stray floating accent; emoji built from several joined code points can lose a modifier and change appearance entirely. If truncation lands in front of users, cutting on a whitespace boundary and appending an ellipsis is safer than cutting mid-word at all. ### The habit Whenever you see a numeric index into a string — a slice bound, an offset, a length check — ask what unit it is in and whether the text is guaranteed ASCII. If it is not, the index has to come from decoding rather than from arithmetic.
- Why did the truncation not fail at the point it happened?Because Go never validates string contents. A string is just bytes with no encoding tag, so slicing mid-character produces a perfectly legal value that stores and compares fine. Only a decoder — the template, the terminal, the browser — notices the incomplete sequence, and it reports it as a replacement character rather than an error.
- Would converting to []rune and keeping the first 20 elements fix it?It produces valid text, but it answers a different question. Twenty code points can be up to eighty bytes, so if the twenty was a storage limit you have traded corruption for an oversize value. Use a rune limit when the rule is genuinely about characters, and a boundary-aware byte cut when the rule is about bytes.
- How would you confirm the diagnosis on data already in the database?Dump the bytes rather than the rendered string — `%x` for a hex view or `%q` for a quoted form that escapes invalid bytes. Rows whose stored value ends in an unpaired leading byte, and whose byte length is exactly the truncation limit, are the affected set. That also gives you the query for the backfill.
- Is silently truncating a signup name the right behaviour at all?Usually not. The user ends up with a name they did not choose and no explanation. For identity fields, validate and reject with a message that states the limit and its unit. Keep truncation for places where losing the tail is acceptable, such as log lines and list previews.
saying these in an interview costs you the question
- Thinks a Go slice expression cuts characters
- Expects the truncation itself to panic or error
- Assumes a 20-rune cut also satisfies a 20-byte column
- Blames the renderer for producing the replacement character
- Validates in characters while storing in bytes