skip to content

UTF-8 Decoding and Runes

Counting characters, validating input and classifying runes with the unicode and unicode/utf8 packages, which is where non-ASCII text stops behaving like len() suggests.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

questions

5

Which Go standard library call counts the runes in a string, and why can its result surprise a user?

level: juniorimportance: must knowfreq 72%

answer

  1. one package owns the counting
  2. unicode/utf8, not len
  3. it walks the bytes, O(n)
  4. a rune is a code point
  5. accents and emoji are several

basics

~20 s

utf8.RuneCountInString from the unicode/utf8 package returns the number of Unicode code points in a Go string. It can still surprise: accented letters, emoji and flags are often several code points that a user sees as one character.

solid answer

~40 s

`utf8.RuneCountInString(s)` walks the string decoding UTF-8 and returns the number of runes, i.e. Unicode code points; `utf8.RuneCount(b)` is the `[]byte` twin. Both are O(n), because UTF-8 is variable width and there is no stored count. The trap is that a rune is not a character a human sees. "e" followed by U+0301 COMBINING ACUTE ACCENT renders as one é but counts as two runes; a family emoji joined with zero-width joiners counts as five or more; a flag is two regional-indicator runes. So a rune count is the right limit for "how many code points" and the wrong one for "how many characters the user typed". Go's standard library has no grapheme-cluster segmenter, so if you truly need user-perceived characters you have to go outside it.

code

go · 3 lines
go
s := "e\u0301" // "e" + U+0301 COMBINING ACUTE ACCENT, renders as é
fmt.Println(len(s))                    // 3
fmt.Println(utf8.RuneCountInString(s)) // 2

go deeper

for a junior

Be ready to name unicode/utf8 and RuneCountInString on the spot, and to say plainly that it counts code points, not the characters a person sees.

for a middle

Explain why the count is a linear scan of variable-width UTF-8, and give a concrete case where runes and visible characters disagree, such as a letter written with a combining accent.

for a senior

Show the judgment: state what a length limit is protecting before picking bytes or runes, and say how you would word the validation error so it does not promise a character count you cannot deliver.

for a principal

Own the policy question. Decide once, for the whole platform, what a length limit means, where it is enforced, and whether the cost of a grapheme-accurate count justifies a dependency and the migration of every existing limit.

## What the call is `unicode/utf8` provides two counters: - `utf8.RuneCountInString(s string) int` - `utf8.RuneCount(p []byte) int` Both return the number of **runes** in the argument. A `rune` in Go is a Unicode **code point** — a number in the Unicode codespace, such as U+0041 LATIN CAPITAL LETTER A or U+1F600 GRINNING FACE. Because a Go string holds UTF-8 bytes and UTF-8 is a **variable-width** encoding (1 byte for ASCII, 2 for most European accented letters and Greek/Cyrillic, 3 for most CJK, 4 for emoji and other supplementary characters), there is no stored rune count to look up. Both functions scan the whole input, so they are O(n) in the number of bytes. If you are counting in a hot loop over the same string, count once and keep the number. Ill-formed bytes still count: each byte that cannot start or continue a valid sequence is counted as one rune, which is the same convention UTF-8 decoding uses everywhere in Go. ## Why the answer can surprise a user Unicode has three different notions of "how long is this text", and a rune count is the middle one: 1. **Bytes** — what storage and network limits care about. 2. **Code points (runes)** — what `utf8.RuneCountInString` returns. 3. **Grapheme clusters** — what a person calls a character: the smallest unit the cursor moves over. The three diverge constantly in real user-supplied text: - **Combining marks.** The text `"e\u0301"` is the letter e followed by U+0301 COMBINING ACUTE ACCENT. It renders as é, occupies 3 bytes and counts as **2 runes**. The very same visible letter can also arrive as the single precomposed rune U+00E9, which is 2 bytes and **1 rune**. Two visually identical names, two different counts. - **Emoji built by joining.** A family emoji is several people-emoji joined by U+200D ZERO WIDTH JOINER; a skin-toned emoji is a base emoji plus a U+1F3FB..U+1F3FF modifier. One picture, many runes. - **Flags.** A flag is two REGIONAL INDICATOR SYMBOL runes, so a single flag counts as 2. - **Scripts that stack.** Devanagari, Thai and Korean jamo routinely compose several code points into one visible cluster. So "name must be at most 20 characters" cannot be enforced honestly by a single number. Decide what the limit is actually protecting: - Protecting a **database column or a wire frame**? Limit **bytes** — but then truncate on a rune boundary, never mid-sequence. - Protecting a **rendering budget or an abuse surface**? Limit **runes**, and accept that a determined user can spend many runes on one glyph. - Promising the user a **character count that matches what they see**? The standard library will not give you that; grapheme segmentation lives outside it. ## Related calls in the same package - `utf8.ValidString(s) bool` — is every byte part of a well-formed sequence? - `utf8.RuneLen(r rune) int` — how many bytes this rune needs when encoded (−1 for a rune that cannot be encoded). - `utf8.UTFMax` — the constant 4, the maximum bytes any rune encodes to. - `utf8.DecodeRuneInString(s)` — one rune at a time, with its byte width. ## What an interviewer is checking They want to hear that you know a rune is a code point rather than "a character", that counting is a scan rather than a field read, and that you would ask what the limit is for before choosing between bytes and runes. Candidates who answer only "use `utf8.RuneCountInString`" have half the answer; candidates who add "and I would not call the result *characters* in the error message shown to the user" have all of it.

  • Why is utf8.RuneCountInString O(n) rather than a constant-time field read?
    A Go string is a pointer and a byte length; it stores no rune count. UTF-8 is variable width, so the only way to know how many runes are present is to decode from the start and step over each sequence. If you need the number repeatedly, compute it once and cache it, or keep a `[]rune` if you also need random access by index.
  • A product manager asks for a 20-character limit on display names. What do you ask back?
    Whether the limit protects storage or presentation. A byte limit protects the column and the wire but lets a user fit only five emoji; a rune limit is fairer but still lets one visible glyph cost five runes. I would enforce a rune limit for the user-facing rule, enforce a byte limit as a second, larger backstop for storage, and truncate only on a rune boundary.
  • Does utf8.RuneCountInString return an error for text containing invalid UTF-8?
    No. It has no error result. Each byte that is not part of a well-formed sequence is counted as one rune, matching how the rest of Go decodes UTF-8. If you need to know whether the input is well formed, that is a separate check with `utf8.ValidString`.

A rune count is like counting brush strokes rather than letters: an accented letter can be written with two strokes and still look like one letter on the page.

saying these in an interview costs you the question

  • Says a rune equals one visible character
  • Claims the count is stored on the string
  • Assumes every rune is one byte
  • Believes an emoji is always exactly one rune
  • Uses RuneCountInString as a storage-size limit
open as a page

In a Go signup validator, what do unicode.IsLetter, IsDigit and IsSpace actually accept?

level: middleimportance: should knowfreq 34%

basics

~20 s

They test Unicode categories over a whole rune, not ASCII. IsLetter accepts letters from every script, IsDigit accepts only decimal digits in category Nd including Arabic-Indic ones, and IsSpace accepts the Unicode white-space set including the no-break space.

open as a page

What does utf8.DecodeRuneInString return for invalid bytes, and how do you spot a genuine U+FFFD?

level: middleimportance: should knowfreq 42%

basics

~10 s

It returns utf8.RuneError with size 1 for invalid bytes, and size 0 for an empty string. A genuine U+FFFD decodes with size 3, so the size, not the rune, tells corruption apart.

open as a page

Why does truncating a Go display name with name[:32] leave a replacement box, and how do you fix it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Slicing a Go string cuts at a byte offset, so a multi-byte rune gets split in half and the stored bytes are no longer valid UTF-8. Back the cut up to a rune boundary and check utf8.ValidString before storing.

open as a page

How do you decide whether to depend on golang.org/x/text and normalise stored names to NFC?

level: principalimportance: nice to knowfreq 20%

basics

~20 s

Go's standard library validates UTF-8 but does not normalise, so normalisation means adding golang.org/x/text to go.mod. Decide it before user data exists, because normalising later rewrites stored bytes and can collide rows that were previously distinct.

open as a page