What does len("héllo") return in Go, and why isn't it 5?
answer
- count what is stored, not what is seen
- Go source files are UTF-8
- one letter is not always one byte
- range decodes, len does not
basics
~10 sIt returns 6. Go's len on a string counts bytes, not characters, and the letter é takes two bytes in UTF-8 while each ASCII letter takes one. Counting characters means decoding the bytes instead.
solid answer
~40 s`len("héllo")` is 6. A Go string is an immutable sequence of bytes, and `len` reports how many bytes it holds; Go source is UTF-8, so `é` is encoded as two bytes and the four ASCII letters take one each. Indexing follows the same rule: `s[1]` gives the first byte of `é`, not the letter. To count characters you have to decode, and `for range s` does that — it steps one character at a time, so counting its iterations gives 5. The byte-oriented `len` is deliberate: length, indexing and slicing stay direct reads of stored bytes. The practical consequence is that any code treating `len(s)` as a character count is wrong the moment a non-ASCII name arrives.
code
go · 8 liness := "héllo"
fmt.Println(len(s)) // prints: 6
n := 0
for range s { // one iteration per decoded character
n++
}
fmt.Println(n) // prints: 5go deeper
Be ready to say the number and the reason in one breath: 6, because len counts bytes and the accented letter is encoded in two of them. Know that ASCII text is the only case where byte count and character count agree.
Explain the mechanics: source is UTF-8, one code point takes one to four bytes, and indexing and slicing use the same byte positions len reports. Show the two honest ways to count characters and say what each costs.
Show where the byte/character gap causes production bugs — length validation, truncation, column widths — and say which unit each constraint really wants. Note that even a rune count is not a glyph count for combining marks and emoji.
Frame it as a contract question: whichever unit a limit is expressed in must be the same one storage, validation and the UI agree on. An interviewer expects you to argue for defining the limit once, at the edge, rather than letting each layer count differently.
### The answer `len("héllo")` evaluates to **6**, not 5. ### What a Go string is A Go string value is an immutable sequence of **bytes**. It is not a sequence of characters, and it carries no encoding tag. By convention — and by rule for string literals written in source files, which Go requires to be UTF-8 — those bytes hold UTF-8 text, but the language itself only guarantees you a byte sequence. Everything the language does to a string is therefore defined on bytes: - `len(s)` is the number of bytes. - `s[i]` is the byte at position `i`, as a numeric value. - `s[a:b]` is a substring taken at byte positions. - `==`, `<` and map lookups compare the bytes. ### Why 6 UTF-8 encodes each Unicode code point in one to four bytes. Plain ASCII letters and digits take one byte. Accented Latin letters, Greek and Cyrillic take two. Most CJK characters take three. Many emoji take four. So `"héllo"` is `h`(1) + `é`(2) + `l`(1) + `l`(1) + `o`(1) = 6 bytes for 5 characters. `"日本語"` is 9 bytes for 3 characters. Pure ASCII is the only case where the two numbers agree — which is exactly why this bug survives testing and shows up the first time a real user signs up. ### If you are porting from another language This is the single most common surprise for someone arriving from a language whose string is a sequence of fixed-width code units. There, `length` counts units of the internal representation and the count is stable for European text; in Go the count is bytes, and it changes with the script the text is written in. Go took the other tradeoff on purpose: one representation everywhere, no hidden re-encoding when you write to a file or a socket, and length and indexing that are direct reads of what is stored. ### Counting characters instead If you want the number of characters, you must decode the UTF-8: - `for range s` walks one decoded character per iteration, so counting the iterations gives the character count without allocating. - `[]rune(s)` decodes the whole string into a freshly allocated slice of code points; its `len` is the character count, at the cost of a copy. Be precise about what "character" means before you compute it. A `rune` is one Unicode code point, which is still not always what a user perceives as one letter. `é` can be written as a single code point, or as `e` followed by a combining acute accent — two runes, three bytes, one visible glyph. A flag or a family emoji is several code points joined together. So there are three plausible answers to "how long is this string": bytes, runes, and user-perceived glyphs, and only the first two are cheap. ### Where the confusion bites in real code - A form rule that says "at most 20 characters" implemented as `len(name) > 20`, which quietly rejects a shorter name written in a non-Latin script. - Truncation with `s[:20]`, which can cut a multi-byte character in half. - A database column sized in characters while the validation counts bytes, or the reverse. - Padding or column alignment computed from `len`, which misaligns for any accented text. ### What is not affected Comparison, equality, use as a map key, concatenation and `switch` on a string are all byte-wise and stay correct whatever the script, because two strings that encode the same characters the same way have the same bytes. The byte orientation only becomes visible when you ask a question about *positions* or *counts*. ### The habit to build Read `len(s)` as "how many bytes am I storing" and never as "how long does this look". When a limit, an offset or a width matters to a human, decode first; when it matters to storage or a wire protocol, bytes are the right unit and `len` already gives you it.
- So how do you count the characters a user actually typed?Decode. `for range s` yields one decoded code point per iteration, so counting iterations gives the count with no allocation; `len([]rune(s))` gives the same number but allocates a copy of the whole string. Be aware that a code point is still not always one visible glyph — a letter plus a combining accent, or an emoji built from several joined code points, counts as more than one rune.
- Is a Go string always valid UTF-8?No. String literals in source are UTF-8 because the source file must be, but a string is just bytes and can hold anything — a chunk read from a file, a slice cut mid-character, a byte slice converted with `string(b)`. Nothing validates on assignment. The invalid bytes only become visible when something decodes them, at which point they surface as the replacement character U+FFFD.
- Does len report something different for []byte(s) than for s?No — the same number. Converting a string to a byte slice copies the bytes without changing them, so the lengths match. `len([]rune(s))` is the one that differs, because that conversion decodes to code points: for "héllo" you get 6, 6 and 5 respectively.
len is a scale, not a headcount: it tells you how much text weighs in bytes, and some letters simply weigh more than others.
saying these in an interview costs you the question
- Says len returns the number of characters
- Assumes every character occupies one byte
- Thinks s[i] yields a character rather than a byte
- Believes Go stores strings as UTF-16 internally
- Claims Go strings are fixed-width arrays of runes