Why does ranging over the Go string "héllo" yield indices 0, 1, 3, 4, 5?
answer
- the first value is a position, not a counter
- positions come from the stored bytes
- a two-byte letter skips an offset
- the loop decodes, so it can jump
basics
~20 sBecause the loop hands back the byte offset where each character starts, not a sequential counter. The accented letter occupies bytes 1 and 2, so the next letter starts at 3. The second loop value is the decoded character.
solid answer
~40 sRanging over a string decodes UTF-8 as it goes. Each iteration produces two values: the **byte offset** at which the current character's encoding begins, and the decoded code point as a `rune`. In `"héllo"` the `é` is encoded in two bytes at offsets 1 and 2, so the loop skips 2 entirely and resumes at 3 — the offsets step by the encoded width of each character, not by one. That is what makes the index useful: `s[i:]` from any yielded offset starts on a character boundary. Contrast it with a hand-written `for i := 0; i < len(s); i++`, which walks every byte and hands you half of a multi-byte character. If a byte is not valid UTF-8, the loop yields U+FFFD, the replacement character, and advances a single byte.
code
go · 8 linesfor i, r := range "héllo" {
fmt.Printf("%d %c\n", i, r)
}
// 0 h
// 1 é
// 3 l
// 4 l
// 5 ogo deeper
Remember that the loop gives you a position and a character, and that the position comes from the stored bytes. If a letter takes two bytes, the next position jumps by two.
Explain the decoding step: each iteration reads one UTF-8 sequence, yields the code point, and advances by its width. Be able to name what one, two and zero loop variables each give you.
Show that the yielded offset is boundary information you can slice with, and that invalid bytes degrade to the replacement character instead of failing. That is the basis of every safe truncation routine you will write.
Own the convention: decide whether your codebase treats text as bytes or as characters at each layer, and make that explicit rather than leaving each author to pick a loop form and hope the input stays ASCII.
### What the loop actually does When the operand of `range` is a string, the loop does not walk bytes and it does not walk a precomputed array of characters. It **decodes UTF-8 on the fly**. On each iteration it looks at the bytes starting at the current position, works out how many of them form one code point, produces that code point, and advances the position by exactly that many bytes. That gives the two loop values: 1. **First value — the byte offset** at which the current character's encoding starts, counted in the same units `len` uses. 2. **Second value — the decoded code point**, of type `rune`. ### Walking through "héllo" The stored bytes are: `h` at 0; `é` at 1 and 2 (two bytes); `l` at 3; `l` at 4; `o` at 5. Six bytes, five characters. So the loop yields the pairs `(0, 'h')`, `(1, 'é')`, `(3, 'l')`, `(4, 'l')`, `(5, 'o')`. Offset 2 never appears, because it is not the *start* of anything — it is the continuation byte of the accented letter. The gaps are the point. The offsets are not a counter; they are positions in the underlying bytes. A five-character string yields five iterations, but the largest offset can be as high as `len(s)-1` for a four-byte character. ### The three forms of the loop - `for i, r := range s` — offset and decoded code point. - `for i := range s` — offsets only. This is the cheap way to enumerate character boundaries. - `for range s` — neither; useful only to count characters. With one variable you get the **index**, not the value — the same rule as ranging over a slice, and a frequent slip when someone means to iterate characters. ### Why the offset is the useful half Because the offset always lands on a character boundary, it is exactly what you need to slice safely. `s[i:]` is the remainder of the text from that character onwards, and `s[:i]` is everything before it — both guaranteed to be whole characters, unlike an arbitrary numeric cut. Scanning for the position of the last character that fits inside a byte budget is the standard use. ### Contrast with an index loop ``` for i := 0; i < len(s); i++ { ... s[i] ... } ``` This walks **bytes**. Each `s[i]` is a single byte value, and for a multi-byte character you see its pieces one at a time, none of which is a valid character on its own. Both loops are legitimate — byte scanning is right when you are looking for an ASCII delimiter such as a comma or a newline, because UTF-8 guarantees no ASCII byte appears inside a multi-byte sequence — but only the range form gives you characters. ### Invalid bytes Nothing guarantees a string is valid UTF-8. When the decode fails, the loop does not panic and does not stop: it yields **U+FFFD**, the Unicode replacement character, and advances exactly one byte, then tries again. That is why a corrupted string produces a run of replacement characters rather than an error. If you need to know whether the text was valid, you have to check explicitly rather than infer it from the loop. One consequence worth remembering: because U+FFFD is also a perfectly legal character that a string can genuinely contain, seeing one in the output does not by itself prove the input was corrupt — though in practice it almost always does. ### Printing what you get The second loop value is a numeric code point. `fmt.Println(r)` prints a number; `fmt.Printf("%c", r)` prints the character and `%q` prints it quoted. Someone debugging this loop for the first time usually prints the number by accident and concludes the decoding is broken. ### The mental model Read `for i, r := range s` as: *give me each character together with where it starts*. The character is for logic; the offset is for slicing. Once both halves have names, the jumpy indices stop looking like a bug and start looking like the boundary information they are.
- What does `for i := range s` give you when s is a string — the character or the offset?The offset. With a single loop variable you get the first value only, which is the byte position where each character's encoding starts — the same rule as ranging a slice, where one variable gives the index. If you want the characters you must write both variables, `for _, r := range s`.
- How does the loop behave on a string containing bytes that are not valid UTF-8?It neither panics nor stops. For each byte it cannot decode it yields U+FFFD, the replacement character, and advances exactly one byte before trying again. So invalid input turns into a run of replacement characters, and if you actually need to know whether the text was valid you have to check for that separately.
- Why is the yielded offset safer to slice with than an arbitrary number?Because every offset the loop produces is the first byte of a character, so `s[:i]` and `s[i:]` always cut between characters. An arbitrary numeric cut can land inside a multi-byte encoding and leave a partial sequence, which is where truncation bugs come from.
saying these in an interview costs you the question
- Expects the indices to be 0, 1, 2, 3, 4 always
- Thinks the first loop value counts characters
- Believes `for i := range s` yields characters
- Assumes invalid bytes make the loop panic
- Says the second value is a one-character string