skip to content

String and []byte Representation

Strings are immutable UTF-8 byte views, which is why conversions to and from []byte copy, why s[i] gives you a byte, and why ranging gives you runes. Expect a question about counting characters in a non-ASCII string, or about removing that copy from a hot conversion path.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

questions

5

What does a Go string value hold at runtime, and what does assigning one to another variable copy?

level: juniorimportance: must knowfreq 62%

answer

  1. It is a small struct, not the text
  2. Two words wide
  3. A pointer and a length
  4. No capacity field, unlike a slice
  5. 16 bytes on a 64-bit build

basics

~20 s

A Go string value is a two-word header: a pointer to an array of bytes plus a length. Assigning or passing a string copies only that header (16 bytes on a 64-bit platform), never the text it points at.

solid answer

~40 s

At runtime a `string` is a small struct of two words: a pointer to the first byte of a byte array, and an `int` length. There is no capacity field, because a string is immutable and can never grow in place. So `b := a`, passing a string to a function, storing it in a map or sending it on a channel all copy 16 bytes on a 64-bit build, whatever the string's length — `unsafe.Sizeof(s)` reports 16 for both `""` and a 10 MB log. Slicing is just as cheap: `s[2:5]` builds a new header pointing into the same bytes and allocates nothing. What is *not* free is anything that has to touch the bytes — comparison, hashing as a map key, concatenation, and conversion to `[]byte` or `[]rune`.

code

go · 7 lines
go
s := "hello"
big := strings.Repeat("x", 10<<20)

fmt.Println(unsafe.Sizeof(s), unsafe.Sizeof(big)) // 16 16 on a 64-bit build

sub := big[100:200] // new header into the same bytes; no allocation
fmt.Println(len(sub)) // 100

go deeper

for a junior

Be ready to say a string value is a pointer plus a length, and that copying or passing one copies only that header. Knowing the number 16 on a 64-bit build is a nice touch.

for a middle

Explain why there is no capacity field, and split operations into header-only ones (assignment, slicing) and byte-touching ones (comparison, concatenation, conversion).

for a senior

Show where this decides real code: choosing slicing over conversion in hot paths, refusing *string parameters offered as an optimisation, and knowing that comparison and map hashing are the O(n) costs.

for a principal

Own the guidance that stops teams from micro-optimising the wrong thing: string copies are never the problem, conversions and concatenation are, and the standard should point at measurement rather than folklore.

## The value itself A Go `string` is not an object, not a class instance, and not a byte slice. At runtime a string value is a two-word header: - a pointer to the first byte of an array of bytes, and - an `int` holding how many bytes are in the string. On a 64-bit platform each word is 8 bytes, so every string value is exactly 16 bytes wide. `unsafe.Sizeof(s)` returns 16 whether `s` is the empty string, `"hi"`, or the contents of a gigabyte log file — `Sizeof` measures the header, not what it points at. The bytes themselves live wherever they were created: in the binary's read-only data section for a string literal, or on the heap for a string built at run time. ## No capacity, and why that is not an omission A slice value carries three words — pointer, length and capacity — because a slice can grow into spare room already reserved behind it. A string carries only two, because a string is immutable: there is no operation that writes another byte into it, so reserved spare room would be unusable. "Growing" a string means building a new one; `s + t` allocates a fresh array and copies both operands into it. Immutability is what makes the two-word design safe. Because no code can write through a string, the runtime is free to let many string values point into the same underlying bytes without anyone being able to observe interference. ## What copying a string costs Every place Go copies a value, it copies the string header and nothing else: - `b := a` - passing a string as a function argument or returning one - storing a string in a slice, a map value, or a struct field - sending a string on a channel - capturing a string in a closure All of these are 16-byte copies. This is why "pass a `*string` to avoid copying a long string" is bad advice in Go — the pointer you would pass is also 8 bytes, and the extra indirection costs a load. Pointers to strings exist for other reasons (distinguishing "absent" from "empty", or mutating a caller's variable), not for performance on large text. ## Slicing shares, conversion copies `s[2:5]` produces a new two-word header whose pointer is `s`'s pointer advanced by 2 and whose length is 3. No allocation, no byte copying — the substring and the original share the same array, which is safe precisely because neither can be written. The offsets are byte offsets, so it is on you to make sure they fall on boundaries that mean something. Conversions are the opposite. `[]byte(s)` and `string(b)` allocate a new array and copy the bytes across, because the result of the conversion is writable (or must stay unwritable) independently of the source. That is an O(n) operation and it shows up in profiles of code that converts in a loop. ## What is genuinely O(n) Cheap (header-only): assignment, argument passing, slicing, taking the length. Proportional to length: `==` comparison (though it can exit early on a length mismatch or a first-byte difference), hashing when used as a map key, concatenation, `[]byte(s)` and `[]rune(s)`, and copying into a `[]byte` with `copy`. ## Things people get wrong - **"A string is a `[]byte` under the hood."** It shares the pointer-and-length idea but has no capacity and no write path; the two are different types and converting between them copies. - **"Strings are null-terminated."** They are not. The length is explicit, so a Go string may contain interior zero bytes. This matters when handing text to C. - **"A string stores characters."** It stores bytes, conventionally UTF-8. The length is a byte count. - **"Passing a big string is expensive."** It is 16 bytes. ## How to check any of this `unsafe.Sizeof` on a string always yields the header size, and a benchmark that passes strings of wildly different lengths through the same function shows a flat cost. Both are quick ways to convince a reviewer who is still writing `*string` parameters out of habit.

  • Why does a string header have no capacity field when a slice header does?
    Capacity exists so `append` can write into reserved room behind the length. A string can never be written to, so reserved room could never be used. Building a longer string always allocates a new array and copies, which is why concatenating in a loop is expensive and `strings.Builder` exists.
  • Does `s[2:5]` on a string allocate?
    No. It builds a new two-word header whose pointer is offset by 2 and whose length is 3, sharing the original bytes. That sharing is safe because neither value can be written. Converting the result with `[]byte(...)` is what allocates.
  • Is it ever worth taking a `*string` parameter to avoid copying?
    Not for size. The header is 16 bytes and a pointer is 8, so you trade a tiny copy for an indirection. Use `*string` only when you need to distinguish absent from empty, or to write back to the caller's variable.

A string value is a library call slip: it names a shelf position and how many pages to read. Handing someone the slip is cheap; photocopying the pages is not.

saying these in an interview costs you the question

  • Says a string is a []byte with three words including capacity
  • Thinks passing a long string copies all of its bytes
  • Believes s[2:5] allocates a new byte array
  • Claims Go strings are null-terminated like C strings
  • Says a string stores UTF-16 code units or fixed-width characters
open as a page

Why does converting between string and []byte in Go allocate and copy, and when does the compiler skip it?

level: middleimportance: must knowfreq 58%

basics

~20 s

Strings are immutable, so []byte(s) and string(b) each allocate a new array and copy the bytes, keeping the two independent. The compiler elides that copy only in read-only patterns such as m[string(b)] and string(b) == "literal".

open as a page

How does strings.Builder produce its final string without copying the bytes it accumulated?

level: middleimportance: should knowfreq 44%

basics

~20 s

strings.Builder appends into an internal byte slice and its String method builds a string header over that same array with unsafe.String, so there is no final copy. It is safe because a Builder only ever appends, never rewrites bytes it already handed out.

open as a page

A Go log pager truncates each line with s[:80] and users see mangled characters — what is happening and how do you cut safely?

level: seniorimportance: should knowfreq 38%

basics

~20 s

s[:80] cuts at byte 80, which can land inside a multi-byte UTF-8 sequence and leave a half-encoded rune that the terminal draws as a replacement character. Cut at a rune boundary instead, then worry about display width separately.

open as a page

What conditions would make you allow unsafe.String zero-copy conversions in your team's Go codebase?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Allow it only where a profile shows the conversion copy actually costs, where the bytes provably can never be written again, and where the trick stays inside one package behind an ordinary safe API. Everywhere else the standard says no.

open as a page