skip to content

In Go, why can keeping strings.TrimSpace's result hold a whole 10 MB document in memory?

level: seniorimportance: nice to knowfreq 30%

answer

  1. what does slicing a string cost
  2. the collector frees allocations, not bytes
  3. a short header over a long array
  4. one call in strings makes a real copy
  5. strings.Clone at the retention boundary

basics

~20 s

Slicing a Go string does not copy its bytes, and strings.TrimSpace returns a slice of its input. A short trimmed field therefore points into the original document's backing array and keeps all of it reachable. strings.Clone makes an independent copy.

solid answer

~40 s

A Go string is a header — a pointer and a length — over immutable bytes, and slicing produces a new header over the same bytes. `strings.TrimSpace` only moves the two ends inward, so its result shares the input's backing array. If a normaliser reads a 10 MB document, trims an 80-byte title out of it and stores that title in a long-lived index, the garbage collector cannot free the document: reachability is per allocation, not per byte, so the whole array stays alive behind that one small header. The fix is `strings.Clone(s)`, which returns a fresh copy with no shared storage, applied at the boundary where a short value is retained beyond the document's lifetime. The same sharing is what makes `TrimSpace` allocation-free, so clone deliberately at retention points rather than defensively everywhere.

code

go · 8 lines
go
// doc holds a 10 MB document that is dead after this call
title := strings.TrimSpace(doc[titleStart:titleEnd])

// title shares doc's bytes: storing it keeps all 10 MB reachable
// index[id] = title

// a copy with its own storage releases the document
index[id] = strings.Clone(title)

go deeper

for a junior

Remember that a Go string is immutable and that slicing one is cheap because it does not copy. Every function in the strings package returns a value; none of them modify the string you passed in.

for a middle

Be able to explain the string header — pointer plus length — and why a substring shares the original bytes, then name strings.Clone as the call that breaks the sharing deliberately.

for a senior

Show that you would diagnose this from a heap profile before changing code, and that you place the clone narrowly at the point where a small value outlives a large buffer rather than defensively across the pipeline.

for a principal

Decide where in the pipeline documents stop being shared: whether the reading layer hands out independent values by contract, or whether every consumer is responsible for copying what it retains. Leaving that unstated is what makes the leak recur after each rewrite.

## What a Go string actually is A `string` value is a two-word header: a pointer to bytes and a length. The bytes are immutable, which is precisely what lets the runtime share them freely. `s[a:b]` therefore allocates nothing — it produces a new header pointing `a` bytes into the same array with length `b-a`. `strings.TrimSpace(s)` is defined entirely in terms of moving those two ends inward past white space, so its result is a substring of `s` sharing the same storage. So are `strings.TrimPrefix`, `strings.TrimSuffix`, `strings.Trim` and the rest of the trim family. None of them copy. ## Why that becomes a memory problem Go's garbage collector reclaims **whole allocations**. A byte array is either reachable or it is not; there is no way to free the unreferenced middle of one. So a single 80-byte header pointing into a 10 MB array keeps all 10 MB alive for as long as that header is reachable. The shape that produces it is ordinary and looks harmless: a normaliser reads a document, trims a few fields out of it, writes those fields into a long-lived structure — a cache, an index, a batch being accumulated — and drops the document. Every retained field is an anchor. Memory grows in proportion to the number of *documents processed*, not the size of what you kept, and it looks exactly like a leak: the live heap climbs, a heap profile shows a large share of `inuse_space` in the byte slices produced by whatever read the input, and nothing in the retaining code allocates enough to explain it. The tell in a heap profile is that the memory is attributed to the reading path, not to the code that holds the data — that mismatch between where it was allocated and what is keeping it alive is the signature of retention rather than over-allocation. ## The fix `strings.Clone(s) string` returns a copy of `s` with fresh backing storage, so the copy shares nothing with the original. Apply it at the boundary where a short value outlives the large one it came from: ```go title := strings.TrimSpace(doc[start:end]) index[id] = strings.Clone(title) ``` Any operation that genuinely builds a new string has the same effect as a side-benefit — `strings.ToLower` on text that actually contains uppercase, or a `Replacer` that matched something, both produce independent storage. That is worth knowing precisely because it makes the bug *intermittent*: a pipeline that lowercases every key never sees this, until someone adds a fast path that skips the transformation when the input is already lowercase. ## The flip side — do not clone everywhere The sharing is a feature, and it is why this package is cheap in the common case: - `strings.TrimSpace` never allocates; it returns a substring. - `strings.ToLower` returns **the input string itself** when no character needs changing, which is the common case for already-normalised text. - `strings.Replace` returns the input unchanged when there is no match, so a replacement pass over text that does not contain the pattern is a scan with no copy. Cloning by reflex throws all of that away and adds an allocation per field to a hot path. The discipline is narrow: clone when a **small** value derived from a **large, otherwise-dead** buffer is going to be **retained**. Values that are compared, written out and discarded within the call need nothing. ## Confirming it rather than guessing Two cheap checks settle it. A benchmark run with `-benchmem` reports bytes and allocations per operation, which tells you whether a normalisation step is copying at all — a step reporting 0 allocs/op is sharing, and sharing is what creates retention. And a heap profile taken while live heap is climbing shows whether the surviving bytes belong to the read buffers rather than to the structure that appears to hold the data. The same golden-file corpus you use to check the normaliser's output is a good input for both, because it is the text the service actually sees rather than a synthetic string.

  • Which normalisation calls in the strings package can return the input string itself?
    `TrimSpace` and the whole trim family always return a substring of the input. `ToLower` and `ToUpper` return the input unchanged when no rune needs changing, and `Replace`/`ReplaceAll` return it unchanged when there is no match. That is why an already-normalised corpus costs almost nothing to run through the pipeline — and why the retention hazard is intermittent.
  • How would you confirm this is retention rather than simply allocating too much?
    Take a heap profile while the live heap is climbing and look at what the surviving bytes were allocated by. Retention shows up as a large `inuse_space` share attributed to the code that read the input, while the code that appears to hold the data allocates almost nothing. Over-allocation instead concentrates in the transforming code itself.
  • Does the same reasoning apply when the document arrives as a []byte?
    Yes, and more sharply, because a byte slice also carries capacity: a sub-slice can reach bytes beyond its own length, so it retains at least as much. The `bytes` package mirrors these functions with the same sharing behaviour, and the copy at the retention boundary is an explicit `append` into a fresh slice or a conversion to `string`.

Tearing a headline out of a newspaper is a copy; circling it in pen is not — and if you keep the circle, you keep the whole paper.

saying these in an interview costs you the question

  • Assumes every function in strings allocates a fresh copy
  • Thinks slicing a string copies the bytes it covers
  • Expects the collector to free the unreferenced tail of an array
  • Calls strings.Clone on every value out of habit
  • Changes the normaliser for speed without a benchmark or a profile
  • Blames the reader for over-allocating when the problem is retention