skip to content

In Go, how does strings.Trim's cutset differ from strings.TrimPrefix's prefix?

level: juniorimportance: must knowfreq 72%

answer

  1. one argument is a set
  2. order in the argument never matters
  3. characters, not a word
  4. TrimPrefix strips one literal occurrence

basics

~20 s

strings.Trim takes a cutset: a set of individual characters, stripped repeatedly from both ends in any order. strings.TrimPrefix takes a literal string and removes it once from the front, returning the input unchanged when that prefix is absent.

solid answer

~40 s

`strings.Trim(s, cutset)` reads its second argument as a *set of characters*. It decodes the cutset into runes and strips any of them from both ends, repeatedly, until it hits a rune that is not in the set, so `"abc"` and `"cba"` behave identically. `strings.TrimPrefix(s, prefix)` reads its second argument as a literal string: it removes exactly one leading occurrence and returns `s` untouched if the prefix is not there. That is why `strings.TrimLeft("user_users_report", "user_")` returns `"port"` and not `"users_report"` — it keeps eating characters as long as each one is in {u, s, e, r, _}. The rule I use in review: if the second argument is a word, you want `TrimPrefix` or `TrimSuffix`; a cutset is for punctuation and whitespace classes.

code

go · 9 lines
go
s := "user_users_report"

// cutset {u, s, e, r, _} stops only at the 'p'
fmt.Println(strings.TrimLeft(s, "user_"))
// port

// literal prefix, removed once
fmt.Println(strings.TrimPrefix(s, "user_"))
// users_report

go deeper

for a junior

Be ready to say plainly that Trim, TrimLeft and TrimRight take a set of characters while TrimPrefix and TrimSuffix take a literal word, and that a missing prefix simply returns the input unchanged.

for a middle

Explain the mechanics: the cutset is decoded into runes, membership is tested repeatedly from the end inwards, and TrimSpace uses unicode.IsSpace rather than the ASCII space alone. Name TrimFunc as the predicate escape hatch.

for a senior

Show that you recognise this as a data-dependent bug: the wrong call passes on most inputs. Talk about how you find it in an existing codebase and how you prove the fix over a captured corpus rather than over three hand-picked examples.

for a principal

Own the convention. Decide in review that cutsets are only for character classes and that literals go through TrimPrefix/TrimSuffix, and require normalisation code to carry a golden-file test so any future change to it lands as a visible diff.

## Two different arguments that look alike Go's `strings` package has two families that both remove characters from the ends of a string, and they take arguments that look identical but mean different things. **The cutset family** — `strings.Trim(s, cutset)`, `strings.TrimLeft(s, cutset)`, `strings.TrimRight(s, cutset)`. The `cutset` is a **set of Unicode code points**, written as a string only because that is a convenient way to spell a set. The function decodes the cutset into runes and then removes leading (and/or trailing) runes for as long as each one is a member of that set. Consequences that follow directly: - Order inside the cutset is meaningless. `Trim(s, "\"'")` and `Trim(s, "'\"")` do the same thing. - Duplicates are meaningless. `Trim(s, "xx")` equals `Trim(s, "x")`. - Removal repeats. `strings.Trim("xxxnamexxx", "x")` is `"name"`, not `"xxnamexx"`. - It can consume the whole string. `strings.Trim("aaa", "a")` is `""`. **The literal family** — `strings.TrimPrefix(s, prefix)` and `strings.TrimSuffix(s, suffix)`. Here the second argument is an ordinary substring compared in order. Exactly one occurrence is removed, and if the string does not start (or end) with it, `s` comes back completely unchanged. There is no error and no boolean — a no-op is the documented behaviour. ## The failure this distinction causes The classic bug appears in normalisation code — the layer that cleans user-entered text before it becomes a search key or an identifier. Someone writes: ```go id := strings.TrimLeft(raw, "user_") ``` intending "drop the `user_` prefix". With `raw = "user_users_report"` the cutset is {u, s, e, r, _}, and every character of `user_users_re` is in that set, so trimming stops only at the `p`. The result is `"port"`. Worse, on many inputs the wrong call *appears* to work — `strings.TrimLeft("user_alpha", "user_")` gives `"alpha"` because the first character after the prefix happens not to be in the set. The bug is data-dependent, which is why it survives code review and shows up as mangled or empty identifiers in production. The same trap is why `strings.TrimLeft(url, "https://")` is always wrong even when it looks right: the set is {h, t, p, s, :, /}, so a host beginning with any of those letters loses them. ## strings.TrimSpace and the predicate variants `strings.TrimSpace(s)` needs no argument: it removes leading and trailing white space as defined by `unicode.IsSpace`, which covers the ASCII space, tab, newline, carriage return, vertical tab and form feed, and also Unicode spaces such as the non-breaking space U+00A0. That makes it strictly stronger than `strings.Trim(s, " ")`, which removes only the ASCII space character — a real difference for text pasted out of a word processor or a web page. When the set you want is a class rather than an enumeration, use the predicate variants: `strings.TrimFunc(s, f)`, `strings.TrimLeftFunc`, `strings.TrimRightFunc`, each taking a `func(rune) bool`. Trimming everything that is not a digit, for example, is a `TrimFunc` job, not a cutset job. Every one of these has a byte-slice mirror in the `bytes` package (`bytes.Trim`, `bytes.TrimPrefix`, `bytes.TrimSpace`, …) with identical semantics, for when the input arrives as a `[]byte` from a reader. ## How to keep it out of the codebase Three habits cover it. First, a naming rule: the second argument of `Trim`/`TrimLeft`/`TrimRight` should never be a word a human would read aloud; if it is, the call is almost certainly meant to be `TrimPrefix` or `TrimSuffix`. Second, when a prefix may legitimately be missing and you need to know which happened, the trim functions are the wrong tool entirely — they cannot tell you. Third, put the normaliser under a golden-file test: run the current corpus of real user-entered text through it and check the before/after diff into the repository. Then a change to a trim call shows up as a reviewable diff of thousands of lines instead of a support ticket about missing search results. All of these functions return a new string value; none of them mutate the input, because Go strings are immutable.

  • What does strings.TrimSuffix("report.txt.txt", ".txt") return, and why?
    It returns `"report.txt"`. `TrimSuffix` removes a single trailing occurrence of the literal suffix, not as many as it can find. If you genuinely want every repeated occurrence stripped, loop until the result stops changing — a cutset of `".txt"` would instead remove any mix of `.`, `t` and `x` from the end and mangle a name like `"exit"`.
  • How do you trim by a rule rather than by a fixed set of characters?
    Use the predicate variants: `strings.TrimFunc(s, f)`, `strings.TrimLeftFunc` and `strings.TrimRightFunc`, each taking a `func(rune) bool`. They strip runes from the ends while the function returns true, so "strip everything that is not a letter or digit" becomes one predicate instead of an impossible-to-enumerate cutset.
  • You inherit a normaliser that uses TrimLeft on user-entered titles. How do you prove it is mangling data before you change anything?
    Capture a real corpus of the text that actually flows through it, run it through the current normaliser, and check the output in as a golden file. Then apply the fix and read the diff: the entries where the cutset over-trimmed show up as lines that gain characters back. That turns a data-dependent bug into a reviewable artefact for whoever is onboarding next.

A cutset is a list of characters the doorman will turn away, checked one at a time; a prefix is a password checked as a whole phrase.

saying these in an interview costs you the question

  • Thinks Trim's second argument is a substring removed as a unit
  • Uses TrimLeft to strip a URL scheme or a field prefix
  • Believes the order of characters in a cutset matters
  • Expects TrimPrefix to strip every repeated occurrence
  • Assumes strings.Trim(s, " ") is the same as strings.TrimSpace(s)