skip to content

Why is strings.Cut safer than strings.Index for parsing a "key=value" header line?

level: middleimportance: should knowfreq 55%

answer

  1. one call answers two questions
  2. the sentinel is the whole problem
  3. minus one plus one is zero
  4. a found bool instead of an offset

basics

~20 s

strings.Cut returns the text before and after the separator plus a found boolean, so a missing separator is impossible to ignore. strings.Index returns -1, and the offset arithmetic around that sentinel either panics or silently slices the wrong text.

solid answer

~50 s

`strings.Cut(s, sep)` returns `(before, after string, found bool)` in one call, so the "was the separator there?" question is answered by a value you must consciously discard, not by a sentinel. `strings.Index` returns `-1` when the separator is absent, and both natural ways of using that offset go wrong: `s[:i]` panics with a slice-bounds error, while `s[i+1:]` quietly becomes `s[0:]` — the whole line — so a malformed header line turns into a plausible-looking value that flows on into the rest of the program. The `Cut` form also reads better, because the two halves get names at the point they are produced instead of being reconstructed from arithmetic. The same argument applies to `strings.CutPrefix` and `strings.CutSuffix`, which replace the `HasPrefix` plus `s[len(prefix):]` pair, and to the `bytes` mirrors when you are working on a buffer.

code

go · 9 lines
go
// Broken: strings.Index returns -1 when the line has no "=".
i := strings.Index(line, "=")
rawValue := line[i+1:] // i == -1 makes this line[0:] — the whole line

// Correct: strings.Cut reports whether the separator was found.
key, value, found := strings.Cut(line, "=")
if !found {
	return fmt.Errorf("malformed header line %q", line)
}

go deeper

for a junior

Remember the two return shapes: Index gives one int and -1 means not found; Cut gives before, after and a found boolean. Reach for Cut whenever you were about to slice around an offset.

for a middle

Be ready to walk through what line[i+1:] evaluates to when i is -1, and why that is worse than the panic you get from line[:i]. Naming both failure modes is the answer.

for a senior

Show the review instinct: any slice expression built from an Index result without a negative guard is a defect, and the fix is to change the API used, not to add the guard. Pair it with a table test over separator-absent and separator-repeated inputs.

for a principal

The wider call is picking library shapes that make the error state unrepresentable in code your team writes daily. Sentinel returns invite arithmetic; multi-value returns force a decision, and that is worth standardising on.

## The signatures ```go func Index(s, substr string) int func Cut(s, sep string) (before, after string, found bool) func HasPrefix(s, prefix string) bool func CutPrefix(s, prefix string) (after string, found bool) ``` `Index` reports the byte offset of the first occurrence of `substr`, or `-1` if it is not there. `Cut` does the same search but hands back the two halves and a boolean instead of an offset. ## What goes wrong with the sentinel Consider a parser for a message-broker handshake frame, whose header block is a sequence of `key=value` lines. The obvious code is: ```go i := strings.Index(line, "=") key, value := line[:i], line[i+1:] ``` On a well-formed line this is correct. On `"CONNECT"` — a line with no `=` at all, which a truncated or malformed frame can easily produce — `i` is `-1`, and the two halves fail in *different* ways: - `line[:i]` is `line[:-1]`, which panics at runtime with `slice bounds out of range`. A panic in a connection-handling goroutine is bad, but at least it is loud. - `line[i+1:]` is `line[0:]`, which is the **entire line**. No panic, no error: the parser records a value of `"CONNECT"` under whatever key it computed, and the bad data flows downstream. Debugging that later means working backwards from a wrong value to a missing separator several layers away. The second failure is the one worth naming in a review. It is silent, it is plausible, and it survives every test whose inputs all contain the separator. ## What Cut changes ```go key, value, found := strings.Cut(line, "=") if !found { return fmt.Errorf("malformed header line %q", line) } ``` Three properties matter. 1. **The not-found case is a value, not an encoding.** To ignore it you have to write `_`, which is visible in review. There is no arithmetic that can accidentally paper over it. 2. **No offsets appear in your code.** Every `+1` and `-1` around a separator length is a chance to be wrong, and the errors are worst when the separator is more than one byte — a `": "` separator needs `i+2`, and `i+1` leaves a stray space at the front of every value. 3. **The behaviour when the separator is absent is defined and useful:** `Cut` returns `before == s`, `after == ""`, `found == false`. So `before` is still the whole input, which is often exactly the right fallback for an optional suffix such as a port or a fragment. `Cut` splits at the **first** occurrence only. That is what you want for `key=value`, where the value is allowed to contain `=`. It is deliberately not a general splitter. ## The prefix and suffix pair The same shape exists for prefixes. The old idiom needed the length repeated: ```go if strings.HasPrefix(line, "CONNECT ") { rest := line[len("CONNECT "):] // literal repeated, easy to drift } ``` `strings.CutPrefix(line, "CONNECT ")` returns `(rest, true)` and, when the prefix is absent, `(line, false)` — the literal appears once, and the not-found case is again a boolean rather than an unchecked assumption. `strings.CutSuffix` is the mirror. `HasPrefix` and `Index` are not obsolete. `HasPrefix` is right when you only need the yes/no answer and are not going to slice; `Index` is right when the offset itself is the answer, for example when you want to report the column at which something was found, or when you are scanning forward through a buffer and need to advance a cursor. ## Reviewing for this A useful review heuristic: **any expression of the form `s[i+1:]` or `s[:i]` where `i` came from an Index call is suspect until you find the `if i < 0` guard.** If the guard is missing, the fix is nearly always `Cut` rather than adding the guard, because `Cut` makes the same class of bug unwriteable next time. The corresponding test is a table that enumerates the shapes of the input rather than the happy path: separator absent, empty input, separator first (`"=value"`), separator last (`"key="`), and separator repeated (`"key=a=b"`). Those five rows are what actually distinguishes correct parsing code from code that happens to work, and they are cheap to write once the parser is a function taking a line and returning `(key, value string, err error)`. ## Byte slices When the frame is still a `[]byte` from the connection, use `bytes.Cut`, `bytes.CutPrefix` and `bytes.Index`, which have identical semantics on `[]byte`. Converting to `string` first only to search is a needless copy.

  • When is strings.Index still the right tool?
    When the offset itself is the answer rather than a step towards slicing — reporting the column of a bad byte, advancing a cursor through a buffer in a loop, or checking a position against a limit. `Cut` throws the offset away, so it cannot express those. Use `Index` with an explicit `if i < 0` guard, and `strings.HasPrefix` when you need only a yes or no.
  • What does strings.Cut do when the separator appears more than once in the line?
    It cuts at the first occurrence only: `strings.Cut("k=a=b", "=")` returns `"k"`, `"a=b"`, `true`. That is usually what a key/value format wants, since the value may legitimately contain the separator. If you need the last occurrence instead, Go 1.27 added `strings.CutLast`; before that, `strings.LastIndex` with an explicit guard.
  • How would you test the parser so this class of bug cannot come back?
    A table test whose rows are input shapes, not happy paths: separator absent, empty line, separator first (`"=v"`), separator last (`"k="`), and separator repeated (`"k=a=b"`). Assert the returned key, value and error for each. Those five rows catch every offset mistake, and they run in microseconds because the parser is a pure function over a string.

saying these in an interview costs you the question

  • Thinks strings.Index returns 0 when the substring is absent
  • Assumes a missing separator always panics rather than sometimes slicing silently
  • Uses i+1 for a multi-byte separator
  • Believes strings.Cut splits at every occurrence
  • Converts a []byte to string just to call strings.Index