skip to content

How do str.find and str.index differ when the substring is absent?

level: middleimportance: should knowfreq 48%

answer

  1. One returns a sentinel, one raises
  2. The sentinel is a valid index too
  3. Zero is falsy and minus one is truthy
  4. For yes/no, skip both methods
  5. startswith accepts a tuple of candidates

basics

~20 s

str.find returns -1 when the substring is absent; str.index raises ValueError instead. Both return the index of the first occurrence otherwise. For a yes/no test use the in operator, never the truthiness of find's result.

solid answer

~50 s

Both return the lowest index of the first occurrence of the substring. They differ only in the missing case: `str.find` returns `-1`, while `str.index` raises `ValueError`. `str.rfind` and `str.rindex` are the same pair searching from the right. The choice is about whether absence is expected: use `find` when a miss is a normal branch, and `index` when a miss is a bug you want to fail loudly. The classic trap is truthiness — `if s.find(x):` is broken in both directions, because `-1` (not found) is truthy and `0` (found at the start) is falsy. Compare explicitly with `!= -1`, or, for a plain membership test, use `x in s`, which is clearer and does not hand you an index you do not need. Both methods accept optional `start` and `end` arguments so you can search a window without slicing a copy.

code

python · 9 lines
python
line = 'sku-1120,blue widget'

if line.find('sku'):          # 0 is falsy - this never runs
    print('prefix present')
if line.find('zzz'):          # -1 is truthy - this always runs
    print('absent, yet here we are')

print(line.find(','))         # 8
print('sku' in line)          # True - the right yes/no test

go deeper

for a junior

Remember the one-line difference: one of them returns -1 for a miss and the other raises ValueError. For a simple yes/no check, use the in operator rather than either method.

for a middle

Explain why -1 is a dangerous sentinel — it is itself a valid index — and why testing the result for truth is wrong on both the found-at-zero and the not-found case. Know the right-hand variants and the start/end window.

for a senior

Show the judgement call: a sentinel for an expected miss, an exception for a contract violation you want to surface at the bad record. Point out that an unchecked -1 flowing into a slice yields plausible wrong data that a fixture set full of well-formed records will never catch.

for a principal

Own the convention: decide as a team whether parsing code fails loudly on malformed input or degrades, and make the search API follow that choice. An unchecked sentinel is a silent-corruption class, which is the expensive kind to find in production.

### The same search, two failure contracts `str.find(sub)` and `str.index(sub)` do identical work: scan left to right and report the lowest index at which `sub` begins. They differ only in what happens when it is not there. - `find` returns the sentinel `-1`. - `index` raises `ValueError: substring not found`. `str.rfind` and `str.rindex` are the mirror pair, reporting the **highest** starting index. All four accept optional `start` and `end` arguments that restrict the search window — `s.find(',', 20)` searches from index 20 onward without building the slice `s[20:]`, which matters in a loop over a long string because a slice copies. ### Choosing between a sentinel and an exception The choice is a design decision about whether absence is *normal*. Use `find` when a miss is an ordinary branch you intend to handle: an optional comment marker in a line, an optional delimiter in a free-form field. Use `index` when the substring is part of the contract and its absence means the input is malformed — you want a traceback pointing at the bad record rather than a `-1` that flows onward as an array index. The worst outcome is neither: using `find` and then forgetting to check, so `-1` reaches a slice and quietly means "one from the end". That is the real hazard of the sentinel. `s[:s.find(',')]` looks harmless and is correct whenever the comma is present; when it is absent it becomes `s[:-1]`, silently dropping the last character instead of failing. The result is plausible enough to survive review and to sail through a large regression pack whose fixtures all contain the delimiter. ### The truthiness trap The single most common bug on these methods is treating the return value as a boolean: ```python if line.find('sku'): # WRONG in both directions ... ``` `-1` is truthy, so the branch fires when the substring is **absent**. `0` is falsy, so the branch is skipped when the substring is found **at the very start** — the most likely position for a prefix. The condition is wrong on exactly the two cases that matter. Write `if line.find('sku') != -1:` when you need the index, and `if 'sku' in line:` when you do not. ### `in` is usually the right answer For a yes/no test, the membership operator reads better and cannot be misused: `'sku' in line` returns a `bool`. It is implemented by the same underlying search, so there is no performance argument for `find`. Reach for `find` or `index` only when you actually need the position. Similarly `str.count(sub)` answers "how many" without a loop, and counts **non-overlapping** occurrences left to right — `'aaaa'.count('aa')` is 2, not 3. ### Prefix and suffix checks: `startswith` takes a tuple For an anchored test, `str.startswith` and `str.endswith` are both clearer and cheaper than a `find` comparison, because they stop as soon as the anchored comparison fails instead of scanning the whole string. Both also accept a **tuple** of candidates: ```python if code.startswith(('SKU-', 'EAN-')): ... ``` That is one call instead of a chain of `or`, and it is the idiomatic form. Note that it must be a tuple: passing a list raises `TypeError`. Both methods also take `start` and `end` arguments, so you can anchor the test at an offset inside the string. A frequent mistake is checking a suffix with `endswith` and then trimming it with `strip`, which takes a character set — `removesuffix` is the matching remover. ### `bytes` has the same API `bytes` and `bytearray` carry `find`, `index`, `rfind`, `rindex`, `startswith`, `endswith` and `count` with the same contracts, but the arguments must be bytes-like, not `str`. Mixing the two raises `TypeError` rather than silently comparing, which is deliberate: an implicit conversion there would be an encoding guess. When a sync between two systems hands you raw bytes off a socket and you search them with a `str` literal, that `TypeError` is the language protecting you from exactly that guess. ### Interview framing State the difference in one sentence — sentinel versus exception — then show judgement by naming when each is appropriate, and finish with the truthiness trap and the `in` operator. Mentioning the `start`/`end` window and the tuple form of `startswith` marks you as someone who has read the method list rather than only the two entries the question named.

  • Why is `if s.find(sub):` wrong in both directions?
    Because find returns an index, not a boolean. When the substring is absent it returns -1, which is truthy, so the branch fires on a miss. When the substring is found at position 0 it returns 0, which is falsy, so the branch is skipped on the most common hit for a prefix. Compare explicitly against -1, or use the membership operator when you only need yes or no.
  • What do the start and end arguments to str.find buy you over slicing first?
    They restrict the search window without copying. s[20:].find(',') builds a new string first, so scanning a long string in a loop is quadratic in allocation; s.find(',', 20) searches in place and, importantly, returns an index into the original string rather than into the slice, so you do not have to add the offset back. str.index, rfind, rindex, count, startswith and endswith all take the same pair.
  • Does str.startswith accept a list of prefixes?
    No — it accepts a single string or a tuple of strings, and a list raises TypeError. The tuple form checks several candidates in one call, which is cleaner than chaining or. endswith behaves identically. The tuple-only rule catches people out because a list looks interchangeable everywhere else in Python.

saying these in an interview costs you the question

  • Uses find's result directly as a boolean condition
  • Thinks find raises when the substring is absent
  • Lets -1 flow into a slice or an index
  • Says index returns None on a miss
  • Uses find when a plain membership test would do
  • Passes a list of prefixes to startswith

context