skip to content

Splitting, Stripping and Searching

str.strip removes a set of characters, not a suffix, which is why removeprefix and removesuffix exist. split, partition, find and startswith round out the methods interviewers expect on sight.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

Why doesn't 'scores.csv'.strip('.csv') simply remove the file extension?

level: juniorimportance: must knowfreq 62%

answer

  1. The argument is not a word
  2. Both ends, character by character
  3. Order of the characters is irrelevant
  4. A dedicated pair arrived in 3.9
  5. removeprefix and removesuffix

basics

~10 s

str.strip takes a set of characters, not a suffix. It peels any of '.', 'c', 's' or 'v' off both ends, so 'scores.csv'.strip('.csv') returns 'ore'. Use str.removesuffix('.csv'), added in Python 3.9.

solid answer

~40 s

`str.strip(chars)` reads its argument as a **set of characters to remove from both ends**, not as a substring to match. It deletes leading characters while they are in the set, stops at the first that is not, then repeats from the right. So `'scores.csv'.strip('.csv')` strips the leading `s` and `c`, stops at `o`, then eats the trailing `v`, `s`, `c`, `.` and one more `s`, leaving `'ore'`. The order of characters in the argument is irrelevant. The correct tool for an exact suffix is `str.removesuffix('.csv')` (Python 3.9, PEP 616), which removes one exact trailing occurrence and returns the string unchanged when it is absent; `str.removeprefix` is its mirror. Both also exist on `bytes` and `bytearray`. For filesystem paths, `pathlib` is a better fit than any string method.

code

pycon · 8 lines
pycon
>>> 'scores.csv'.strip('.csv')
'ore'
>>> 'banana.csv'.strip('.csv')
'banana'
>>> 'scores.csv'.removesuffix('.csv')
'scores'
>>> 'notes.txt'.removesuffix('.csv')
'notes.txt'

go deeper

for a junior

Memorise the one sentence that matters: strip takes a set of characters, not a suffix. Be able to say removesuffix is the right method for an exact ending, and remember that string methods return a new string you must assign.

for a middle

Be ready to trace the algorithm out loud on a concrete input, explain why the character order is irrelevant, and date removeprefix/removesuffix to Python 3.9. Know why replace and a hardcoded slice are both wrong answers.

for a senior

Show why this defect is dangerous rather than merely wrong: it produces silently corrupted values on a subset of inputs and passes a happy-path suite. Flag any multi-character strip argument in review, and know the empty-affix edge in the hand-rolled guard.

for a principal

Own the class, not the instance: this is a whole family of APIs whose contract reads naturally but means something else. Argue for banning multi-character strip arguments in a linter rule and for pushing filename work into pathlib so the string method never gets the chance.

### `strip` trims a character set, not a suffix `str.strip(chars)` removes characters from **both ends** of a string, and `chars` is a *bag of characters that are allowed to be removed* — never a substring to be matched. The implementation walks in from the left deleting every character that appears in the set, stops dead at the first character that does not, then does the same walk from the right. Trace `'scores.csv'.strip('.csv')` with the set `{'.', 'c', 's', 'v'}`: - From the left: `s` is in the set, gone. `c` is in the set, gone. `o` is not — stop. We now have `'ores.csv'`. - From the right: `v` gone, `s` gone, `c` gone, `.` gone, and now the trailing character is the `s` of `ores`, which is also in the set — gone. `e` is not in the set — stop. - Result: `'ore'`. The order of characters in the argument never matters: `'vs.c'` behaves identically to `'.csv'`. `str.lstrip` and `str.rstrip` are the one-sided versions with the same set semantics. With no argument, `s.strip()` removes whitespace — spaces, tabs, newlines, carriage returns, form feeds, vertical tabs, and the characters Unicode classifies as whitespace. That is the job the method was designed for: trimming padding off a field read out of a file or a form. The set semantics are exactly right for that use, which is why the API looks the way it does. ### Why this bug survives testing The trap is that the wrong code frequently produces the right answer. `'banana.csv'.strip('.csv')` returns `'banana'`, because no leading character of `banana` is in the set and the trailing `.csv` happens to be entirely inside it. A test suite whose sample names never begin or end with `c`, `s`, `v` or a dot passes every case. The failure surfaces later on the one record whose name starts with `scan-` or ends in `-specs`, and it surfaces as *silently wrong data*, not as an exception — the worst failure shape there is. When you review a `strip` call whose argument is longer than one character, treat it as a defect until proven otherwise. ### `removeprefix` and `removesuffix` Python 3.9 added `str.removeprefix` and `str.removesuffix` (PEP 616) precisely because this mistake was endemic. Their contract is narrow and therefore safe: - The argument is matched as an **exact substring**, not a set. - **One** occurrence is removed, and only at that end. - If the affix is not present, the string comes back unchanged — no exception, no partial removal. - They exist on `str`, `bytes` and `bytearray`, and like every string method they return a **new** object; `str` is immutable, so nothing is edited in place. ### The alternatives, and why each is worse `str.replace('.csv', '')` removes **every** occurrence anywhere in the string, so `'csv_export.csv'` becomes `'_export'`. Slicing a hardcoded count — `name[:-4]` — is a magic number that silently breaks the moment the extension changes length. The hand-rolled guard is closest to correct: ```python if name.endswith(suffix): name = name[:-len(suffix)] ``` but it carries one real edge: when `suffix` is the empty string, `endswith('')` is `True` and `name[:-0]` is `name[:0]`, which is `''`. The whole string vanishes. `removesuffix('')` returns the string unchanged, as you would expect. That empty-affix case is not academic when the suffix comes from configuration or from a caller's argument. For filenames specifically, neither string method is really the right answer: `pathlib` understands directory separators and multi-dot names, and using it removes the whole class of bug. ### Where `strip` is still the right call `strip` earns its keep whenever you genuinely mean "peel any of these characters off the ends": `field.strip()` to drop surrounding whitespace, `value.strip('"')` to drop any number of wrapping quotes, `line.rstrip()` to drop a line terminator regardless of whether it is a bare newline or a carriage-return pair. The rule of thumb that keeps you out of trouble: **if the argument is a word, you want `removeprefix` or `removesuffix`; if it is a set of throwaway characters, you want `strip`.** ### Interview framing This is a first-screen question, and the interviewer is watching for two things: that you know the semantics rather than the folklore, and that you reach for the modern method by name and can date it. Saying "`strip` takes a character set, so use `removesuffix` — that landed in 3.9" answers it completely in one sentence.

  • What does str.removesuffix do when the suffix is not present at the end?
    It returns a copy of the original string, unchanged. There is no exception and no partial removal, which is what makes it safe to call unconditionally on a mixed batch of names. The same holds for str.removeprefix, and for both methods on bytes and bytearray. Note that it removes at most one occurrence: 'a.csv.csv'.removesuffix('.csv') gives 'a.csv', not 'a'.
  • Does str.strip modify the string in place, and what does it return when nothing matches?
    It never modifies in place — str is immutable, so every string method returns a new object and the original binding is untouched unless you reassign it. When nothing at either end is in the character set, strip returns the original content; CPython may hand back the same object, but that is an implementation detail you must not rely on. Forgetting to reassign the result is the other classic strip bug.
  • You need to remove a prefix that might appear twice, like 'tmp_tmp_report'. How do you write that?
    Loop while the prefix is still present: `while name.startswith('tmp_'): name = name.removeprefix('tmp_')`. removeprefix deliberately strips one occurrence, so repetition is your decision to make explicit. Do not reach for lstrip('tmp_') as the shortcut — that is a character set again, and it would also chew through the leading letters of a name like 'ptm_totals'.

strip is a bouncer with a list of names who turns people away at both doors until someone not on the list arrives; removesuffix is a doorman told to remove one specific person standing at the back.

saying these in an interview costs you the question

  • Claims strip('.csv') removes the substring '.csv'
  • Thinks the order of characters in the argument matters
  • Reaches for replace('.csv', '') to drop an extension
  • Believes strip mutates the string in place
  • Expects removesuffix to raise when the affix is absent
  • Cannot name removeprefix or removesuffix at all

context

open as a page

How does str.split() with no argument differ from str.split(' ')?

level: middleimportance: must knowfreq 66%

basics

~20 s

With no argument, str.split treats any run of whitespace as one separator and discards leading and trailing whitespace, so it never yields empty strings. With ' ' it splits on each single space, so runs of spaces produce empty strings.

open as a page

How do str.find and str.index differ when the substring is absent?

level: middleimportance: should knowfreq 48%

basics

~20 s

str.find returns -1 when the substring is absent; str.index raises ValueError instead. Both return the index of the first occurrence otherwise. For a yes/no test use the in operator, never the truthiness of find's result.

open as a page

When is str.translate with str.maketrans better than chained str.replace calls?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

str.translate applies a whole per-character mapping in one pass, so replacements never cascade into each other and cost stays linear in the text. Chained str.replace calls each scan the string again and can re-transform an earlier call's output.

open as a page