skip to content

Common String API Methods

The methods you actually use daily: substring, indexOf, split with its regex argument, format, join, and the chars and codePoints stream views. Interviewers probe split's regex nature and substring's copying behavior since Java 7.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

6

What do indexOf and lastIndexOf do in Java, and how do their fromIndex variants and return values behave?

level: juniorimportance: must knowfreq 70%

answer

  1. indexOf = first/leftmost; lastIndexOf = last/rightmost
  2. not found → -1 (basis of contains())
  3. indexOf(x, from) searches forward ≥ from; lastIndexOf(x, from) searches backward ≤ from
  4. empty "" matches: indexOf→0/clamped, lastIndexOf→length
  5. case-sensitive, literal — NOT regex

basics

~10 s

indexOf finds the first position of a character or substring and returns that index, or -1 if not found. lastIndexOf does the same but finds the last (rightmost) occurrence.

solid answer

~40 s

indexOf(target) returns the zero-based index of the first occurrence of a char or substring; lastIndexOf(target) returns the index of the last occurrence. Both return -1 when the target is absent — that sentinel is how you test 'contains', and indeed contains() is implemented via indexOf. The fromIndex overloads control where the search starts: indexOf(target, fromIndex) searches forward from fromIndex, while lastIndexOf(target, fromIndex) searches backward, treating fromIndex as the highest position allowed to start a match. For an empty target string, indexOf returns fromIndex (clamped to length) and lastIndexOf returns fromIndex too. The char overloads actually take an int code point, and search is case-sensitive and exact — there's no regex here, unlike split or matches.

go deeper

for a junior

Knows indexOf finds the first occurrence, lastIndexOf the last, and that -1 means not found.

for a middle

Correctly uses both fromIndex overloads, knows the search is case-sensitive and literal, and uses -1 (not 0) to test absence.

for a senior

Explains the opposite meaning of fromIndex in the two methods, the empty-string and clamping edge cases, the int/code-point overload, and that it's not regex.

for a principal

Advises on performance (naive scan complexity, when to switch to Pattern/automaton) and on code-point-correct searching for international text across a codebase.

## Purpose `indexOf` and `lastIndexOf` answer 'where does this character or substring appear?' They return a **zero-based index** (position counting from 0) or **-1** if the target does not occur. This -1 'sentinel' (a special return value signalling 'not found') is the idiomatic way to test presence; `String.contains(cs)` is literally `indexOf(cs.toString()) > -1`. ## The four core forms For both a single character and a substring: - `indexOf(target)` — first (leftmost) occurrence, scanning left→right. - `indexOf(target, fromIndex)` — first occurrence **at or after** `fromIndex`. - `lastIndexOf(target)` — last (rightmost) occurrence, scanning right→left. - `lastIndexOf(target, fromIndex)` — last occurrence that **starts at or before** `fromIndex`. Example with `"banana"` (b=0,a=1,n=2,a=3,n=4,a=5): - `"banana".indexOf('a')` → 1 - `"banana".lastIndexOf('a')` → 5 - `"banana".indexOf('a', 2)` → 3 (first 'a' at index ≥ 2) - `"banana".lastIndexOf('a', 4)` → 3 (last 'a' at index ≤ 4) - `"banana".indexOf("na")` → 2 (substring match returns the start index) - `"banana".indexOf('z')` → -1 (not present) ## fromIndex direction is the tricky part The same parameter name means opposite things: - For `indexOf`, `fromIndex` is the **lower bound** — the search moves **forward** and won't return an index below it. A `fromIndex` below 0 is treated as 0; above `length` yields -1 (or, for an empty target, the clamped length). - For `lastIndexOf`, `fromIndex` is the **upper bound** on where a match may **start** — the search moves **backward** from there. A `fromIndex` above `length` is clamped down; below 0 yields -1. ## The empty-string special case Searching for the empty string `""` always 'matches' at a position. `"abc".indexOf("")` returns 0; `"abc".indexOf("", 2)` returns 2; `"abc".indexOf("", 99)` returns 3 (clamped to length). `"abc".lastIndexOf("")` returns the length, 3. This falls out of the matching definition and occasionally surprises people. ## char overloads take an int The character overloads are declared as `indexOf(int ch)`, accepting a **code point** (the numeric Unicode value), so they can match supplementary characters above the BMP that don't fit in a single `char`. The substring overloads match exact char sequences. All matching is **case-sensitive** and **literal** — these methods do **not** interpret regular expressions, unlike `split`, `matches`, `replaceAll`. To find ignoring case, lowercase both sides first or use a regex method. ## Complexity These run in O(n·m) worst case (naive scan of length-n text for length-m needle) — fine for typical inputs; for heavy repeated searching consider a precompiled `Pattern` or specialized algorithm.

  • How is String.contains implemented, and why does that matter?
    contains(cs) returns indexOf(cs.toString()) > -1. It matters because contains is just a convenience over indexOf, so it is also case-sensitive and literal (no regex), and shares indexOf's performance.
  • What does "abc".indexOf("", 99) return and why?
    It returns 3. The empty string matches everywhere, and fromIndex 99 is clamped to the string length 3, so the reported match position is 3.

saying these in an interview costs you the question

  • Thinking lastIndexOf's fromIndex is a lower bound like indexOf's
  • Expecting these to interpret regex patterns
  • Returning a default of 0 instead of treating -1 as 'not found'
  • Assuming the search is case-insensitive
  • Confusing the returned start index of a substring match with a length

context

open as a page

How does String.substring() work in Java, including its index arguments and what happens to the original string?

level: juniorimportance: must knowfreq 78%

basics

~10 s

substring(begin, end) returns the part of the string from begin up to but not including end. The original string is unchanged because strings are immutable; you get a new string back.

open as a page

How does String.split() work, and why is its delimiter argument a regular expression with surprising trailing-empty-string behavior?

level: middleimportance: must knowfreq 74%

basics

~10 s

split breaks a string into an array of pieces around a delimiter. The delimiter is a regex, not plain text, and by default empty pieces at the end of the result are dropped.

open as a page

Given that Java Strings are immutable, what are the performance implications of the String API in loops, and how do StringBuilder, the '+' operator, and intern() fit in?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Every String operation creates a new String, so building one piece-by-piece with '+' in a loop is slow because it makes many throwaway strings. Use StringBuilder to build in a loop; it edits a buffer in place.

open as a page

What do String.format() and String.join() do, and when would you choose each over manual concatenation?

level: middleimportance: should knowfreq 58%

basics

~10 s

String.format() builds a string from a template with placeholders like %s and %d filled by arguments. String.join() glues several strings together with a separator between them.

open as a page

What is the difference between String.chars() and String.codePoints(), and why does it matter for characters like emoji?

level: seniorimportance: should knowfreq 40%

basics

~20 s

chars() gives a stream of UTF-16 code units, so a character stored as two units (like an emoji) shows up as two values. codePoints() gives a stream of full Unicode characters, so each emoji is one value.

open as a page