skip to content

New String Methods (Java 11+)

isBlank, strip, lines, repeat, indent and transform, added since Java 11. The one interviewers probe is strip versus trim: strip is Unicode-aware, trim only removes characters below U+0020.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

5

What is the difference between String.strip() and String.trim() in Java, and why was strip() added?

level: middleimportance: must knowfreq 70%

answer

  1. trim = code point <= U+0020 (pre-Unicode 1.0)
  2. strip = Character.isWhitespace (Unicode-aware, Java 11)
  3. stripLeading / stripTrailing for one side
  4. U+00A0 non-breaking space NOT stripped by either
  5. both return a new String (immutable)

basics

~10 s

Both remove leading and trailing whitespace. trim() (old) only removes characters up to and including space (U+0020). strip() (Java 11) is Unicode-aware, so it also removes Unicode whitespace like non-breaking-aware spaces and ideographic spaces.

solid answer

~40 s

trim(), from Java 1.0, removes any leading/trailing character whose code point is <= U+0020 (space). It predates Unicode awareness, so it misses many real whitespace characters and also strips some non-whitespace control characters below space. strip(), added in Java 11, uses Character.isWhitespace() to decide what to remove, so it correctly handles the full Unicode whitespace set (e.g. U+2028 line separator, U+3000 ideographic space) and leaves non-whitespace controls alone. There are also directional variants stripLeading() and stripTrailing(). Prefer strip() in new code for correct, locale/Unicode-safe trimming; trim() remains only for legacy behavior compatibility. Note neither removes the non-breaking space U+00A0, because Character.isWhitespace() excludes it by design.

code

java · 7 lines
java
String s = "   hello   "; // em space + ASCII spaces
System.out.println("[" + s.trim() + "]");  // [   hello   ] -> em spaces left, only inner trim of ASCII? actually leaves em spaces
System.out.println("[" + s.strip() + "]"); // [hello]

String html = " value ";        // non-breaking spaces
System.out.println("[" + html.strip() + "]"); // [ value ] -- NOT removed
System.out.println("[" + html.replace(' ',' ').strip() + "]"); // [value]

go deeper

for a junior

Knows both remove leading/trailing whitespace and that strip() is the newer Unicode-aware one to prefer.

for a middle

Articulates the exact rules: trim() uses code point <= U+0020; strip() uses Character.isWhitespace; knows stripLeading/stripTrailing exist.

for a senior

Explains why trim() is wrong for Unicode (misses U+3000/U+2028, strips control chars) and the U+00A0 gotcha; chooses strip() in new code with reasoning.

for a principal

Discusses backward-compat constraints that froze trim(), implications for parsing/normalization pipelines, and sets a team convention (strip() + explicit NBSP handling) for text ingestion.

## Background: what "whitespace" means Text is made of **code points** — numeric identifiers for characters defined by the Unicode standard. ASCII space is code point U+0020. But Unicode defines many other characters that visually act as blank space: the tab (U+0009), the line separator (U+2028), the ideographic space used in CJK text (U+3000), and more. "Trimming" a string means removing such blank characters from its start and end while leaving the meaningful middle intact. ## The old method: trim() `String.trim()` has existed since Java 1.0 (1996), before Java fully embraced Unicode. Its rule is brutally simple: remove any leading or trailing character whose **code point value is less than or equal to U+0020** (decimal 32, the space). This has two consequences: 1. It removes some characters that are NOT whitespace — e.g. control characters like U+0000 (null) or U+0007 (bell), because their code point is below 32. 2. It FAILS to remove many characters that ARE whitespace in Unicode but have code points above 32 — e.g. U+2028 (line separator) or U+3000 (ideographic space). Those would be left in place. So `trim()` is neither a strict whitespace remover nor Unicode-correct; it is a historical artifact kept for backward compatibility. ## The new method: strip() (Java 11) `String.strip()` removes leading and trailing characters for which `Character.isWhitespace(int codePoint)` returns true. `Character.isWhitespace` is the JDK's official Unicode-aware definition of whitespace. As a result: - It correctly removes Unicode whitespace like U+2028 and U+3000. - It does NOT remove non-whitespace control characters below U+0020 (it leaves them, which is usually what you want). There are two companion methods for one-sided trimming: - `stripLeading()` — removes only from the start (left). - `stripTrailing()` — removes only from the end (right). All three return a new String (Java strings are immutable; methods never mutate in place). ## A subtle gotcha: the non-breaking space The non-breaking space U+00A0 (`&nbsp;` in HTML) is, perhaps surprisingly, **not** removed by `strip()`, because `Character.isWhitespace(0x00A0)` returns false (Unicode classifies it as a space character but specifically a *non-breaking* one, which `isWhitespace` excludes). If you scrape HTML, you may still see leading/trailing U+00A0 after `strip()`. To remove it you must handle it explicitly (e.g. `replace(' ', ' ').strip()`). ## When to use which - New code: use `strip()` / `stripLeading()` / `stripTrailing()` — correct Unicode behavior. - Existing code where you must reproduce the exact legacy behavior (e.g. an established data format that relied on the <= U+0020 rule): keep `trim()`. ## Deriving the answer at any level Given the above: trim() = pre-Unicode, code-point <= 32 rule; strip() = Unicode-aware via Character.isWhitespace, plus directional variants; the U+00A0 exception trips people up. From those facts you can reconstruct a junior, senior, or principal answer.

  • Does strip() remove the HTML non-breaking space (U+00A0)?
    No. Character.isWhitespace(0x00A0) is false, so strip() leaves it. You must replace it explicitly first.
  • What do trim() and strip() return when the string is all whitespace?
    An empty string (""). They return a new String; the original is unchanged because Strings are immutable.

saying these in an interview costs you the question

  • Claiming trim() and strip() are identical
  • Saying strip() removes the non-breaking space U+00A0
  • Thinking trim() is Unicode-aware
  • Believing these methods mutate the original string

context

open as a page

What does String.isBlank() do (Java 11), and how does it differ from isEmpty()?

level: juniorimportance: should knowfreq 55%

basics

~20 s

isEmpty() is true only when the string has length 0. isBlank() (Java 11) is true when the string is empty OR contains only whitespace characters. So " ".isBlank() is true but " ".isEmpty() is false.

open as a page

What does String.lines() return, and why is it preferable to split("\n") for processing text line by line?

level: middleimportance: should knowfreq 50%

basics

~10 s

lines() (Java 11) returns a Stream<String> of the lines in the text, splitting on line terminators (\n, \r, \r\n) without keeping them. It is lazy and handles all line-ending styles, unlike split("\n").

open as a page

What does String.repeat(int) do, and what are its edge cases and uses?

level: juniorimportance: nice to knowfreq 40%

basics

~10 s

repeat(n) (Java 11) returns a new string formed by concatenating this string n times. "ab".repeat(3) is "ababab". repeat(0) gives "", and a negative count throws IllegalArgumentException.

open as a page

What do String.indent(int) and String.transform(Function) do (Java 12), and when would you use each?

level: seniorimportance: nice to knowfreq 35%

basics

~20 s

indent(n) (Java 12) adds n spaces of indentation to each line (or removes up to n with a negative value) and normalizes line endings to \n, ensuring each line ends with a newline. transform(fn) applies a function to the whole string and returns its result, letting you chain custom operations fluently.

open as a page