What does String.lines() return, and why is it preferable to split("\n") for processing text line by line?
answer
- lines() -> Stream<String> (Java 11)
- handles \n, \r, \r\n; strips the terminator
- lazy / stream-friendly (filter, map, count)
- split("\n") leaves stray \r on Windows text
- no spurious trailing empty for a trailing terminator
basics
~10 slines() (Java 11) returns a Stream<String> of the lines in the text, splitting on line terminators (\n, \r, \r\n) without keeping them. It is lazy and handles all line-ending styles, unlike split("\n").
solid answer
~40 sString.lines(), added in Java 11, returns a Stream<String> where each element is one line of the source string. It recognizes all common line terminators — \n (Unix), \r (old Mac), and \r\n (Windows) — and the terminators themselves are not included in the returned lines. It is lazy: lines are produced on demand as the stream is consumed, which is efficient for large text and composes with the Stream API (filter, map, count). Compared to split("\n"), lines() is more correct because split("\n") only handles the \n terminator (leaving stray \r on Windows text), is eager (builds a whole array), and uses regex. A trailing line terminator does not produce a spurious trailing empty line with lines(). It is the idiomatic way to iterate text by line in modern Java.
go deeper
Knows lines() breaks a string into its lines and returns a Stream you can loop over.
Explains the Stream<String> return type, terminator handling (\n/\r/\r\n) with stripping, and laziness.
Contrasts with split("\n") (stray \r, eager, regex) and trailing-empty behavior; chooses lines() for cross-platform correctness.
Considers performance on very large inputs (lazy streaming, short-circuiting), encoding/normalization upstream, and standardizing line handling across a text-processing module.
## The problem: text has multiple line-ending conventions A block of text is a sequence of lines separated by **line terminators**. Historically platforms disagree on what that terminator is: - Unix/Linux/macOS (modern): line feed `\n` (U+000A) - Windows: carriage return + line feed `\r\n` (U+000D U+000A) - Classic Mac OS (pre-OS X): carriage return `\r` (U+000D) Java's `lines()` also recognizes a few additional Unicode line/paragraph separators per its spec, but the three above are the ones that matter day to day. ## The old approach and its bugs: split("\n") A common idiom was `text.split("\n")`. Problems: 1. **Only handles \n.** On Windows text (`\r\n`), each resulting line keeps a trailing `\r`, which silently corrupts comparisons and output. 2. **Eager + regex.** `split` compiles a regex and builds the whole `String[]` array up front, even if you only need the first few lines. 3. **Trailing-empty quirks.** `split` has its own rules about trailing empty strings (the zero-limit form drops trailing empties), which can surprise. ## The modern method: lines() `String.lines()` returns a `Stream<String>`: - **Splits on \n, \r, and \r\n** (and a few Unicode separators), treating any of them as a line boundary. - **Strips the terminator** — the returned line strings never contain `\n` or `\r`. - **Is lazy** — it produces a `Stream`, so lines are computed as you pull them. You can do `text.lines().filter(...).findFirst()` and stop early without scanning the whole string. - **No spurious trailing empty line** for a string that ends in a terminator: `"a\nb\n".lines()` yields `["a", "b"]`, not `["a", "b", ""]`. (An empty line *between* terminators, like `"a\n\nb"`, is preserved as an empty element.) ### What is a Stream? A `java.util.stream.Stream<T>` is a lazy pipeline of elements supporting operations like `filter`, `map`, `count`, `collect`. "Lazy" means nothing happens until a terminal operation (like `count()` or `collect()`) runs, and intermediate operations are fused so the text is traversed minimally. If you need a `List`, call `text.lines().toList()` (Java 16+) or `collect(Collectors.toList())`. ## Example contrasts - `"a\r\nb".split("\n")` -> `["a\r", "b"]` (stray `\r`!) - `"a\r\nb".lines().toList()` -> `["a", "b"]` (clean) ## When to use which Use `lines()` for nearly all line-by-line processing in modern Java — it is correct across platforms, lazy, and stream-friendly. Reach for `split` only when you need to split on a non-line delimiter or specifically want regex behavior. ## Deriving the answer at any level Facts: returns Stream<String>; handles \n/\r/\r\n; strips terminators; lazy; no trailing empty for trailing terminator; safer than split("\n") which leaves stray \r and is eager. Those reconstruct an answer at any level.
- How do you get a List<String> from lines()?Call a terminal collector: text.lines().toList() (Java 16+) or text.lines().collect(Collectors.toList()).
- What does "a\r\nb".split("\n") produce versus lines()?split gives ["a\r", "b"] with a stray carriage return; lines() gives ["a", "b"] cleanly, because it understands \r\n.
saying these in an interview costs you the question
- Saying lines() returns a String[] or List
- Claiming it only splits on \n like split
- Thinking the returned lines still contain the \n/\r
- Believing it is eager rather than lazy