skip to content

Compare limit/skip with takeWhile/dropWhile (Java 9). How do they behave on ordered vs unordered and infinite streams?

level: seniorimportance: should knowfreq 52%

answer

  1. limit/skip = positional; takeWhile/dropWhile = prefix by predicate
  2. limit + takeWhile short-circuit (bound infinite streams)
  3. dropWhile keeps later matches; filter keeps matches anywhere
  4. skip/dropWhile leave an infinite tail infinite
  5. 'first n' / 'prefix' is meaningless without encounter order

basics

~20 s

limit(n) keeps the first n elements and skip(n) drops the first n, both by position. takeWhile keeps elements from the start as long as a condition is true and stops at the first failure; dropWhile drops the leading run that matches and keeps the rest. limit and takeWhile can stop a stream early, so they work on infinite streams; skip and dropWhile generally must process the leading part.

solid answer

~50 s

limit(n) and skip(n) are positional: limit keeps the first n elements (short-circuiting — it can stop the source early), skip discards the first n and passes the rest. takeWhile(predicate) and dropWhile(predicate) (Java 9) are condition-based on a prefix: takeWhile emits the leading run where the predicate holds and stops at the first false (short-circuiting); dropWhile discards that leading run and emits everything after, including later elements that wouldn't match. The crucial difference from filter is that takeWhile/dropWhile look only at the contiguous prefix, not every element. On infinite streams, limit and takeWhile are the standard ways to bound them; skip and dropWhile only help if the kept part is finite. Ordering matters: on an ordered stream these ops use encounter order deterministically; on an unordered (or parallel-unordered) stream, 'first n' or 'the prefix' is not deterministic, so limit/takeWhile may return any valid subset and incur less coordination cost. In ordered parallel streams they add overhead because the framework must respect position.

code

java · 15 lines
java
var data = Stream.of(2, 4, 6, 7, 8, 10);

// positional
Stream.of(2,4,6,7,8,10).skip(2).limit(3);   // 6, 7, 8

// prefix by predicate (Java 9)
Stream.of(2,4,6,7,8,10).takeWhile(n -> n % 2 == 0); // 2, 4, 6  (stop at 7)
Stream.of(2,4,6,7,8,10).dropWhile(n -> n % 2 == 0); // 7, 8, 10 (keep 8,10!)

// contrast with filter (every element, anywhere)
Stream.of(2,4,6,7,8,10).filter(n -> n % 2 == 0);    // 2, 4, 6, 8, 10

// bounding an infinite stream
Stream.iterate(1, x -> x + 1).takeWhile(x -> x < 5).toList(); // [1,2,3,4]
Stream.iterate(1, x -> x + 1).limit(4).toList();              // [1,2,3,4]

go deeper

for a junior

Knows limit keeps the first n and skip drops the first n; can use them for simple 'take 10' cases.

for a middle

Correctly distinguishes takeWhile/dropWhile (prefix by predicate) from filter (everywhere) and knows limit/takeWhile can bound infinite streams.

for a senior

Explains short-circuiting precisely, the dropWhile-keeps-later-matches subtlety, and how ordered vs unordered and parallel execution affect determinism and cost.

for a principal

Designs pipelines that exploit short-circuiting on large/infinite sources, chooses ordered vs unordered deliberately for parallel performance, and recognizes when skip-based pagination is O(skipped) and a different access pattern is warranted.

## Two families, two selection strategies All four operations *select a contiguous portion* of the stream, but they differ in how they decide where the cut is. ### Positional: limit / skip - `limit(long n)`: keep the **first n** elements, drop the rest. It is **short-circuiting** — once n elements have passed, it tells the upstream to stop, so it can bound an infinite stream. - `skip(long n)`: drop the **first n** elements, pass everything after. It must traverse (and discard) those first n; it does not short-circuit the tail. They decide purely by **count/position**, ignoring element values. ```java Stream.iterate(1, x -> x + 1) // infinite: 1,2,3,... .skip(2) // drop 1,2 .limit(3) // keep next 3 -> 3,4,5 .forEach(System.out::println); ``` ### Condition-based prefix: takeWhile / dropWhile (Java 9+) - `takeWhile(Predicate p)`: emit elements from the start **as long as p is true**; the moment p returns false, **stop** and emit nothing further. Short-circuiting. - `dropWhile(Predicate p)`: **discard** the leading run where p is true; once p first returns false, emit that element and **all** remaining elements — even ones that would have matched p. ```java Stream.of(2, 4, 6, 7, 8, 10) .takeWhile(n -> n % 2 == 0); // 2, 4, 6 (stops at 7) Stream.of(2, 4, 6, 7, 8, 10) .dropWhile(n -> n % 2 == 0); // 7, 8, 10 (keeps 8 and 10 even though even) ``` ## takeWhile/dropWhile vs filter — a key distinction `filter` tests **every** element and keeps all matches anywhere in the stream. `takeWhile`/`dropWhile` only consider the **contiguous leading prefix** and then stop testing. - `filter(even)` on `2,4,6,7,8,10` -> `2,4,6,8,10` (all evens, anywhere). - `takeWhile(even)` -> `2,4,6` (stops at the first odd). - `dropWhile(even)` -> `7,8,10` (drops the leading evens, keeps the rest verbatim). This makes `takeWhile`/`dropWhile` ideal for **sorted/ordered data** where the predicate marks a boundary (e.g. "take log lines until the first ERROR"). ## Infinite streams | Op | Bounds an infinite stream? | |----|----------------------------| | `limit(n)` | **Yes** — short-circuits after n | | `takeWhile(p)` | **Yes** *if* p eventually becomes false | | `skip(n)` | No — the tail is still infinite | | `dropWhile(p)` | No — after dropping the prefix, the rest is still infinite | So `limit` and `takeWhile` are the tools to tame infinite generators; `skip`/`dropWhile` only reposition the start. ## Ordered vs unordered, and parallelism These operations are **order-sensitive by definition** — "first n" and "the leading prefix" only make sense relative to an **encounter order**. - On an **ordered** stream they are **deterministic**: you always get the same elements. - On an **unordered** stream (e.g. one sourced from a `HashSet`, or after `unordered()`), there is no defined first element, so `limit`/`takeWhile`/`dropWhile` may legally return **any** valid subset of the right size/shape — results can vary between runs. - In **parallel ordered** streams, `limit` and `skip` add **coordination overhead**: the framework must figure out positions across threads to honor encounter order, which can serialize work and hurt the speedup. Marking the stream `unordered()` (when correctness allows) relaxes this and improves parallel performance. `takeWhile`/`dropWhile` in parallel are also more expensive because the prefix boundary depends on order; they are cheapest on sequential ordered streams. ## Practical guidance - Use `limit` to cap infinite/large streams and for pagination-style 'first N'. - Use `skip`+`limit` together for windowing/pagination (`skip(page*size).limit(size)`), aware that `skip` still traverses the skipped prefix. - Reach for `takeWhile`/`dropWhile` on **sorted** data when a predicate marks a natural cut point — it short-circuits where `filter` would scan everything. - Be deliberate about ordering: don't expect deterministic 'first n' from an unordered source, and consider `unordered()` to speed up parallel `limit`/`distinct`.

  • On the stream [2,4,6,7,8,10], why does dropWhile(even) yield [7,8,10] and not [7]?
    dropWhile discards only the leading contiguous run where the predicate is true (2,4,6), then stops testing entirely and emits everything from the first failing element onward. So 7 is the first kept element, and 8 and 10 are kept verbatim even though they're even — dropWhile does not re-apply the predicate after the prefix, unlike filter.
  • Why can limit() hurt the performance of a parallel ordered stream?
    limit must return the first n elements in encounter order. In parallel, work is split across threads, so the framework must coordinate to determine which elements are positionally first and stop the rest, adding synchronization and reducing parallelism. On an unordered stream (or after unordered()), it can take any n elements with far less coordination, so it's cheaper.

limit/skip are like 'serve the first 5 customers' / 'skip the first 5' regardless of who they are. takeWhile/dropWhile are like 'serve customers while it's still morning' — you stop (or start) at a boundary defined by a condition, not a count, and once the boundary passes you don't re-check it.

saying these in an interview costs you the question

  • Saying takeWhile/dropWhile test every element like filter — they only consider the leading prefix
  • Expecting dropWhile to remove all matching elements (it keeps later matches)
  • Thinking skip or dropWhile can bound an infinite stream (they can't)
  • Assuming 'first n' is deterministic on an unordered or parallel stream
  • Believing limit is free in parallel ordered pipelines

context