skip to content

Why does peek sometimes print fewer elements than the source contains, for example in stream.peek(...).limit(2).collect(...)? Explain in terms of the execution model.

level: middleimportance: should knowfreq 58%

answer

  1. peek runs per CONSUMED element, not per source element
  2. limit/findFirst/anyMatch/takeWhile short-circuit the pull
  3. Lazy pull stops once downstream has enough
  4. Infinite stream + limit works because of this
  5. peek count depends on its position + downstream selectivity

basics

~20 s

Because peek only runs for elements that are actually pulled through the pipeline. A short-circuiting op like limit(2) stops after two elements are produced, so the engine never asks the source for more, and peek never sees the rest.

solid answer

~40 s

peek is a per-element intermediate operation, so its action runs only for elements that the pipeline actually consumes. Streams pull elements lazily, one at a time, and short-circuiting terminal/intermediate operations like limit, findFirst, anyMatch, or takeWhile can stop the pull early. With peek(...).limit(2), the engine pulls element 1 (peek prints it), passes it down, pulls element 2 (peek prints it), reaches the limit of 2, and stops — it never pulls elements 3+, so peek never prints them. This is a direct consequence of laziness plus fusion: there is no eager loop over the whole source. The same reasoning explains why findFirst can make an expensive map run only once, and why an infinite stream (Stream.iterate / generate) works fine as long as a downstream limit bounds it.

go deeper

for a junior

Recognizes that peek only sees elements that pass through and that limit can cut traversal short.

for a middle

Explains lazy pull + short-circuit, names short-circuiting ops, and walks through peek().limit(2) showing later elements are never pulled.

for a senior

Reasons about how peek's invocation count depends on its position and downstream selectivity, why infinite streams work, and why peek is unreliable for side effects per the Javadoc.

for a principal

Discusses cancellation propagation and short-circuit semantics across sequential and parallel pipelines, and the API-design rationale for marking peek debug-only.

## Setup: what peek is `peek(Consumer)` is an **intermediate** operation that, for each element flowing through it, runs the given consumer (typically for debugging/logging) and then passes the element along unchanged. Being intermediate, it is **lazy** and runs *only as elements are pulled through it*. ## The pull/short-circuit model Streams do not eagerly loop over the entire source. The terminal operation drives a **demand-driven traversal**: it pulls elements one at a time, pushing each through the fused chain. Crucially, some operations are **short-circuiting** — they can declare "I have enough; stop pulling": - `limit(n)` — emit at most n, then stop the upstream. - `findFirst` / `findAny` — stop at the first match. - `anyMatch` / `allMatch` / `noneMatch` — stop as soon as the answer is decided. - `takeWhile` — stop at the first element failing the predicate. When such an op decides it is done, the engine signals **cancellation** upstream and stops pulling new elements from the source. ## Walking through peek(...).limit(2) Consider a 5-element source `[a,b,c,d,e]` and the pipeline `source.peek(print).limit(2).collect(toList())`: 1. Terminal `collect` asks for an element. Source yields `a`. `peek` prints `a`. `limit` accepts it (count 1). 2. Asks again. Source yields `b`. `peek` prints `b`. `limit` accepts it (count 2) — **limit reached**. 3. `limit` signals upstream that no more elements are needed. The engine **stops pulling**. `c`, `d`, `e` are never requested → `peek` is **never called** for them. Result: only `a` and `b` are printed, even though the source had five elements. peek printing "fewer than the source size" is therefore *expected* and is direct evidence that the pipeline is lazy and short-circuiting — not a bug. ## Why ordering of operations matters If instead you wrote `source.limit(2).peek(print)`, the limit happens *before* peek in the chain, so peek still only ever sees 2 elements. But `source.peek(print).filter(...).limit(2)` could print *more* than 2 elements: peek runs on every element pulled, and the pipeline must keep pulling until `filter` has passed 2 elements through to `limit`. So the count peek prints depends on where peek sits and how selective the upstream filter is. ## Infinite streams: the same principle, flipped Because of lazy pull + short-circuit, an *infinite* source is usable: ``` Stream.iterate(0, n -> n + 1) // infinite .peek(System.out::println) .limit(3) .collect(Collectors.toList()); // prints 0,1,2 then stops ``` Without laziness this would loop forever. The `limit(3)` bounds the pull, so only 3 elements are ever generated and peeked. ## The trap people fall into Using `peek` to count or to drive side effects is unreliable precisely because of this: the number of times peek runs equals the number of elements *consumed*, which depends on downstream short-circuiting, not on the source size. The Javadoc explicitly warns that `peek` is mainly for debugging and that an implementation may even skip it if it can prove the element's value isn't needed. ## Summary peek prints fewer than the source size when a downstream short-circuiting op (limit/findFirst/anyMatch/takeWhile) stops the lazy pull early. peek runs *per consumed element*, not per source element.

  • Why is using peek to mutate state or count elements considered an anti-pattern?
    Because peek only runs for elements actually consumed, and the JVM may even elide it when the value is provably unneeded. The number of invocations depends on downstream short-circuiting, so any count or side effect is unreliable; use map/reduce/collect for real work.
  • In source.peek(p).filter(f).limit(2), can peek print more than 2 elements? Why?
    Yes. peek runs on every element pulled, and the pipeline keeps pulling until filter has passed 2 elements to limit. If many elements fail the filter, peek prints all of those too — possibly many more than 2.

saying these in an interview costs you the question

  • Assuming peek always runs once per source element
  • Using peek to count or for required side effects
  • Thinking limit filters AFTER processing the whole source
  • Believing an infinite stream must hang regardless of limit

context