Why does peek sometimes print fewer elements than the source contains, for example in stream.peek(...).limit(2).collect(...)? Explain in terms of the execution model.
answer
- peek runs per CONSUMED element, not per source element
- limit/findFirst/anyMatch/takeWhile short-circuit the pull
- Lazy pull stops once downstream has enough
- Infinite stream + limit works because of this
- peek count depends on its position + downstream selectivity
basics
~20 sBecause peek only runs for elements that are actually pulled through the pipeline. A short-circuiting op like limit(2) stops after two elements are produced, so the engine never asks the source for more, and peek never sees the rest.
solid answer
~40 speek is a per-element intermediate operation, so its action runs only for elements that the pipeline actually consumes. Streams pull elements lazily, one at a time, and short-circuiting terminal/intermediate operations like limit, findFirst, anyMatch, or takeWhile can stop the pull early. With peek(...).limit(2), the engine pulls element 1 (peek prints it), passes it down, pulls element 2 (peek prints it), reaches the limit of 2, and stops — it never pulls elements 3+, so peek never prints them. This is a direct consequence of laziness plus fusion: there is no eager loop over the whole source. The same reasoning explains why findFirst can make an expensive map run only once, and why an infinite stream (Stream.iterate / generate) works fine as long as a downstream limit bounds it.
go deeper
Recognizes that peek only sees elements that pass through and that limit can cut traversal short.
Explains lazy pull + short-circuit, names short-circuiting ops, and walks through peek().limit(2) showing later elements are never pulled.
Reasons about how peek's invocation count depends on its position and downstream selectivity, why infinite streams work, and why peek is unreliable for side effects per the Javadoc.
Discusses cancellation propagation and short-circuit semantics across sequential and parallel pipelines, and the API-design rationale for marking peek debug-only.
## Setup: what peek is `peek(Consumer)` is an **intermediate** operation that, for each element flowing through it, runs the given consumer (typically for debugging/logging) and then passes the element along unchanged. Being intermediate, it is **lazy** and runs *only as elements are pulled through it*. ## The pull/short-circuit model Streams do not eagerly loop over the entire source. The terminal operation drives a **demand-driven traversal**: it pulls elements one at a time, pushing each through the fused chain. Crucially, some operations are **short-circuiting** — they can declare "I have enough; stop pulling": - `limit(n)` — emit at most n, then stop the upstream. - `findFirst` / `findAny` — stop at the first match. - `anyMatch` / `allMatch` / `noneMatch` — stop as soon as the answer is decided. - `takeWhile` — stop at the first element failing the predicate. When such an op decides it is done, the engine signals **cancellation** upstream and stops pulling new elements from the source. ## Walking through peek(...).limit(2) Consider a 5-element source `[a,b,c,d,e]` and the pipeline `source.peek(print).limit(2).collect(toList())`: 1. Terminal `collect` asks for an element. Source yields `a`. `peek` prints `a`. `limit` accepts it (count 1). 2. Asks again. Source yields `b`. `peek` prints `b`. `limit` accepts it (count 2) — **limit reached**. 3. `limit` signals upstream that no more elements are needed. The engine **stops pulling**. `c`, `d`, `e` are never requested → `peek` is **never called** for them. Result: only `a` and `b` are printed, even though the source had five elements. peek printing "fewer than the source size" is therefore *expected* and is direct evidence that the pipeline is lazy and short-circuiting — not a bug. ## Why ordering of operations matters If instead you wrote `source.limit(2).peek(print)`, the limit happens *before* peek in the chain, so peek still only ever sees 2 elements. But `source.peek(print).filter(...).limit(2)` could print *more* than 2 elements: peek runs on every element pulled, and the pipeline must keep pulling until `filter` has passed 2 elements through to `limit`. So the count peek prints depends on where peek sits and how selective the upstream filter is. ## Infinite streams: the same principle, flipped Because of lazy pull + short-circuit, an *infinite* source is usable: ``` Stream.iterate(0, n -> n + 1) // infinite .peek(System.out::println) .limit(3) .collect(Collectors.toList()); // prints 0,1,2 then stops ``` Without laziness this would loop forever. The `limit(3)` bounds the pull, so only 3 elements are ever generated and peeked. ## The trap people fall into Using `peek` to count or to drive side effects is unreliable precisely because of this: the number of times peek runs equals the number of elements *consumed*, which depends on downstream short-circuiting, not on the source size. The Javadoc explicitly warns that `peek` is mainly for debugging and that an implementation may even skip it if it can prove the element's value isn't needed. ## Summary peek prints fewer than the source size when a downstream short-circuiting op (limit/findFirst/anyMatch/takeWhile) stops the lazy pull early. peek runs *per consumed element*, not per source element.
- Why is using peek to mutate state or count elements considered an anti-pattern?Because peek only runs for elements actually consumed, and the JVM may even elide it when the value is provably unneeded. The number of invocations depends on downstream short-circuiting, so any count or side effect is unreliable; use map/reduce/collect for real work.
- In source.peek(p).filter(f).limit(2), can peek print more than 2 elements? Why?Yes. peek runs on every element pulled, and the pipeline keeps pulling until filter has passed 2 elements to limit. If many elements fail the filter, peek prints all of those too — possibly many more than 2.
saying these in an interview costs you the question
- Assuming peek always runs once per source element
- Using peek to count or for required side effects
- Thinking limit filters AFTER processing the whole source
- Believing an infinite stream must hang regardless of limit