How does a Stream differ from a Collection in Java?
answer
- Collection stores; Stream computes
- Stream carries no storage — it's a view over a source
- Collection = reusable; Stream = one-shot
- External iteration (loop) vs internal iteration (library)
- Streams can be lazy, parallel, even infinite
basics
~20 sA Collection stores data in memory and you can access it many times. A Stream stores nothing — it just describes a computation over a source and can only be used once, after which it is finished.
solid answer
~50 sA Collection is a data structure: it holds elements in memory, you can add, remove, index, and traverse them repeatedly. A Stream is not a data structure — it carries no storage. It is a one-shot view over some source (which may be a collection, an array, a file, or a generator) that describes a computation to perform on those elements. Collections are about *storing and managing* data; streams are about *expressing computations* over it declaratively. Because a stream holds nothing, it can be lazy, can be processed in parallel without you managing threads, and can even be backed by an infinite source. A collection is finite, eager, and externally iterated (you pull elements with a loop), whereas a stream is internally iterated (the library pulls for you) and runs only when a terminal operation is applied — once.
go deeper
States that a collection stores data and a stream does not, and that a stream is used once.
Contrasts storage vs computation, single-use vs reusable, and lazy vs eager, and knows how to convert between the two.
Explains internal vs external iteration and how it enables fusion and parallelism, and the view-not-snapshot semantics with its concurrency hazard.
Reasons about when stream abstraction overhead is worth it vs a plain loop, and the API design rationale (Stream as a separate type rather than methods on Collection).
## Two different jobs The `Collection` interface (`List`, `Set`, `Map.values()`, etc.) and the `Stream` interface look superficially similar — both are 'a bunch of elements you do things to' — but they exist for *opposite* purposes: - A **Collection** is a **data structure**: its job is to *store* and *organize* elements in memory so you can put them in, take them out, look them up, and walk over them as many times as you like. - A **Stream** is a **computation pipeline**: its job is to *describe a sequence of operations* to perform on elements coming from some source. It is a means of *processing*, not *storing*. ## Concrete differences | Aspect | Collection | Stream | |---|---|---| | Storage | Holds elements in memory | Holds nothing; pulls from a source | | Reuse | Traverse/iterate any number of times | One-shot — consumed by its terminal op | | Evaluation | Eager — elements exist now | Lazy — computed only when terminal runs | | Iteration | External (you write the loop) | Internal (the library iterates) | | Size | Finite, known | May be infinite (generator sources) | | Mutation | Add/remove/set elements | Cannot modify the source elements | ## What 'carries no storage' means When you write `list.stream()`, you do **not** copy the list. You get a lightweight object that knows *how to ask the list for its elements* and what operations to apply. The data still lives in the list. This is why a stream is cheap to create and why mutating the backing collection while a stream over it is being processed is unsafe (it can throw `ConcurrentModificationException` or give undefined results — the stream is a *view*, not a snapshot). ## Internal vs external iteration With a collection you iterate **externally**: ```java int sum = 0; for (int n : list) { // you control the loop if (n > 0) sum += n; } ``` With a stream you iterate **internally** — you say *what* to do and the library controls *how* and *when* to loop: ```java int sum = list.stream() .filter(n -> n > 0) .mapToInt(Integer::intValue) .sum(); ``` Internal iteration is what lets streams transparently optimize (fuse operations, short-circuit) and parallelize (`list.parallelStream()`) without you rewriting the loop. ## Going between the two - **Collection → Stream:** `collection.stream()` or `parallelStream()`. - **Stream → Collection:** a terminal `collect(Collectors.toList())` / `toList()` / `Collectors.toSet()` materializes the result back into a data structure. The mental model: *collection = noun (the data); stream = verb (the processing).* Use a collection when you need to keep and revisit data; use a stream to express a transformation/aggregation over it once. ## Terms defined - **Eager evaluation:** results are computed as soon as the operation is invoked. - **Lazy evaluation:** computation is deferred until the result is actually demanded. - **Internal iteration:** the looping logic lives inside the library, not your code. - **View:** an object that exposes another object's data without copying it.
- Why can a Stream be infinite but a Collection cannot?A Collection must physically store all its elements in memory, so it must be finite. A Stream stores nothing and is lazy — it pulls elements on demand from a generator like Stream.iterate, so it can represent an unbounded sequence as long as a short-circuiting op (limit, findFirst) eventually stops it.
- Is a stream a snapshot of the collection at the time stream() was called?No. It is a lazy view, not a copy. Elements are pulled from the source when the terminal op runs, so structurally modifying the source between stream() and the terminal op generally causes a ConcurrentModificationException or undefined behavior.
saying these in an interview costs you the question
- Saying stream() copies the collection's elements into the stream
- Treating a stream as a data structure you can re-traverse
- Claiming streams can mutate (add/remove) elements in the source
- Believing collections can be infinite or that streams must be finite