skip to content

When processing a very large or streaming input, how do windowed/chunked/zip behave on a List vs a Sequence, and what are the allocation and correctness implications?

level: seniorimportance: should knowfreq 30%

answer

  1. List = eager, materializes everything
  2. Sequence = lazy, on-demand windows
  3. Transform overload skips sublist/Pair allocation
  4. Window sublist is a transient snapshot
  5. zip still truncates; windowed buffers size elements

basics

~20 s

On a List these operators run eagerly and build all the result lists at once. On a Sequence they are lazy and produce windows/chunks on demand, so you can process huge or streaming data without holding everything in memory.

solid answer

~40 s

On Iterable/List, chunked, windowed, zip, zipWithNext, and flatten are eager: they materialize the full result (a List<List<T>> or List<Pair>) immediately, allocating every sublist/Pair up front. On Sequence the same operators are lazy and return a Sequence, generating each window/chunk only when the terminal operation pulls it — essential for large or infinite/streaming sources. Use the transform overloads (chunked(n) { ... }, windowed(size, step) { ... }, zip(other) { ... }, zipWithNext { ... }) to fold each sublist immediately and skip Pair/sublist allocation. Caveats: the lambda's sublist is a transient snapshot — do not retain it across iterations; zip on two sequences consumes both lazily but still truncates to the shorter; windowed buffers up to size elements internally. Prefer .asSequence() before windowing big data, then a single terminal op.

code

kotlin · 6 lines
kotlin
val firstFiveTriples = generateSequence(1) { it + 1 }   // infinite
    .windowed(3)                                        // lazy sliding triples
    .map { it.sum() }
    .take(5)
    .toList()                                           // [6, 9, 12, 15, 18]
// Works because the Sequence is never fully realized.

go deeper

for a junior

Recognizes Sequence is lazy and List is eager at a high level.

for a middle

Explains when to use asSequence and that transform overloads avoid allocation.

for a senior

Details snapshot lifetime, buffering, truncation on lazy zip, and bounding work with take/first.

for a principal

Designs streaming pipelines with memory/throughput trade-offs, one-shot consumption hazards, and chooses operators (or a manual buffer) for hot paths.

## Eager (Iterable/List) vs lazy (Sequence) All these operators exist in two worlds: - On **`Iterable`/`List`** they are **eager** — the entire result (`List<List<T>>` for windowed/chunked, `List<Pair<A,B>>` for zip) is built and returned immediately, allocating **every** sublist and `Pair` up front. - On **`Sequence`** they are **lazy** — they return a `Sequence` and emit each window/chunk/pair **only when a terminal operation** (e.g. `toList`, `forEach`, `sum`, `first`) pulls it. ```kotlin // Eager: builds the whole List<List<Int>> right away val eager = (1..1_000_000).toList().windowed(3) // Lazy: nothing computed until the terminal op; windows produced on demand val lazy = (1..1_000_000).asSequence() .windowed(3) .map { it.sum() } .take(5) // only 5 windows ever materialized .toList() ``` ## Why laziness matters - **Memory**: with a `List`, `windowed(3)` over a million items allocates ~a million 3-element lists at once. With a `Sequence` and an early `take`/`first`, only the windows you consume are built. - **Streaming / unbounded**: a `Sequence` (e.g. from `generateSequence`) can be windowed/zipped without ever existing fully in memory; an eager `List` cannot represent an infinite source. ## Cut allocation with transform overloads Each operator has a `transform` overload that folds the sublist/pair immediately, so **no intermediate `List`/`Pair` is allocated**: ```kotlin seq.windowed(3) { it.sum() } // emits Int, not List<Int> seq.chunked(1000) { it.average() } a.zip(b) { x, y -> x + y } seq.zipWithNext { p, n -> n - p } ``` ## Correctness caveats - **Snapshot lifetime**: the sublist passed to a `windowed`/`chunked` transform is a **transient view valid only during that call** — never stash the reference and read it later; copy with `.toList()` if you must retain it. - **zip still truncates** on sequences: it stops at the shorter side and consumes both lazily in lockstep. - **Buffering**: `windowed(size, step)` must buffer up to `size` elements internally even when lazy; a huge `size` still costs that much memory per window. - **Single consumption**: a `Sequence` is typically one-shot; windowing it then iterating twice re-runs the upstream. ## Rule of thumb For big or streaming data: `.asSequence()` first, prefer the **transform overload**, and end with **one** terminal operation; reach for an early `take`/`first` to bound work.

  • Why can you windowed() an infinite generateSequence but not an infinite List?
    windowed on a Sequence is lazy and pulls elements on demand, so a terminal op like take(5) bounds the work. An infinite List cannot exist in memory to call the eager overload on.
  • What's the risk of storing the sublist passed to windowed(size){ ... } for later use?
    It is a transient snapshot reused/invalidated across iterations; retain a copy via toList() instead of keeping the reference.

Eager is photocopying the whole book into stacks before reading; lazy is reading page-by-page and only copying the pages you actually need.

saying these in an interview costs you the question

  • Believing windowed/chunked on a plain List are lazy
  • Retaining the transform's sublist reference across iterations
  • Thinking lazy zip stops truncating to the shorter input
  • Assuming Sequence laziness removes the per-window buffering cost
  • Forgetting a Sequence is one-shot and re-iterating re-runs upstream work

context