You own a shared data-processing library. How do you decide your public API returns List vs Sequence, and what pitfalls do you warn callers about?
answer
- Return type = contract on timing + reuse
- List: small, reusable, eager-failure, snapshot
- Sequence: large/infinite, composable, short-circuit
- Warn: lazy timing, single-use, recompute, accidental sorted()
- Async stream -> Flow, not Sequence
basics
~20 sReturn a List when results are small, finite, and reused; return a Sequence when data may be large or infinite and callers can stream and stop early. Warn that Sequences are lazy and may be single-use.
solid answer
~50 sChoose the return type by how data flows. Return an eager `List` (or `Set`) when the result set is bounded and small, callers will iterate it multiple times, or you must guarantee the work is already done (no surprise lazy exceptions later). Return a `Sequence` when the source is large/unbounded, you want callers to compose more operators without forcing intermediate allocations, or short-circuiting matters. Pitfalls to document: (1) Sequences are **lazy** — exceptions and side effects fire at the terminal op, not at call time; (2) some Sequences are **single-use** (builder-based or `constrainOnce()`) and throw on re-iteration; (3) re-iterating a non-cached Sequence **recomputes** everything; (4) terminal ops like `sorted`/`toList` defeat laziness. Often the cleanest contract is to return `Sequence` for streaming pipelines but expose convenience eager overloads, and never leak a one-shot Sequence where callers expect a reusable collection.
go deeper
Knows List is eager and Sequence is lazy and that the choice matters for callers.
Picks List for small reusable results and Sequence for large/streaming ones with basic rationale.
Articulates reuse, eager-failure, and short-circuit trade-offs and the single-use/recompute pitfalls.
Treats the return type as an API contract, offers layered eager/lazy overloads, and routes async streams to Flow with documented semantics.
## Framing the API choice The return type is a **contract about evaluation timing and reuse**, not just a container choice. ### Return an eager `List`/`Set` when: - The result is **bounded and small** — callers gain nothing from laziness. - Callers will **iterate multiple times** (a Sequence would recompute or throw). - You need **eager failure**: validation/IO errors should surface at call time, not deferred to a downstream terminal op. - The result is a **snapshot** that must not change if the source mutates. ### Return a `Sequence` when: - The source is **large or unbounded** (`generateSequence`, streaming IO). - You want callers to **compose further** `map`/`filter` without forcing intermediate lists. - **Short-circuiting** is valuable (callers may `take`/`first`). - You're modeling a pull-based stream and want backpressure-like, on-demand evaluation. ## Pitfalls to document for callers ```kotlin // 1) Laziness defers side effects/exceptions to the terminal op fun rows(): Sequence<Row> = source.asSequence().map { parse(it) } // parse throws at terminal, not here // 2) Single-use sequences val s = sequence { yield(1); yield(2) } s.toList() // ok s.toList() // IllegalStateException: already consumed (constrainOnce semantics) ``` - **Lazy timing:** Document that nothing runs until a terminal op; failures appear there. - **Single-use:** Builder sequences and `constrainOnce()` throw on re-iteration. If your Sequence is multi-pass-safe, say so explicitly. - **Recomputation:** A multi-pass Sequence re-runs the whole pipeline each iteration — expensive or non-deterministic if it reads IO/clock. - **Thread-safety / mutation:** A Sequence over a mutable source can observe concurrent changes; a returned `List` is a stable snapshot. - **Accidental materialization:** A caller adding `sorted()`/`groupingBy` silently buffers everything, erasing the streaming benefit. ## Pragmatic library patterns - Offer a `Sequence`-returning core for composition plus eager convenience overloads (`...List()` / `...Set()`). - Prefer returning `List` from small, frequently re-read accessors; reserve `Sequence` for genuine pipelines. - If returning a Sequence that must be re-iterable, back it by a stored collection (`elements.asSequence()`), not a one-shot builder. - Consider Kotlin `Flow` instead when the stream is **asynchronous/suspending** — Sequences are synchronous and blocking. - Name and KDoc the laziness explicitly so callers don't assume eager semantics. ## Decision summary Default to `List` for small, reusable, must-be-eager results; choose `Sequence` for large/streaming/composable/short-circuiting ones; reach for `Flow` when async. Always document evaluation timing and reuse guarantees.
- When would you return a Flow instead of a Sequence?When elements are produced asynchronously or via suspending work (network, DB streaming) — Flow supports suspension and structured concurrency; Sequence is synchronous and blocking.
- How do you make a returned Sequence safe to iterate multiple times?Back it with a stored collection and return collection.asSequence(), or materialize once internally — never return a one-shot builder/constrainOnce sequence to callers who expect reuse.
saying these in an interview costs you the question
- Returning a one-shot builder Sequence as a reusable collection
- Ignoring that exceptions fire at the terminal op
- Defaulting all APIs to Sequence for 'performance'
- Confusing Sequence with Flow for async streams
- Not documenting evaluation timing or reuse semantics