What is the difference between external (pull) and internal (push) iteration, and what are the trade-offs of each?
answer
- who owns the loop?
- external = pull, break works, cursors mergeable
- internal = push, parallel + fusion + cleanup
- inverted control, Hollywood principle
- generators: write push, expose pull
basics
~20 sWith external iteration the caller drives the loop and pulls elements one by one (while (it.hasNext()) ...). With internal iteration the caller hands a function to the collection (forEach { ... }) and the collection runs the loop, pushing each element into that function.
solid answer
~50 sExternal iteration: the client owns the loop and asks for the next element, so the client controls pacing and can stop early, `break`, `return`, interleave two sequences (merge/zip), or hold the cursor across method calls. Cost: more boilerplate, and the traversal strategy is fixed as sequential — the client sees only "one at a time". Internal iteration: the client passes a callback (`forEach`, `map`, `each`) and the aggregate runs the loop. Because the library owns the loop, it can choose the order, parallelize, fuse several operations into a single pass, manage resources (open/close a file, hold a lock), and hide recursion for trees. Cost: early exit and multi-sequence interleaving become awkward (you need exceptions, sentinel return values, or dedicated short-circuit operations like `anyMatch`/`find`), and control flow is inverted, which complicates debugging, checked exceptions and mutable local state. Most modern libraries offer both: an external `Iterator` primitive plus internal combinators layered on top.
code
pseudocode · 11 lines// External (pull): client owns the loop; break is trivial
it = users.iterator()
while (it.hasNext()) { u = it.next(); if (u.isAdmin) return u }
// Internal (push): library owns the loop; early exit needs a convention
users.forEach { u -> if (u.isAdmin) /* ...how do I return? */ }
found = users.first { u -> u.isAdmin } // short-circuiting combinator
// Merging two sequences: only external iteration expresses this directly
a = xs.iterator(); b = ys.iterator()
while (a.hasNext() && b.hasNext()) { /* advance whichever is behind */ }go deeper
Define both with one line of code each and say who runs the loop; mention break works naturally only in the external form.
Give at least two concrete advantages per side — early exit and merging two sequences for external; parallelism, fusion and resource cleanup for internal — and note that most libraries provide both.
Discuss pull-vs-push as a backpressure question, why external→internal is easy and internal→external needs coroutines, and how generators give you both.
Frame it as a published-contract decision across services: pull cursors give consumers pacing and resumability, push/streaming needs an explicit credit protocol; choosing wrong shows up as either N+1 chattiness or an overwhelmed consumer under load.
## Two directions of control The question is simply: **who owns the loop?** ### External iteration (pull) ``` it = collection.iterator() while (it.hasNext()) { e = it.next() if (bad(e)) break // client decides to stop use(e) } ``` The client calls the iterator. The iterator is *passive*; nothing happens until asked. This is also called a **pull model**: elements are pulled out on demand. ### Internal iteration (push) ``` collection.forEach { e -> use(e) } ``` The client hands over a function. The collection runs its own loop and *pushes* each element into the function. The client's code is now a callback — control is inverted ("don't call us, we'll call you", the Hollywood Principle). ## Why internal iteration exists — what the library can do once it owns the loop 1. **Choose the traversal strategy.** Order can be reversed, chunked, or unordered. Most importantly the library can **parallelize**: split the source into chunks and run them on several cores. A `while (hasNext)` loop is inherently sequential — the client's loop body cannot be split by the library. (Java's `Spliterator`, with its `trySplit`, exists precisely to give the library a splittable source instead of a strictly sequential one.) 2. **Fuse passes.** `filter(...).map(...).sum()` expressed as internal operations can be executed in a single traversal with no intermediate collections; naively written external loops often build a temporary list per step. 3. **Manage resources and locks correctly.** `withEachLine(file) { ... }` opens the file, guarantees the loop runs, and closes the file in a `finally` — the client cannot leak the handle by forgetting to close or by exiting early. Similarly a collection can hold a lock for the whole traversal. 4. **Hide complex traversal.** Walking a tree or graph internally is a plain recursive function. Exposing the *same* walk externally requires materializing the recursion into an explicit stack/queue inside the iterator object (or a coroutine), which is real work. 5. **Less boilerplate and fewer off-by-one/typo bugs** (the classic external-iteration bug is calling `next()` twice in one loop body). ## Why external iteration survives — what the client keeps 1. **Early exit is natural.** `break`, `return`, `throw` all work. With a callback you need a boolean return convention (`each { return false to stop }`), an exception used as control flow, or a purpose-built short-circuiting combinator (`find`, `anyMatch`, `takeWhile`). 2. **Interleaving multiple sequences.** A merge of two sorted inputs, a zip, or a diff needs to advance *whichever* cursor is behind. With two external iterators this is trivial; with two internal loops it is not expressible without buffering one side or spawning threads. 3. **Pacing / backpressure.** The consumer decides when to take the next element, so a slow consumer naturally throttles a fast producer. In a pure push model, a fast producer can overwhelm a slow consumer unless the protocol adds an explicit request signal (this is exactly what Reactive Streams' `request(n)` re-introduces: async push *with* pull-style backpressure). 4. **Suspendable state.** The iterator is an object you can store in a field, pass to another method, or park while you wait for I/O — a paused position. A callback loop's position is on the stack and cannot be handed around. 5. **Debuggability and plain control flow.** Stack traces stay shallow, step-debugging is linear, mutable locals and (in some languages) checked exceptions inside the loop body work without ceremony. ## The synthesis used by real libraries Most ecosystems keep external iteration as the low-level primitive and layer internal operations on top: Java's `Iterator` plus `forEach`/Streams; C#'s `IEnumerator` plus LINQ; Python's iterator protocol plus comprehensions and `itertools`; Ruby's `each` (internal-first) plus `Enumerator` to externalize it. Note the conversion direction: turning an external iterator into internal iteration is trivial (`while (hasNext) f(next())`); turning an internal `each` into an external cursor requires a coroutine, generator, thread or buffering — which is why libraries whose only primitive is `each` need machinery (Ruby's `Enumerator`, generator functions) to go the other way. **Generators/coroutines** (`yield`) blur the line usefully: you *write* the traversal as an internal-style recursive push loop, and the compiler produces an external pull-style iterator by turning the function into a resumable state machine. You get readable traversal code and consumer-controlled pacing. ## Choosing - Need early exit, merging of several sequences, or consumer-paced consumption → external. - Need parallelism, single-pass fusion, guaranteed resource cleanup, or you are hiding a gnarly recursive walk → internal. - Publishing a library API → offer the external primitive (it is the more fundamental one) and add internal convenience operations, including short-circuiting ones so users are not forced back to exceptions for early exit.
- Why is it easy to build internal iteration on top of an external iterator, but hard to go the other way?Internal-from-external is a three-line loop calling the callback. External-from-internal must *pause* the library's loop between elements and hand control back — the position lives on the library's call stack, so you need a coroutine/generator, a separate thread with a handoff queue, or full buffering of the results.
- Which model gives you backpressure, and how do push-based systems get it back?Pull gives it for free: the consumer only asks when ready. Push-based systems re-introduce it explicitly — Reactive Streams' `request(n)` means the subscriber tells the publisher how many elements it can currently absorb, making it push-with-credit rather than uncontrolled push.
- Why does parallel execution favour internal iteration?Because the library must be free to split the source and run chunks concurrently. With `hasNext`/`next` the client's sequential loop is the schedule, so the library has no place to inject splitting; with a callback plus a splittable source it can partition, run, and combine.
External iteration is a buffet — you take one plate at a time and stop whenever you're full. Internal iteration is a set menu with waiters — the kitchen decides the order and pace and can serve many tables in parallel, but leaving halfway through the courses is awkward.
saying these in an interview costs you the question
- Saying internal iteration is "just syntactic sugar for a for-loop" — it enables parallelism, operation fusion and guaranteed resource cleanup that a client-owned loop cannot provide.
- Claiming you cannot exit early from internal iteration at all — you can, via short-circuiting operations (`find`, `anyMatch`, `takeWhile`) or a documented boolean-return convention; it is just not plain `break`.
- Assuming internal iteration is always faster — for a small in-memory collection the callback/lambda and pipeline setup can be slower than a simple indexed loop.
- Confusing internal iteration with asynchrony; a `forEach` is normally synchronous and blocking.
- Believing a push API automatically handles a slow consumer — without an explicit request/credit signal, push floods the consumer or forces unbounded buffering.