skip to content

In Ruby 4.0, how would you wrap an endless, paginated stock-tick API as an Enumerator using Enumerator.new or Enumerator.produce?

level: seniorimportance: nice to knowfreq 22%

answer

  1. the block runs only when iterated
  2. yielder << returns the yielder
  3. each(&y) via Yielder#to_proc
  4. produce: previous value in, next out
  5. size: keyword, new in 4.0

basics

~20 s

Enumerator.new { |y| ... } runs its block only when iterated, pushing each tick with y << tick as it fetches pages. Enumerator.produce(first) { |prev| ... } derives each value from the last and stops on StopIteration; Ruby 4.0 added its size: keyword.

solid answer

~40 s

With `Enumerator.new`, the block receives an `Enumerator::Yielder`; the block runs only when a consumer iterates, and each `y << tick` hands one value to the consumer and returns the yielder, so calls chain. A page loop inside it (fetch, push each tick, follow the cursor) becomes an endless Enumerator that fetches only as many pages as `first(50)` or a `lazy` chain needs. `page.ticks.each(&y)` works because `Yielder#to_proc` exists. `Enumerator.produce(initial) { |prev| ... }` fits a state machine: each call turns the previous value (say, the previous page) into the next, and raising `StopIteration` ends it. Its `size` defaults to `Float::INFINITY`; Ruby 4.0 added a `size:` keyword taking an Integer, `Float::INFINITY`, a callable or `nil`. Both re-run from the start on every new iteration, so they are not caches.

code

ruby · 21 lines
ruby
Page = Data.define(:ticks, :next_cursor)

# stand-in for an HTTP client: three pages, then no cursor
PAGES = { nil => Page.new([1, 2], :a), a: Page.new([3], :b), b: Page.new([4, 5], nil) }

feed = Enumerator.new do |y|
  cursor = nil
  loop do
    page = PAGES.fetch(cursor)
    page.ticks.each(&y)
    cursor = page.next_cursor or break
  end
end
p feed.first(3)   # => [1, 2, 3]; stops after the second page

pages = Enumerator.produce(PAGES[nil], size: nil) do |page|
  page.next_cursor or raise StopIteration
  PAGES.fetch(page.next_cursor)
end
p pages.size                               # => nil
p pages.lazy.flat_map { |pg| pg.ticks }.to_a # => [1, 2, 3, 4, 5]

go deeper

for a junior

Recall that Enumerator.new takes a block with a yielder and that y << value hands one value to whoever iterates.

for a middle

Explain when the block runs, how produce derives each value from the previous one, and how StopIteration ends a produce chain.

for a senior

Design paged or endless feeds that fetch only what consumers pull, keep throttling inside the generator, and avoid eager consumers on endless sources.

for a principal

Decide where streaming abstractions belong in a service, weighing pull-based enumerators against explicit batch jobs for observability and retry behaviour.

## The task A market-data API returns stock ticks in pages: each response carries a batch of ticks and a cursor for the next batch, and the feed never ends. Callers want to write `ticks.lazy.select { ... }.first(10)` without knowing about pages. Ruby offers two constructors for that. ## Enumerator.new with a yielder ```ruby def tick_feed(client, symbol) Enumerator.new(Float::INFINITY) do |y| cursor = nil loop do page = client.ticks(symbol, after: cursor) page.ticks.each(&y) # Yielder#to_proc cursor = page.next_cursor end end end ``` How it works: 1. The block does **not** run when `Enumerator.new` returns. It runs when a consumer iterates: `each`, `first`, `take`, `next`, a `lazy` chain. 2. The block's argument is an **`Enumerator::Yielder`**. `y << value` passes one value to the consumer and returns the yielder, so `y << a << b` works. `y.yield(a, b)` passes several values and returns what the consumer's block returned. 3. `Yielder#to_proc` lets you pass the yielder as a block, so `page.ticks.each(&y)` forwards a whole page. 4. The optional argument to `Enumerator.new` is the size hint (a value or a callable); `Float::INFINITY` documents that the feed is endless. 5. Because the consumer pulls, `tick_feed(client, "ACME").first(50)` fetches only the pages needed for 50 ticks, then stops. ## Enumerator.produce `Enumerator.produce(initial) { |prev| next_value }` builds an Enumerator whose first element is `initial` and whose every later element is the block applied to the previous one. Raising `StopIteration` inside the block ends the sequence. ```ruby pages = Enumerator.produce(client.ticks("ACME")) do |page| page.next_cursor or raise StopIteration client.ticks("ACME", after: page.next_cursor) end pages.lazy.flat_map { |page| page.ticks }.first(50) ``` It is the natural fit when the state *is* the value, as with "next page from the previous page", a date that advances by a day, or a parent pointer walked upward. ## Choosing between them | | `Enumerator.new { \|y\| ... }` | `Enumerator.produce(init) { \|prev\| ... }` | |---|---|---| | Shape | any loop, any number of yields per step | one value per step, derived from the previous | | Emits several values per fetch | yes, via `y <<` or `each(&y)` | no, emit pages and `flat_map` them lazily | | How it ends | the block returns | the block raises `StopIteration` | | Size | constructor argument | `size:` keyword (Ruby 4.0), default `Float::INFINITY` | ## The Ruby 4.0 change Ruby 4.0's NEWS adds an optional **`size:`** keyword to `Enumerator.produce`. It accepts an Integer, `Float::INFINITY`, a callable such as a lambda, or `nil` for unknown; left out, the size is `Float::INFINITY`. The rdoc warns that this default is inaccurate for a producer whose block raises `StopIteration`, so pass `size: nil` for a finite walk whose length you do not know. ## Traps with endless Enumerators - **Eager consumers never return.** `tick_feed(...).map { ... }`, `select`, `to_a`, `sort_by` or `count` on an endless feed loop forever. Use `first(n)`, `find`, `take_while` on a `lazy` chain, or external `next`. - **Every iteration starts over.** Calling `first(10)` twice fetches the first pages twice. Store the result, or iterate once with `next`. - **Errors surface in the consumer.** An HTTP error raised inside the block propagates out of `first` or `next` in the caller. - **Throttling belongs inside the block.** A `sleep` or rate limiter between page fetches runs only as fast as the consumer pulls. ## Summary `Enumerator.new` with a yielder turns any loop into a pull-based stream; `Enumerator.produce` turns a step function into one. Both are lazy about *starting* but not about their consumers, so pair endless ones with `first`, `lazy` or `next`.

  • What does y.yield(tick) return inside Enumerator.new, and when does that matter?
    It returns the value the consumer's block returned for that element. Under `map` that is the mapped value; under `next` it is whatever `feed` supplied, or nil. It matters only for generators that adapt to their consumer, which is rare; `y << tick` discards it and returns the yielder.
  • Why would you pass size: nil to Enumerator.produce in Ruby 4.0?
    Because the default is `Float::INFINITY`, which is wrong for a producer that ends by raising `StopIteration`, such as walking pages until the cursor runs out. `size: nil` tells callers the length is unknown instead of claiming it is infinite.

saying these in an interview costs you the question

  • Enumerator.new runs its block immediately and buffers the values
  • y << tick returns the tick, so it cannot be chained
  • Enumerator.produce ends when its block returns nil
  • Calling first(10) twice reuses the pages fetched the first time
  • select on an endless Enumerator returns once the matches run out