In Ruby, how do each_slice and each_cons differ when walking a marathon results table, and what do they return?
answer
- disjoint batches vs sliding windows
- last slice may be short
- too few elements: no window
- n <= 0 raises ArgumentError
- return self since 3.1
basics
~10 seach_slice(n) yields disjoint groups of n, with a shorter last group; each_cons(n) yields overlapping windows of n consecutive elements, sliding by one. Both return the receiver in Ruby 4.0; before 3.1 they returned nil.
solid answer
~40 s`each_slice(n)` cuts the collection into **disjoint** groups: `(1..5).each_slice(2)` yields `[1, 2]`, `[3, 4]`, `[5]`, so the last group can be short and nothing is padded. It fits batching, such as inserting results 500 rows at a time. `each_cons(n)` yields **overlapping** windows that slide by one: `(1..5).each_cons(2)` yields `[1, 2]`, `[2, 3]`, `[3, 4]`, `[4, 5]`, which fits comparing each finisher with the one before. With fewer than `n` elements `each_cons` never calls the block. Both raise `ArgumentError` for `n <= 0`, and without a block both return a sized `Enumerator`. In Ruby 4.0 both return the receiver; they returned `nil` until Ruby 3.1 changed that.
code
ruby · 18 linesFinisher = Data.define(:name, :secs)
results = [
Finisher.new(name: "Abebe", secs: 7530),
Finisher.new(name: "Chen", secs: 7612),
Finisher.new(name: "Okafor", secs: 7701)
]
results.each_slice(2) { |batch| p batch.map(&:name) }
# ["Abebe", "Chen"]
# ["Okafor"]
results.each_cons(2) { |ahead, behind| p behind.secs - ahead.secs }
# 82
# 89
results.each_slice(2) { }.equal?(results) # => true in Ruby 4.0
results.first(1).each_cons(2) { raise "never" } # block never runs
(1..10).each_slice(3).size # => 4go deeper
Recall that each_slice makes separate groups and each_cons makes sliding windows of consecutive elements.
Explain the short last slice, the zero-call case for each_cons, ArgumentError for n <= 0, and that both return the receiver since 3.1.
Batch streaming sources with each_slice so memory stays bounded, and watch for code written against the pre-3.1 nil return when upgrading.
Choose batch sizes as a policy trade-off between round trips and memory per batch, and make that size a named, tunable value rather than a literal.
## Two ways to take n at a time Ruby's `Enumerable` module has two methods that hand the block several consecutive elements at once, packed in an array. They differ in whether the groups **overlap**: | | `each_slice(n)` | `each_cons(n)` | |---|---|---| | Groups | disjoint, back to back | overlapping, sliding by one | | `(1..5)` with `n = 2` | `[1, 2]`, `[3, 4]`, `[5]` | `[1, 2]`, `[2, 3]`, `[3, 4]`, `[4, 5]` | | Number of calls | size divided by n, rounded up | size minus n plus 1, or zero | | Short final group | yes, the remainder | never | | `n <= 0` | `ArgumentError` (invalid slice size) | `ArgumentError` (invalid size) | | Returns with a block | the receiver | the receiver | ## `each_slice`: batching A results table with 40,000 finishers should not be written to storage one row at a time, nor all at once. `each_slice` gives fixed-size batches: ```ruby results.each_slice(500) { |batch| store.insert_all(batch) } ``` Points worth stating: - The **last batch is shorter** when the size is not a multiple of `n`. It is not padded with `nil`. - Each yielded batch is an ordinary `Array`, so the block can call any array method on it. - `each_slice` works on any `Enumerable`, including one that streams from a file, so a batch job does not have to load everything first. ## `each_cons`: neighbours `each_cons` answers questions about **adjacent** elements. For finishers sorted by time, the gap to the runner ahead is: ```ruby results.each_cons(2) { |ahead, behind| puts behind.secs - ahead.secs } ``` - A block with two parameters destructures the two-element window. - There are `size - n + 1` windows, so a table with one finisher produces **no** calls for `each_cons(2)`; the block simply never runs. - Wider windows work the same way: `each_cons(3)` gives every run of three consecutive checkpoints. ## What they return This is where version matters. Ruby 3.1's release notes list the change: `each_cons` and `each_slice` now **return the receiver**. In Ruby 3.0 and earlier they returned `nil`. 1. In Ruby 4.0, `[1, 2, 3].each_slice(2) {}` returns `[1, 2, 3]`. 2. In Ruby 3.0 the same call returned `nil`. 3. Code that chained on the result, or a method ending in `each_slice`, behaves differently across that boundary. The change brought them in line with the rest of the each family: `each`, `each_with_index` and `reverse_each` all return the receiver. Neither method returns the groups themselves; to get them as data, call without a block and convert, as in `(1..5).each_slice(2).to_a`. ## Without a block Called without a block, both return a **sized `Enumerator`**: `(1..10).each_slice(3).size` is `4` and `(1..10).each_cons(3).size` is `8`, computed without iterating. That lets you chain `with_index(1)` to number the batches, for example `results.each_slice(500).with_index(1) { |batch, n| ... }`. ## Mistakes seen in review A handful of bugs recur around these two methods: - **Assuming a full last batch.** Code that reads `batch[499]` or divides by `n` breaks on the final, shorter slice. - **Using `each_slice(2)` for neighbours.** Disjoint pairs compare runner 1 with 2 and 3 with 4, but never 2 with 3, so half the gaps are missing. - **Expecting a window for short input.** `each_cons(2)` on a one-row table does nothing, which is correct, but a report that assumed at least one gap prints an empty section. - **Relying on the return value.** Code written for Ruby 3.0 that tested `each_slice(...) { }.nil?` behaves differently after 3.1. None of these raise; they produce plausible but wrong output, which is why interviewers ask about the edges rather than the happy path. ## Choosing in an interview Say which question each method answers: `each_slice` is for **processing in chunks**, `each_cons` is for **looking at neighbours**. Then add the edge cases that show experience: the short last slice, the zero-call `each_cons` on a short collection, `ArgumentError` for a non-positive size, and the receiver return value since 3.1.
- How would you batch a results file too large to hold in memory?Slice a source that streams rather than one already loaded: `File.foreach(path).each_slice(500) { |lines| ... }`. `File.foreach` without a block returns an enumerator that reads line by line, and `each_slice` only holds the current batch. Calling `File.readlines(path).each_slice(500)` would read the whole file into an array first.
- How do you number the batches as they are processed?Call `each_slice` without a block to get its enumerator, then add an index: `results.each_slice(500).with_index(1) { |batch, n| log("batch #{n}: #{batch.size} rows") }`. The enumerator is sized, so `results.each_slice(500).size` also tells you the batch count up front.
saying these in an interview costs you the question
- each_slice pads the last group with nil to reach n elements
- each_cons(2) on a one-element array yields that element once
- each_slice and each_cons return nil in current Ruby
- each_slice(0) returns an empty result instead of raising
- each_slice returns the array of groups it built