skip to content

In Ruby, scanning a due-date-sorted list of millions of invoices, which query methods stop early, and why is take_while not a drop-in for select?

level: seniorimportance: should knowfreq 28%

answer

  1. short-circuit vs full walk
  2. find, any?, include? stop early
  3. select, count, partition walk all
  4. take_while stops at first false
  5. only correct on sorted input

basics

~20 s

find, any?, all?, none?, include? and take_while stop once the answer is known; select, reject, count, grep, partition and filter_map walk everything. take_while stops at the first failure, so it matches select only on data ordered by that test.

solid answer

~40 s

Enumerable's query methods split into two families. **Short-circuiting**: `find`/`detect`, `find_index`, `include?`, `any?`, `all?` and `none?` stop at the first decisive element, `one?` at the second match, and `take_while` at the first element whose block is falsy. **Full walk**: `select`, `reject`, `count` with a block, `grep`, `partition` and `filter_map` visit every element. On millions of rows, `select { }.first`, `select { }.any?` and `select { }.size` should become `find`, `any?` and `count`. `take_while` is a **prefix** operation: on invoices sorted by due date, `take_while { |i| i.due_on < today }` returns every overdue one and stops at the first current invoice. On unsorted data it silently misses later matches, which is why it is not a drop-in for `select`. `drop_while` returns the remaining suffix.

code

ruby · 13 lines
ruby
days_late = [3, 12, 1, 9]

days_late.take_while { |d| d < 10 }   # => [3], stops at 12
days_late.select     { |d| d < 10 }   # => [3, 1, 9]
days_late.drop_while { |d| d < 10 }   # => [12, 1, 9]

checked = []
[5, 40, 70, 90].any? { |d| checked << d; d > 30 }  # => true
checked                                            # => [5, 40]

checked = []
[5, 40, 70, 90].count { |d| checked << d; d > 30 } # => 3
checked                                            # => [5, 40, 70, 90]

go deeper

for a junior

Recall that find and any? stop at the first match, while select always checks every element.

for a middle

Explain which query methods short-circuit, and why take_while returns a prefix rather than all passing elements.

for a senior

Rewrite select.first, select.any? and select.size in hot paths, and use take_while only where sort order is guaranteed and documented.

for a principal

Decide where ordering guarantees belong in the data layer so that prefix scans stay correct as data sources change.

## Why stopping early matters When a collection is small, the choice between query methods is about readability. When it holds **millions of invoices**, or when elements come from a streaming source that has to be read, the number of elements visited dominates the cost. Some `Enumerable` methods can answer after a few elements; others must see every one. ## The two families | Stops early | Stops when | Always walks everything | |---|---|---| | `find` / `detect` | first truthy block | `select` / `filter` | | `find_index` | first match | `reject` | | `include?` | first `==` element | `count` with a block or argument | | `any?` / `none?` | first match | `grep` / `grep_v` | | `all?` | first failure | `partition` | | `one?` | second match | `filter_map` | | `take_while` | first falsy block | `drop_while` (after its prefix) | The consequences for code review follow directly: - `invoices.select { |i| i.overdue? }.first` should be `invoices.find { |i| i.overdue? }`. - `invoices.select { |i| i.overdue? }.any?` should be `invoices.any? { |i| i.overdue? }`. - `invoices.select { |i| i.overdue? }.size` should be `invoices.count { |i| i.overdue? }`. This one still visits every element, but it no longer builds an array of matches just to count it. The early-exit methods also run the block **fewer times**, which matters when the block does I/O, logs, or computes something expensive. ## `take_while` is a prefix, not a filter `take_while` calls the block on elements from the start and returns them **until the first falsy result**, then stops without looking further. `drop_while` skips that same prefix and returns every remaining element, without testing them. That makes `take_while` a powerful tool on **ordered** data: 1. Invoices are sorted by `due_on`, oldest first. 2. `invoices.take_while { |i| i.due_on < today }` returns every overdue invoice. 3. It stops at the first invoice that is not yet due, never touching the millions that follow. On **unsorted** data the same call is a bug. `[3, 12, 1, 9].take_while { |d| d < 10 }` returns `[3]`: it stops at 12, and the 1 and 9 that also pass the test are never seen. `select` would return `[3, 1, 9]`. Nothing raises; the result is simply incomplete. ## Designing around order The senior judgement is whether the ordering is **guaranteed** or merely **usual**: - If the list comes from a query with an explicit sort on due date, `take_while` is correct and fast. - If the order is incidental, such as insertion order that usually matches due date, `take_while` works in testing and fails when one late-entered invoice breaks the order. - If the ordering cannot be guaranteed, use `select`, or sort first and accept the cost of the sort. A short comment next to a `take_while` stating the ordering assumption saves the next reader from turning it back into `select`, or from breaking the ordering upstream. ## Validation that stops at the first problem The early-exit family is also the right tool for validation over large imports: - `rows.all? { |r| r.amount.positive? }` stops at the first non-positive amount, so a bad file fails fast. - `rows.none?(nil)` stops at the first `nil`. - `rows.find { |r| !r.valid? }` both detects and returns the offending row for the error message, which is usually more useful than a bare `false`. The opposite mistake, counting failures with `count { ... } > 0`, walks every row even though the first failure already settled the answer. ## Streaming sources The distinction becomes sharper when the collection is an `Enumerable` that reads from a file or a network feed. `find` and `take_while` stop reading once they have their answer; `select` reads to the end. Combining early-exit methods with a lazy pipeline is a separate topic, but the principle is the same: pick the method whose stopping rule matches the question. ## What the interviewer is testing They want to hear the two families named, the rewrites for `select.first`, `select.any?` and `select.size`, and above all the ordering precondition of `take_while`, which separates people who have used it from people who have only read its name.

  • What does drop_while return, and does it keep testing after the prefix?
    It returns every element after the leading run for which the block was truthy. Once the block returns a falsy value, `drop_while` stops calling it and collects all remaining elements, whether or not they would pass. `[1, 5, 2].drop_while { |d| d < 3 }` is `[5, 2]`.
  • Is count { } better than select { }.size if both visit every element?
    Yes, modestly. Both call the block on every element, but `select` also allocates an array of all matches only for `size` to read its length. `count` keeps a running integer, so it produces less garbage, which adds up in hot paths or on very large collections.

saying these in an interview costs you the question

  • take_while returns every element that passes the test, like select
  • select { }.first stops as soon as it finds a match
  • any? evaluates the block on every element before answering
  • drop_while removes every element that passes the test
  • count with a block stops early once it has a count