skip to content

In a stream pipeline, why does a positional take of the first n elements stop the source while a predicate filter does not?

level: middleimportance: should knowfreq 40%

answer

  1. ask what the step knows
  2. content test versus a count
  3. a count proves the rest irrelevant
  4. cancellation travels upstream to the source
  5. placement decides what n counts

basics

~20 s

A positional take counts what passes, so after the nth element it knows nothing further is needed: it cancels upstream and completes downstream. A predicate filter judges each element alone, never learns it is finished, and leaves the source producing.

solid answer

~50 s

The difference is what each step **knows**. A positional take carries a counter, so once the nth element has gone by it can prove that no later element matters: it cancels the subscription upstream and signals completion downstream, and the source stops doing work. A predicate filter only ever answers a yes/no question about the element in front of it. It has no idea whether a later element might match, so it can never end the sequence — it keeps dropping non-matching elements for as long as the source keeps producing. That is why a preview of 50 pay lines built from a filter alone reads the whole payroll, while the same preview with a positional take stops the producer after 50. Placement matters too: a take after the filter counts matching lines, a take before it counts raw rows.

code

pseudocode · 10 lines
pseudocode
// 50 payable lines: the source runs until 50 matches are found, then stops
preview = timesheetRows
    .transformEach(function(row) { return payLineFor(row) })
    .keepIf(function(line) { return line.grossPay > 0 })
    .takeFirst(50)

// 50 raw rows: the source stops after 50, and fewer than 50 may survive
sample = timesheetRows
    .takeFirst(50)
    .keepIf(function(row) { return row.payableHours > 0 })

go deeper

for a junior

Know the shapes: a predicate filter keeps or drops based on the element's content, while a positional take forwards only the first n it sees and ignores content entirely.

for a middle

Explain the mechanism: the take holds a count, so after the nth element it can cancel the subscription upstream and complete downstream, while a content test never reaches that conclusion.

for a senior

Show where it costs money: a preview built from a filter alone reads the whole source, and a positional drop trims output without saving any upstream work at all.

for a principal

The judgement is about bounding work at the right place: a count near the consumer bounds the whole chain behind it, while a limit expressed as a content test bounds nothing.

## Two steps that both shrink the output A **predicate filter** and a **positional take** both produce fewer elements than they receive, which is why they are often lumped together. They are not the same kind of step at all, and the difference shows up as wasted work in production. - A **predicate filter** forwards an element when a yes/no test on that element passes. It looks at one element, in isolation, and forgets it. - A **positional take** forwards the first n elements it sees and then stops. It does not look at the element's content at all; it counts. ## Why counting changes everything Because the take carries a counter, it can reach a conclusion the filter can never reach: *no further element can affect my output*. In a stream pipeline the consumer holds a handle on its subscription and can cancel it, and cancellation travels **upstream** — from the step that no longer wants values, back towards the source. So a positional take does two things once the nth element has passed: it cancels upstream, and it signals completion downstream. The source stops producing, and anything expensive behind it stops too. A predicate filter cannot do this. Even a filter that has rejected a million elements in a row has no evidence about the next one, so it must let the source keep going. Its output may be finite while its *work* is not. Implementations differ slightly over the exact ordering of the cancel and the completion signal around the nth element; what is universal is that a counting step can end a sequence and a content-testing step cannot. ## Take, drop and filter, side by side | Step | Decides on | Ends the sequence early | Stops upstream work | Output size | |---|---|---|---|---| | Positional take of n | position (a count) | yes, after the nth | yes, it cancels upstream | at most n | | Positional drop of n | position (a count) | no | no | source size minus n at most | | Predicate filter | the element's content | no | no | 0 up to the source's size | The middle row is the one candidates get wrong. Dropping the first n elements saves nothing upstream: the step has to **receive** those n elements to know it has passed them, and it never learns that it is done. It is a pure output-side trim. ## Placement decides what n counts Because a positional take counts what reaches it, its position in the chain is part of its meaning: - filter, then take 50 — the 50 are the first 50 lines that **passed** the filter; the source runs until 50 matches have been found; - take 50, then filter — the source produces exactly 50 raw rows, and the filter may leave you with fewer than 50 outputs, possibly none. Both are legitimate; they answer different questions. "Show me a preview of 50 payable lines" is the first. "Sample the first 50 rows and see how many are payable" is the second. ## Edge cases worth knowing 1. **Fewer elements than n.** A take of 50 over a source that produces 30 is not an error; the sequence simply completes when the source does, with 30 elements delivered. 2. **A take of zero.** The step can complete immediately without ever needing a value; implementations differ over whether the source is subscribed at all first. 3. **An endless source.** A positional take is one of the few element-wise steps that turns an endless sequence into a finite one, which is exactly why it is reached for in previews, samples and tests. 4. **A filter that will never match again.** There is nothing to be done about it inside the filter. If you know the condition, express it as a count, or as a stop condition that the step can prove — not as a content test. ## What interviewers listen for The answer that lands names the mechanism rather than the behaviour: the take holds a count, and a count is enough to prove that the rest of the sequence is irrelevant, so it cancels upstream. The follow-ups that separate levels are the drop — which saves nothing and stops nothing — and the placement question, where a candidate who says "it depends what you want n to count" has understood that a positional step is defined by what reaches it.

  • Does a positional drop of the first n elements stop any work upstream?
    No. The step has to receive those n elements in order to know it has passed them, so the source does exactly the same work; only the output is trimmed. And because it never reaches a conclusion about later elements, a drop can never end the sequence — it is the mirror image of a take in name only.
  • What does a positional take of 50 do when the source produces only 30 elements?
    Nothing special: it forwards all 30 and the sequence completes when the source completes. The take is an upper bound, not a requirement, so a short source is not a failure and no padding or error appears. A pipeline that needs exactly 50 has to check the count itself.

saying these in an interview costs you the question

  • Thinks a predicate filter completes once no later element could match
  • Believes a positional take lets the source run and just ignores the surplus
  • Treats take and filter as interchangeable because both shrink the output
  • Says dropping the first n elements avoids the work of producing them
  • Expects a take of 50 to fail when the source produces only 30