skip to content

What is the difference between itertools.takewhile and itertools.dropwhile?

level: middleimportance: should knowfreq 36%

answer

  1. Both act on the leading run only
  2. One keeps the prefix, one skips it
  3. Neither behaves like a filter
  4. The stopping element is tested, then lost
  5. After dropping starts yielding, no more testing

basics

~10 s

itertools.takewhile yields items until the predicate first returns false and then stops permanently. itertools.dropwhile discards items while the predicate is true, then yields the first false item and everything after it without testing again.

solid answer

~50 s

Both take a predicate and an iterable, and both make a single one-way decision rather than filtering. `takewhile(pred, it)` yields items while `pred` holds and stops at the first failure - permanently, even if later items would pass. `dropwhile(pred, it)` throws away the leading run where `pred` holds, then yields the first item that fails **and every item after it**, with the predicate never called again. The detail interviewers push on is the boundary element in `takewhile`: it was pulled from the source in order to be tested, and it is then discarded, so it never reaches you and it is no longer in the source either. Contrast both with the built-in `filter`, which tests every element and keeps all that pass. takewhile and dropwhile only make sense when the data is ordered or monotonic - sorted, timestamped, or an ascending counter.

code

python · 10 lines
python
from itertools import takewhile, dropwhile

readings = iter([1, 2, 3, 10, 4, 5])
print(list(takewhile(lambda n: n < 5, readings)))
# [1, 2, 3]
print(list(readings))
# [4, 5]   the 10 was pulled to be tested, then dropped

print(list(dropwhile(lambda n: n < 5, [1, 2, 3, 10, 4, 5])))
# [10, 4, 5]

go deeper

for a junior

Learn the two one-line definitions and be able to predict the output of each on a short literal list. Knowing that neither one is a filter is most of the answer at this level.

for a middle

Explain the leading-run model, contrast both against the built-in filter with a concrete list, and volunteer that takewhile's failing element is consumed and discarded rather than left in the source.

for a senior

Show where the short-circuit pays for itself - stopping early on a large sorted stream, or bounding an endless counter - and know when losing the boundary element makes an explicit loop with a break the better engineering choice.

for a principal

Judge readability against cleverness: a dropwhile/takewhile chain that extracts a middle band is elegant but loses an item at each seam and is hard for a reviewer to verify. Decide when a plain state machine is the more maintainable expression of the same intent.

### Two prefix operations, not filters `itertools.takewhile(predicate, iterable)` and `itertools.dropwhile(predicate, iterable)` are about a *prefix*: the leading run of items for which the predicate is true. `takewhile` keeps that prefix and discards the rest; `dropwhile` discards it and keeps the rest. Everything surprising about them follows from that one framing. The key word is *leading*. Neither one examines the whole input the way the built-in `filter` does. `filter(lambda n: n < 5, [1, 2, 9, 1])` gives `[1, 2, 1]`, testing all four items. `takewhile(lambda n: n < 5, [1, 2, 9, 1])` gives `[1, 2]` and stops for good at the `9`; the trailing `1` would have passed, but the predicate is never consulted again. Symmetrically, `dropwhile(lambda n: n < 5, [1, 2, 9, 1])` gives `[9, 1]` - once dropping stops, it stops, and the trailing `1` is emitted even though it satisfies the drop condition. ### The element that fails takewhile is gone This is the sharpest question on the pair. To decide whether to stop, `takewhile` must first *pull* the next item from the source and call the predicate on it. When the predicate returns false, that item has already been consumed - and `takewhile` does not yield it and cannot push it back. So it vanishes: it never reaches your loop, and it is no longer available from the source iterator either. ```python readings = iter([1, 2, 3, 10, 4, 5]) print(list(takewhile(lambda n: n < 5, readings))) # [1, 2, 3] print(list(readings)) # [4, 5] <- the 10 is gone ``` If you need that boundary element - the first record after a cutoff, the header that ended a block - `takewhile` is the wrong tool. Reach for an explicit loop with a `break` that keeps the value, or restructure so the boundary is handled by the code that consumes the rest. `dropwhile` does not lose anything, because the item that ends the dropping phase is precisely the first item it yields. Internally it holds that first false item and emits it before switching to a straight pass-through. ### Order is a precondition Both functions assume something about the *arrangement* of the input. Asking "take while the value is below the threshold" is a meaningful question only if values below the threshold come first - a sorted sequence, a monotonically increasing counter, timestamps in order, lines up to a separator. Applied to unordered data they silently return an arbitrary prefix, which looks like a filter that lost most of the data. If your predicate is a property of individual items rather than of a run, you wanted `filter` (or a comprehension) instead. The payoff of the prefix framing is laziness. Over a sorted stream of a million records, `takewhile` stops reading at the first record past the cutoff; a comprehension with an `if` reads all million. That short-circuit is the whole reason to prefer it, and it is why it is the natural partner for an endless source: `takewhile(lambda n: n * n < 500, itertools.count())` terminates, whereas `filter` over the same counter would run forever because it never stops looking for more matches. ### Composing the two Running `dropwhile` and then `takewhile` over the same iterator extracts a middle band: skip the leading run, then keep until the next failure. That is a common way to isolate a section of a log or a block of lines between two markers. Because both are lazy and both consume the same underlying iterator, the composition reads the source exactly once. One caution when composing: because `takewhile` eats its boundary item, the value that terminated the first band is not available to the next stage. If you chain several bands you will lose one item at each seam. When that matters, an explicit loop with a small state variable is clearer and lossless than a chain of one-liners. ### What to say Give the definitions in one breath, then volunteer the two things a weaker answer misses: `takewhile` stops permanently rather than resuming, and the failing element is consumed and discarded. Mention that both need ordered input and that both short-circuit, which is what makes them useful against a sorted stream or an endless counter.

  • How does itertools.takewhile differ from the built-in filter on the same predicate?
    filter tests every element and keeps all that pass; takewhile tests only until the first failure and then stops for good. On `[1, 2, 9, 1]` with `n < 5`, filter yields `1, 2, 1` and takewhile yields `1, 2`. The practical difference is short-circuiting: over a sorted or endless source takewhile stops reading, while filter keeps consuming looking for more matches.
  • What happens to the element that fails takewhile's predicate?
    It is consumed and lost. takewhile has to pull the item from the source to test it, and once the predicate is false it neither yields the item nor pushes it back, so it disappears from the source iterator too. If you need that boundary value, write an explicit loop with a break that captures it, or use dropwhile, which yields the item that ends its dropping phase.
  • Why do takewhile and dropwhile assume ordered input?
    Both reason about the leading run of items satisfying the predicate, which is only a meaningful unit when related items are adjacent - sorted values, ordered timestamps, lines up to a separator. On unordered data they return an arbitrary prefix that looks like a filter which silently lost most of the data. If the predicate describes individual items rather than a run, filter or a comprehension is the correct tool.

takewhile is a queue that closes the door on the first person without a ticket; dropwhile is the same door letting nobody through until that person arrives, then waving through everyone behind them without checking.

saying these in an interview costs you the question

  • Thinks takewhile resumes when the predicate becomes true again
  • Describes takewhile as a filter over the whole input
  • Expects the failing element to appear later in the source
  • Believes dropwhile keeps testing after it starts yielding
  • Says dropwhile removes every element matching the predicate
  • Uses takewhile on unsorted data and trusts the result

context