Why is an itertools.groupby sub-iterator empty after the outer loop advances?
answer
- The groups are not independent
- One shared source behind every group
- Advancing kills the previous group silently
- Empty groups, not an exception
- Materialize inside the loop body
basics
~20 sAll groups share one underlying source iterator. Advancing the outer groupby object invalidates the previous group and skips its unread items, so a group read later yields nothing. Consume each group inside the loop, typically with list().
solid answer
~50 s`itertools.groupby` does not copy anything. It yields `(key, group)` pairs where the group is a thin view over the **same** source iterator, and every call to the outer iterator's `__next__` first invalidates the group it handed you last, then fast-forwards past whatever of that group you did not read, looking for the next key change. A group object you kept a reference to therefore yields nothing at all — not an error, just an empty iteration. So `list(groupby(data))` followed by materializing the groups gives you every key with an empty list, and `{k: g for k, g in groupby(data)}` is the same bug wearing a dict. The rule is: fully consume or store each group before pulling the next pair — `[(k, list(g)) for k, g in groupby(data, key=f)]` works because the comprehension calls `list(g)` before it asks for the next pair.
code
pycon · 6 lines>>> from itertools import groupby
>>> pairs = list(groupby('AABBC'))
>>> [(k, list(g)) for k, g in pairs]
[('A', []), ('B', []), ('C', [])]
>>> [(k, list(g)) for k, g in groupby('AABBC')]
[('A', ['A', 'A']), ('B', ['B', 'B']), ('C', ['C'])]go deeper
Learn the habit before the theory: turn each group into a list inside the loop that produced it. If you see the right keys with empty groups, you read a group after moving on.
Explain the mechanism: one shared source iterator, a grouper that is only valid while current, and an outer advance that invalidates it and skips its unread items. Be ready to spot the dict-of-groupers and list-of-pairs variants.
Demonstrate that you can keep the streaming benefit while staying correct — aggregate a group lazily rather than materializing it, and know when the honest answer is to abandon groupby for an eager mapping of lists.
Own the tradeoff behind the design: independent groups would require unbounded buffering, so the library sells constant memory in exchange for a validity window. Be ready to say when a pipeline should expose lazy views at all and when it should hand callers owned containers.
## One iterator, borrowed by everyone `itertools.groupby` is a single-pass streaming operator. It holds exactly one source iterator, one current item and one current key. When it detects a key change it constructs a lightweight *grouper* object and hands it back as the second half of the `(key, group)` pair. That grouper does not own a buffer of the group's items; it is a window onto the shared source, and it knows it is only valid while groupby considers it the current group. Calling `__next__` on the outer groupby object does two things, in this order: 1. it marks the previously handed-out grouper as no longer current — from that instant the grouper is inert and will raise `StopIteration` on its first `next()`; 2. it drains the source until the computed key changes, discarding any items of the old group you never read, and then builds the next pair. Nothing about step 1 is an error condition, which is what makes it dangerous: a dead grouper does not complain, it simply behaves like an empty iterable, and your report ends up with the right keys and no rows under them. ## The three shapes of the bug ```python from itertools import groupby data = sorted(["a1", "a2", "b1"], key=lambda s: s[0]) # 1. materialise the pairs first -- every group is already dead print([(k, list(g)) for k, g in list(groupby(data, key=lambda s: s[0]))]) # [('a', []), ('b', [])] # 2. same bug in a dict comprehension lookup = {k: g for k, g in groupby(data, key=lambda s: s[0])} print({k: list(g) for k, g in lookup.items()}) # {'a': [], 'b': []} # 3. correct: consume each group before advancing print([(k, list(g)) for k, g in groupby(data, key=lambda s: s[0])]) # [('a', ['a1', 'a2']), ('b', ['b1'])] ``` Shapes 1 and 2 both stash groupers and read them after the outer iterator has moved on — in shape 1, past the end. Shape 3 differs only in evaluation order: the comprehension binds one pair, evaluates `list(g)` immediately, and only then asks groupby for the next pair. The same reasoning covers the wider family: appending groupers to a list for later, returning one from a function, handing one to another thread or scheduling it as work to run later, or storing them so a second pass can re-read a group. Each keeps a reference that outlives its validity window. ## The rule, and how to keep laziness The rule is one sentence: **consume or store a group before you advance the outer iterator.** `list(g)` inside the loop body is the standard way, and it is what the standard library's own documentation recommends. That does not force you to materialize everything. What must happen inside the loop is *use*, not *storage*. If you are aggregating, you can consume a group lazily and keep only the aggregate — `sum(x.amount for x in g)` or `sum(1 for _ in g)` — so peak memory is one group's worth of iteration state rather than the whole grouping. This is the property that makes groupby worth using on a sorted stream you cannot fit in memory: you see each group exactly once, in order, and you are free to discard it immediately. You may also read part of a group and move on. `[next(g) for _, g in groupby(data)]` returns the first item of each run and is perfectly legal — the unread remainder is skipped for you when the outer iterator advances. What you cannot do is come back for it. ## Why not just make groups independent? Because that would require buffering. To let two groupers be alive at once, groupby would have to read ahead and store the items of the earlier group, which turns a constant-memory operator into one whose footprint is bounded by the largest group — and on an unbounded stream, by nothing at all. The standard library chose the streaming contract and documented the consequence. If you genuinely need random access to the groups — several passes, lookups by key, groups consumed out of order — then you want a mapping of lists built in one pass, or `{k: list(g) for k, g in groupby(sorted_data, key=f)}` if the input is already sorted. Both give you real containers with real lifetimes. The choice is between a lazy view with a validity window and an eager structure you own outright; groupby only offers the first. ## Diagnosing it in the wild The symptom is characteristic: correct group keys, correct group count, and empty or short groups. Look for a grouper that escaped its loop iteration — a comprehension over `list(groupby(...))`, a dict keyed on groupers, a group passed to something deferred, or a `break` followed by a later read. The fix is always the same edit: materialize inside the loop.
- Why does the comprehension [(k, list(g)) for k, g in groupby(data)] work while list(groupby(data)) then reading the groups does not?Evaluation order. The comprehension binds one pair, calls `list(g)` on the spot, and only afterwards asks the outer iterator for the next pair, so each group is read while it is still current. `list(groupby(data))` drives the outer iterator to exhaustion first, invalidating every grouper — including the last — before any of them is read, which is why you get the keys with empty groups.
- Does list() inside the loop cost you the streaming property groupby was chosen for?Only if you keep the lists. Peak memory is one group at a time as long as you discard each list before the next iteration, so a huge sorted stream is still fine provided no single group is enormous. If you only need an aggregate, skip the list entirely and consume the group lazily — `sum(1 for _ in g)` or a running total — which keeps the footprint at a single item.
- What happens to the items of a group you only partially read?They are discarded. When you advance the outer iterator, groupby drains the source past the rest of the current run to find the next key change, so those items are consumed and gone. That makes `[next(g) for _, g in groupby(data)]` a legitimate way to take the first item of every run — the remainder is skipped for you — but there is no way to return for it later.
Each group is a view through a moving train window, not a photograph: once the train has rolled on, looking at the old window shows you nothing, and no one tells you the view expired.
saying these in an interview costs you the question
- Thinks each group is an independent list
- Expects an exception when a stale group is read
- Builds a dict of group objects for later use
- Believes the last group survives outer exhaustion
- Claims list() inside the loop defeats streaming entirely
- Calls the shared-source design a standard-library bug