Why can a dict comprehension over `zip(ids, scores)` silently lose rows from a 6,800-row batch?
answer
- Three silent losses, not one
- The shortest input ends the pairing
- Repeated keys overwrite without complaint
- Position is the only thing zip matches on
- strict=True since Python 3.10
basics
~20 sTwo silent losses stack: zip stops at the shorter input, dropping the tail, and a repeated key overwrites its earlier entry. Worse, zip pairs purely by position, so if the two lists were built by different passes the mapping is wrong rather than short.
solid answer
~40 s`{rid: s for rid, s in zip(ids, scores)}` makes three assumptions and fails quietly on all of them. First, `zip` stops when the shortest input is exhausted, so a length mismatch drops the tail with no error. Second, duplicate ids overwrite, so the dict ends up shorter than the batch and holds the last score for each repeat. Third — the dangerous one — `zip` pairs by position only: if the ids and the scores came from separate passes and one of those passes filtered or reordered, every pair after the divergence point is mis-assigned, and the dict is full length and wrong. The fixes are `zip(ids, scores, strict=True)` from Python 3.10, which raises `ValueError` on a mismatch, and building the mapping from one record source so id and score can never drift apart.
code
python · 13 linesids = [101, 102, 103, 102]
scores = [0.91, 0.42, 0.77, 0.13]
by_id = {rid: s for rid, s in zip(ids, scores)}
print(len(by_id), by_id[102])
short = [0.91, 0.42]
print({rid: s for rid, s in zip(ids, short)})
try:
{rid: s for rid, s in zip(ids, short, strict=True)}
except ValueError as exc:
print("ValueError:", exc)go deeper
Remember two facts: zip stops at the shortest input without complaining, and a dict cannot hold the same key twice, so the later pair wins. Both make the result shorter than you expect.
Explain each failure mode and its fix — strict=True from Python 3.10 for length mismatches, a different result shape for legitimate duplicates — and be able to write the length assertion that catches both.
Show that you know the mis-pairing case is the dangerous one because it does not shrink anything, and that you would restructure to pair by record or by key rather than by index. Talk about how you would detect it in a batch already in production.
Frame it as an unchecked invariant living in code that never verifies it. Own the pipeline convention that identifiers travel with their data, and decide where batch reconciliation belongs so that a mis-scored batch cannot reach downstream consumers unnoticed.
A fraud-scoring batch computes a score per row and builds a lookup with a dict comprehension over `zip`: ```python by_id = {rid: s for rid, s in zip(ids, scores)} ``` For a 6,800-row batch, `len(by_id)` comes back at 6,782. Nothing raised. There are three distinct ways this construction loses or corrupts data, and they have different fixes. ## 1. `zip` truncates to the shortest input `zip` stops as soon as any input is exhausted, and reports nothing. If `scores` is short by eighteen entries — a filtered step upstream, a partial retry, a batch boundary off by one — the last eighteen ids are simply absent from the mapping. Every downstream lookup for those ids raises `KeyError` far from the cause, or worse, hits a `dict.get` default and scores them as benign. The fix in the comprehension itself is `strict=True`, added in **Python 3.10** (PEP 618): ```python {rid: s for rid, s in zip(ids, scores, strict=True)} # ValueError on mismatch ``` `itertools.zip_longest` is the other direction — pad to the longest with a fill value — which is right only when a missing score genuinely means *unknown* rather than *bug*. ## 2. Duplicate keys collapse A dict comprehension is not required to produce one entry per input pair. If the same row id appears twice — a replayed record, a join that fanned out — the second write overwrites the first, silently. The result is shorter than the input, and it holds the *last* score for that id, which may not be the one you would have chosen. If duplicates are legitimate, the comprehension should build a different shape (a dict of lists, or a set of scores per id) rather than a last-one-wins mapping. Both of these show up as `len(result) != len(ids)`, which is why the single most valuable line you can add is an assertion or a logged comparison of those two counts at build time. ## 3. Positional pairing — the failure that does not shrink anything The first two make the dict too small; this one keeps it the right size and makes it wrong. `zip` knows nothing about ids or scores. It pairs the first with the first, the second with the second, and so on. That is only correct if both sequences were produced by the same traversal, in the same order, with no filtering on either side. In real pipelines that assumption erodes: - the ids were collected in one pass and the scores computed in a later pass that skipped rows failing validation; - one side was built by iterating a set or by a comprehension over a set, so its order is not the source order at all; - one side was sorted for a report and the other was not. Drop a single row on one side and every pair after it is shifted by one. Row 4,102's score is attached to row 4,101's id, all the way to the end of the batch. The dict has 6,800 entries, every lookup succeeds, and the service confidently scores thousands of records against the wrong evidence. No exception, no shortfall, no signal at all until someone reconciles against another system. ## Diagnosis Compare `len(by_id)` with `len(ids)` and with `len(scores)` — a mismatch localises causes 1 and 2 immediately. Find the duplicates by counting the ids with `collections.Counter` and keeping the keys whose count exceeds one. For the ordering failure, counting does not help; you need a spot check that re-derives the score for a few sampled ids from the row itself and compares. If a re-derived score disagrees while the counts all match, you are looking at mis-pairing, not loss. ## The structural fix Stop building parallel sequences. Keep the identifier and its score together in one object for the whole pipeline, and let the comprehension read both from the same item: ```python by_id = {row["id"]: score(row) for row in rows} ``` Now nothing can drift: a filtered row removes its own key and its own score together, ordering is irrelevant because the pairing is structural rather than positional, and the only remaining failure mode is duplicate ids — which an assertion on `len` catches on the first bad batch. Where two sequences genuinely arrive from different systems, `strict=True` is the minimum, and a shared key on both sides — pairing by matching id rather than by index — is the real answer. The general lesson is that a comprehension over `zip` encodes an *invariant* (these two sequences are aligned) in a place where nothing checks it. Either make the invariant checkable, or restructure so it cannot be violated.
- Which of these failure modes does `strict=True` not protect you against?Both of the ones that do not involve a length mismatch. It cannot see duplicate keys, because the pairs are produced correctly and the dict discards them afterwards, and it cannot see mis-ordering, because two equally long but misaligned sequences look perfectly valid to `zip`. It only turns a silent truncation into a `ValueError`.
- How would you keep every score when the same id legitimately appears more than once?Do not build a last-one-wins mapping. Accumulate into a dict of lists in an explicit loop, or use `collections.defaultdict` with a list factory, so each id keeps all of its scores. If you want the comprehension shape, first group the pairs and then build `{rid: combine(group) for rid, group in grouped.items()}` with an explicit combining rule.
- What single cheap check would have caught the truncation in production?Comparing the length of the built mapping against the input row count at build time, and failing or alerting on a mismatch rather than logging it. It is one line, it catches truncation and duplicate collapse together, and it converts a silent data-correctness bug into a loud batch failure.
Pairing two printed lists by line number: lose one line from either page and every couple below it is married to the wrong partner, yet the register still looks complete.
saying these in an interview costs you the question
- Believes zip raises when the inputs differ in length
- Thinks a dict comprehension keeps every produced pair
- Assumes two parallel lists stay aligned by construction
- Says strict=True also detects mis-ordered inputs
- Only checks for KeyError downstream instead of at build time
- Reaches for zip_longest to hide a genuine mismatch