Why does a `match` guard's helper call run twice per record in a geocoding batch?
answer
- The subject is evaluated once; guards are not
- Every leading case whose pattern matched pays
- A rejected case has already done its work
- Hoist the expensive call above the statement
basics
~20 sBecause guards are evaluated per case, in source order, until one succeeds. Two overlapping cases that both call the lookup in their guards call it twice: the first case's guard ran and paid its side effect before rejecting the record.
solid answer
~50 sThe subject of a `match` is evaluated once, but guards are not: for each `case` in source order, the pattern is tried and, if it matches, the guard is evaluated. A guard that returns falsy rejects only that case, so matching walks on to the next one — whose guard runs too. Two overlapping cases that both call the same lookup therefore call it twice per record. Anything the guard does beyond returning a boolean duplicates with it: network calls, counters, log lines, and memo dictionaries that grow past what the job budgeted. The fixes are to compute the value **once before** the `match` and match on it, to bind it in the first guard with a walrus and reuse the name, to memoize behind a *bounded* cache, or to make the patterns disjoint so only one guard can ever run.
code
python · 17 lineslookups = []
def geocode(addr):
lookups.append(addr)
return len(addr)
def route(record):
match record:
case {"addr": addr} if geocode(addr) > 20:
return "full"
case {"addr": addr} if geocode(addr) > 5:
return "partial"
case _:
return "skip"
print(route({"addr": "12 Mill Lane"})) # partial
print(lookups) # ['12 Mill Lane', '12 Mill Lane'] - called twicego deeper
Know that each case is tried in order and that a guard belongs to a single case. That alone explains why code in a guard can run more than once for one input.
Explain the walk: pattern, then guard, then either commit or move on. Be able to count the evaluations for a given match and say which cases never evaluate their guards at all.
Diagnose it in production: connect the doubled call rate to overlapping patterns with different guards, then choose a fix — hoist the call, reuse a walrus binding, bound the cache, or make the patterns disjoint — and say why each is or is not right for an effectful helper.
Own the convention: guards are predicates, so effectful work belongs above the match. Decide how the team enforces that in review and where a wide guarded match should instead be an explicit dispatch step with its expensive lookup done once.
## The symptom A batch job geocodes address records and dispatches on the result with a `match`. Instrumentation shows the geocoding helper invoked roughly twice per input record, the provider's rate limiter tripping at half the expected throughput, and the process settling at a **2.4 GB working set** because the helper's memo dictionary is unbounded and holds every intermediate it was asked for. Nothing in the code loops twice. ## The mechanism `match` evaluates its subject once. It does **not** evaluate guards once. Matching is a walk over the `case` clauses in source order, and for each clause: 1. the pattern is tried; 2. if — and only if — the pattern matched, the guard expression is evaluated; 3. a truthy guard commits to that body and ends the `match`; a falsy guard rejects **only that case** and the walk continues. So the number of guard evaluations for one subject is the number of leading cases whose *patterns* matched, up to and including the winner. When several cases share a shape and differ only in their guards — the common way people express "same structure, different threshold" — every one of those guards runs until one returns truthy. ```python match record: case {"addr": addr} if geocode(addr) > 20: # runs ... case {"addr": addr} if geocode(addr) > 5: # runs again ... ``` Both patterns match the same records, so `geocode` is called twice for every record that falls through to the second case. That is the duplicated side effect: a *pure* guard costs only CPU, but a guard that talks to a service, increments a counter, writes a log line, or populates a cache duplicates all of it. ## Two secondary hazards in the same code **Bindings survive a failed guard.** Pattern matching binds as it destructures and never rolls back, so after the first case's guard returns falsy, `addr` is still bound. If a later branch — or code after the whole `match` — reads a capture name it did not itself bind, it silently sees a value from a case that lost. This is the quiet cousin of the duplicated call and is much harder to spot in review. **Unbounded memoization turns a correctness smell into a memory problem.** A helper that caches in a plain module-level `dict` will happily hold every key it has ever been asked for. Doubling the call rate does not double the cache, but a job that was already retaining every intermediate now looks like a leak in monitoring, and the working set is the first thing anyone notices. ## The fixes, roughly in order of preference **Hoist the call out of the `match`.** The cleanest form: compute once, then match on the computed value. ```python score = geocode(record["addr"]) match score: case s if s > 20: ... case s if s > 5: ... ``` This is usually the right answer, and it also makes the code readable: the expensive thing is visible on its own line instead of hiding in a guard. **Bind in the first guard with a walrus and reuse it.** A guard is an ordinary expression, so `:=` works and binds in the enclosing function scope, letting the next guard read the name: ```python case {"addr": addr} if (n := geocode(addr)) > 20: ... case {"addr": addr} if n > 5: ... ``` It works and it is a legitimate trick to know, but it is fragile: it depends on the first case's *pattern* matching, so the second guard raises `NameError` if the record ever reaches it without the first pattern having matched. Prefer hoisting unless you are certain of the shapes. **Memoize behind a bounded cache.** `functools.lru_cache` with an explicit `maxsize` makes the repeat call cheap and caps the retained set — the direct answer to the 2.4 GB working set. It does not remove the duplicate *call*, only its cost, so it is a mitigation rather than a fix, and it is wrong outright if the helper's side effect is what you cared about (a counter, an audit record, an outbound request that must not repeat). **Make the patterns disjoint.** If the cases discriminate on structure rather than on a computed threshold, only one guard can ever run. This is the design-level fix and is often what the code wanted in the first place. ## The rule to carry away Treat guards as **predicates**: cheap, pure, and safe to evaluate more than once, because the language will evaluate them more than once per subject as a matter of ordinary control flow. Anything that costs money, mutates state, or must happen exactly once belongs before the `match`, not inside a `case`. This is 3.10 `match` semantics (PEP 634) and is unchanged through 3.14.
- How many times can a guard be evaluated for a single subject, and what determines the count?Each `case` contributes at most one guard evaluation, and only when its pattern matched. So for one subject the count equals the number of leading cases whose patterns matched, up to and including the case that wins. Cases whose patterns fail never evaluate their guards, and everything after the winning case is skipped entirely.
- Is `if (n := geocode(addr)) > 20:` in the first guard a safe way to avoid the second call?It works — a guard is a normal expression, so `:=` binds `n` in the enclosing scope and a later guard can read it — but it is fragile. `n` only exists if the first case's *pattern* matched; a record that reaches the second guard by another route raises `NameError`. Hoisting the call above the `match` is the safer form.
- After a guard rejects its case, what is the state of the names that case's pattern captured?Still bound. Matching binds as it destructures and never unwinds, so captures from losing cases linger in the enclosing scope with stale values. Read a capture only inside the body of the case that bound it; code after the `match`, or in a later case, that touches such a name is a latent bug.
- When is a bounded cache the wrong answer to a duplicated guard call?Whenever the duplication itself matters rather than its cost — an outbound request that must not repeat, a counter or metric, an audit or log record, anything that mutates state. Caching hides the second call's expense but the effect has still happened twice. Then you must hoist the call or make the patterns disjoint.
saying these in an interview costs you the question
- Believes a guard is evaluated only once per match statement
- Thinks a falsy guard rewinds the bindings it made
- Assumes the interpreter caches identical guard expressions
- Blames the subject being re-evaluated for each case
- Reaches for an unbounded memo dict as the fix
- Reads a capture name after the match without checking which case bound it