skip to content

How does `__match_args__` decide what a positional class pattern binds in Python?

level: middleimportance: must knowfreq 45%

answer

  1. Positional patterns need a declared order
  2. A tuple of attribute names on the class
  3. Positional is sugar over keyword form
  4. Dataclasses generate it from __init__
  5. Too many sub-patterns raises TypeError

basics

~20 s

__match_args__ is a class attribute holding a tuple of attribute names. A positional class pattern maps its sub-patterns onto that tuple in order, so case Job(a, b) reads the first two names. More positional sub-patterns than entries raises TypeError.

solid answer

~40 s

Positional sub-patterns carry no attribute names, so the interpreter converts them to keyword form using the class's `__match_args__` — a tuple of strings naming attributes in order. `case Job(a, b)` therefore means `case Job(<name0>=a, <name1>=b)`. A class without `__match_args__` accepts zero positional sub-patterns and raises `TypeError` when one is tried against an instance; supplying more sub-patterns than the tuple has entries raises the same error naming both counts. Fewer is fine. `dataclasses.dataclass` generates `__match_args__` from its generated `__init__` parameters, excluding `init=False` fields, and skips generation under `match_args=False` or when the class already defines the attribute; `collections.namedtuple` and `typing.NamedTuple` set it to the field names. The tuple must really be a tuple — a list raises `TypeError` — and naming the same attribute both positionally and by keyword is also a `TypeError`.

code

python · 21 lines
python
from dataclasses import dataclass


@dataclass
class ShardJob:
    shard: int
    docs: int


print(ShardJob.__match_args__)

match ShardJob(3, 1200):
    case ShardJob(s, n):
        print("positional read as shard, docs:", s, n)

try:
    match ShardJob(3, 1200):
        case ShardJob(s, n, extra):
            pass
except TypeError as exc:
    print(exc)

go deeper

for a junior

Recall that case Job(a, b) only works when the class declares __match_args__, and that dataclasses get it for free while a plain class does not. That single fact explains most first encounters with the error.

for a middle

Explain the rewrite to keyword form, the exact failure modes — no tuple, too many sub-patterns, a list instead of a tuple — and the dataclass generation rules including init=False and match_args=False.

for a senior

Show that you know the arity error is raised at match time behind the isinstance gate, so a bad pattern can lie dormant until an uncommon subject arrives, and that you test dispatch chains against every class they claim to handle.

for a principal

Own the API stance: __match_args__ is a published contract as much as the constructor signature, and a codebase that destructures value types positionally across module boundaries has coupled every consumer to field order.

`__match_args__` exists to answer one question: when a class pattern is written positionally, which attribute does each slot refer to? Keyword sub-patterns say so themselves — `case Job(shard=s)` obviously reads `shard`. Positional sub-patterns do not, so the class has to declare an order. ## The mechanism `__match_args__` is a class attribute holding a **tuple of strings**, each the name of an attribute. When the interpreter meets `case Job(a, b)`, it looks up `type(subject).__match_args__`, takes the first two names, and rewrites the pattern into the keyword form `case Job(<name0>=a, <name1>=b)`. From there it proceeds exactly like any keyword class pattern: read the attribute, match the sub-pattern against it, bind any capture names. There is no separate mechanism; positional is sugar over keyword. Three rules follow, and all three are enforced with `TypeError` rather than a silent non-match: * **A class with no `__match_args__` accepts zero positional sub-patterns.** A plain class with a hand-written `__init__` has none, so `case Plain(x, y)` raises `TypeError: Plain() accepts 0 positional sub-patterns (2 given)`. Keyword sub-patterns on the same class work fine. * **Too many positional sub-patterns is an error, too few is not.** If `__match_args__` is `("shard", "docs", "started")`, then `case Job(s, d)` is legal and binds the first two; `case Job(s, d, x, y)` raises `TypeError` naming both counts. * **`__match_args__` must be a tuple.** Writing it as a list raises `TypeError` saying it must be a tuple, naming the type you supplied. The critical operational detail: **all of this is checked at match time, not compile time, and only after the isinstance test passes.** A pattern with the wrong number of positional sub-patterns compiles cleanly, passes review, and sits dormant until a subject of that exact class reaches it — at which point it raises rather than falling through to the catch-all. A dispatch chain that is exercised in tests only with the common message types can hide such a landmine indefinitely. ## Who supplies the tuple **Dataclasses.** `dataclasses.dataclass` generates `__match_args__` from the parameters of the `__init__` it would generate: the fields in declaration order, inherited fields first, with `init=False` fields excluded. Pass `match_args=False` to suppress it, and if the class body already defines `__match_args__` the decorator leaves yours alone. This is why positional class patterns feel native on dataclasses and immediately fail on ordinary classes. **Named tuples.** `collections.namedtuple` and `typing.NamedTuple` set `__match_args__` to the field names, so positional class patterns work on them out of the box — separately from the fact that a named tuple is also a sequence and so matches sequence patterns too. **Hand-written classes.** Any class can declare `__match_args__ = ("shard", "docs")` in its body. That is the right move when a legacy class is a pattern-matching target: it is one line, it is explicit, and it makes the ordering a deliberate declaration instead of a side effect of the field layout. **Builtins.** A dozen builtin types are special-cased so that a single positional sub-pattern matches the *whole subject* instead of an attribute — that is a different mechanism, and it is why `case str(s)` binds the string itself. ## Errors worth recognising Naming the same attribute twice, once positionally and once by keyword, raises `TypeError` about multiple sub-patterns for that attribute — easy to hit when `__match_args__` starts with the field you also wrote as a keyword. And the attribute lookups themselves are ordinary attribute access: CPython gathers every attribute the pattern names before testing any sub-pattern, so a `property` in `__match_args__` runs its code during dispatch even when an earlier sub-pattern would have rejected the case. ## Design guidance Treat `__match_args__` as part of the class's public surface, exactly like the `__init__` signature it usually mirrors. Reordering a dataclass's fields regenerates the tuple in the new order and silently changes what every positional pattern elsewhere binds, with no error to catch it. In practice, use positional sub-patterns where the ordering is obvious and stable — two or three fields of a small value type — and keyword sub-patterns everywhere else. Keyword form costs a few characters, survives refactors, and reads better at the call site because the attribute name is right there in the pattern.

  • Which dataclass fields end up in the generated `__match_args__`?
    The parameters of the `__init__` the decorator would generate: fields in declaration order with inherited fields first, and `init=False` fields left out. `match_args=False` suppresses generation entirely, and a `__match_args__` written in the class body is never overwritten. That last rule is the supported way to pin an order that does not track field layout.
  • When does a wrong positional arity actually blow up?
    Only when a subject that passes the isinstance test reaches that case. The pattern compiles fine, and a subject of any other type falls through without touching `__match_args__`. So a dispatch chain can carry a latent `TypeError` for months and raise it the first time an uncommon message class arrives in production.
  • Are the attributes read lazily, one sub-pattern at a time?
    No. CPython looks up every attribute the pattern names before testing any sub-pattern, so a `property` listed in `__match_args__` executes even when the first sub-pattern would have failed the case. Keep expensive or side-effecting properties out of the attributes you match on, or match on a plain field that caches the result.

saying these in an interview costs you the question

  • Assumes any class supports positional sub-patterns
  • Thinks `__match_args__` can be a list
  • Believes a wrong arity just fails to match
  • Says `__slots__` or `__annotations__` supplies the order
  • Forgets dataclasses generate the tuple automatically
  • Names one attribute both positionally and by keyword

context