skip to content

Where does validation belong in a dataclass, and what does __post_init__ not cover?

level: middleimportance: should knowfreq 45%

answer

  1. One gate, at construction only
  2. The decorator writes no __setattr__
  3. Mutation afterwards is unchecked
  4. Restoring state skips the constructor
  5. Immutability or descriptors close the gap

basics

~20 s

Put invariant checks in post_init and raise ValueError there; it is the one hook the generated constructor gives you. It runs once, at construction, so a plain dataclass stays mutable and later attribute assignment bypasses every check.

solid answer

~40 s

The construction-time gate is `__post_init__`: every field is already assigned when it runs, so it can check combinations across fields and raise `ValueError` before the object escapes the constructor. What it cannot do is hold that invariant over the object's lifetime. `@dataclass` generates no `__setattr__`, so `job.pages = -1` afterwards just rebinds the attribute; nor does the hook re-run when `copy.deepcopy` or `pickle` restores an instance, since both bypass `__init__`. If the invariant must always hold, make the type immutable with `frozen=True` so mutation raises instead of silently corrupting, or move enforcement onto a descriptor or property rather than a one-shot check. And keep parsing at the system boundary: a dataclass validating a payload deep inside the call stack turns a bad request into a `ValueError` far from where it can be reported well.

code

pycon · 12 lines
pycon
>>> from dataclasses import dataclass
>>> @dataclass
... class Job:
...     attempts: int
...     def __post_init__(self):
...         if self.attempts < 0:
...             raise ValueError("attempts must be >= 0")
...
>>> job = Job(1)
>>> job.attempts = -5
>>> job
Job(attempts=-5)

go deeper

for a junior

Know the mechanical answer: cross-field checks go in post_init and raise ValueError, because the generated constructor gives you no other place to run code at construction time.

for a middle

Explain the limit as well as the hook: the decorator writes no setattr, so the object stays mutable and later assignment is unchecked, and state-restoring paths like copy and pickle never call the constructor at all.

for a senior

Demonstrate the fix you would actually ship: immutability for value types, attribute-level enforcement where mutation is required, and validation at the system boundary so a bad request produces a reportable error rather than a ValueError deep in a call stack.

for a principal

Set the policy: which layer owns input validation, whether domain types may raise from their constructors at all, and how errors are aggregated and reported — a per-field ValueError is a poor contract for an API that must return every problem at once.

## The one hook you get A dataclass's generated `__init__` only assigns fields, so `__post_init__` is the single place where construction-time rules can live. It runs after every assignment, which is what makes it the right place for *cross-field* rules that no per-field annotation can express: an end before a start, a page range wider than the document, a currency that does not match the account's. Raise `ValueError` for a value the type cannot represent and `TypeError` for the wrong kind of thing entirely; both propagate out of the constructor, so the caller never receives a half-valid object. ```python from dataclasses import dataclass @dataclass class PageRange: first: int last: int def __post_init__(self) -> None: if self.first < 1 or self.last < self.first: raise ValueError(f"bad range {self.first}-{self.last}") ``` ## The gap: it fires exactly once The decorator writes `__init__`, `__repr__` and `__eq__` — it does **not** write `__setattr__`. A plain dataclass is a normal mutable object, so after construction any code can assign any attribute and no check runs. In a document-conversion queue this is how a validated job record goes bad: the job is built correctly, and a retry handler later sets `job.attempts = -1` or overwrites a normalized format string with the raw locale-dependent spelling it read from the request. Nothing raises; the corruption is discovered downstream, and in a 27-minute regression suite that runs the whole pipeline it surfaces as a failure nowhere near the assignment that caused it. Two more paths skip the hook entirely. `copy.copy`, `copy.deepcopy` and `pickle` reconstruct an instance by restoring its state directly, without calling `__init__`, so nothing is re-validated on the way back in — a pickled bad object stays bad, and a fixed-up derived value is carried over rather than recomputed. And `dataclasses.replace()` does the opposite: it goes back through the constructor, so `__post_init__` runs again and can reject the new combination, which is usually what you want but is worth knowing before you use it in a hot loop. ## Closing the gap **Make the type immutable.** `@dataclass(frozen=True)` installs a `__setattr__` that raises `dataclasses.FrozenInstanceError`, so the invariant checked once at construction genuinely holds for the object's life; the only way to a different value is a new instance. For a record-shaped value type this is almost always the right answer, and it turns the mutation bug above into an exception at the offending line. **Enforce per attribute.** If the object must stay mutable, the check belongs on the attribute rather than in the constructor — a descriptor with `__set__`, or a property whose setter validates. Combining a property with a dataclass field takes care, because the class attribute the property occupies interacts with how the decorator sees defaults; the descriptor-per-attribute approach composes more cleanly, and either way `__post_init__` remains useful for the cross-field rules a single attribute cannot see. **Or do not validate here at all.** The strongest version of the argument is that a dataclass is a *record*, and coercing or validating untrusted input is a boundary concern: parse the incoming payload where it arrives, report errors as errors of that request, and construct the dataclass only from values already known to be good. A declarative validation library exists precisely to own that boundary and produce structured error reports; a bare `raise ValueError` from a constructor five frames deep gives the caller a message and nothing else — no field name, no path, no accumulation of every problem in one response. ## Practical shape A reasonable default: `frozen=True` for value types, `__post_init__` for cross-field invariants and normalization, `ValueError` with a message that names the offending values, and no per-field type checking at runtime unless the input is untrusted — annotations are not enforced, and re-implementing a type checker inside every constructor costs more than it catches. Above all, be explicit about which of the two questions the class answers: *was this object built correctly* (the hook can answer that) or *is this object still correct* (only immutability or attribute-level enforcement can). ## Validation and normalization are different jobs Both live in the hook, and conflating them causes arguments. *Validation* rejects: it raises when the arguments cannot describe a legal value, and it leaves the object exactly as the caller specified it. *Normalization* accepts and rewrites: it strips whitespace, lowercases a code, sorts a tuple, so that values meaning the same thing become the same value. Normalization is the more dangerous of the two in a mutable dataclass, because it makes the stored state differ from what the caller passed, and the caller can then re-break it with one assignment. If a type normalizes, it usually wants to be frozen as well, so the normalized form is the only form anyone ever sees. Testing follows from the same split. Every invariant the hook checks deserves one construction test asserting that building it raises, and every normalization deserves a test asserting that two spellings of the same input produce equal objects; both are fast, because construction is all they exercise.

  • Which exception type should __post_init__ raise for an invalid combination of field values?
    `ValueError` for a value of the right kind that the type cannot represent — a negative count, an end before a start — and `TypeError` only when the argument is the wrong kind of object altogether. Both leave the constructor, so no partly-valid instance reaches the caller. Include the offending values in the message; the traceback alone points at the dataclass, not at the caller that supplied them.
  • Does dataclasses.replace() re-run the validation in __post_init__?
    Yes. `replace()` builds a new instance by calling the class's `__init__` with the existing field values merged with your changes, so the hook runs again on the result and can reject the new combination. That is the deliberate contrast with `copy.deepcopy` and `pickle`, which restore state directly and never re-validate anything.
  • How would you keep a dataclass invariant true for the object's whole lifetime rather than only at construction?
    Either remove mutation — `frozen=True` makes assignment raise `dataclasses.FrozenInstanceError`, so the checked-once invariant holds and changes go through a new instance — or move the check to where the mutation happens, with a descriptor's `__set__` or a validating property setter. `__post_init__` then keeps only the cross-field rules that no single attribute can evaluate.

saying these in an interview costs you the question

  • Believes __post_init__ guards later attribute assignment
  • Thinks annotations are type-checked at runtime
  • Assumes unpickling re-validates through the constructor
  • Raises a bare Exception instead of ValueError
  • Puts untrusted-payload parsing inside the value type

context