In a dataclass, when would you set field(repr=False) or field(compare=False)?
answer
- Three generated methods, three switches
- Some values should not be logged
- Equality is a design decision
- Timestamps and caches are not identity
- field(repr=False), field(compare=False)
basics
~20 sEach dataclasses.field() flag removes the field from one generated method: repr=False keeps a large or sensitive value out of the printed repr, compare=False keeps a volatile value out of eq, and init=False drops it from the constructor signature.
solid answer
~50 s`dataclasses.field()` lets you opt a field out of each generated method independently. `repr=False` keeps the value out of `__repr__` — the right call for a secret, a credential, or a payload too large to want in every log line. `compare=False` keeps it out of `__eq__`, which is how you say "this is incidental data, not part of the record's identity": an import timestamp, a raw source line, a cached derived value. `init=False` drops the field from the constructor signature, so callers cannot pass it; if it has no default and nothing else assigns it, the attribute simply never exists and reading it raises `AttributeError`. The flags compose — one field can be excluded from all three — and the decision is a design one: `__eq__` defines what "the same record" means, and every field you leave in becomes part of that contract.
code
python · 12 linesfrom dataclasses import dataclass, field
@dataclass
class PayrollRow:
employee_id: int
raw_line: str = field(repr=False, compare=False)
net_pay: float = field(init=False, default=0.0)
r1 = PayrollRow(11, "11,Ada,3900")
r2 = PayrollRow(11, "0011,Ada,3900")
print(r1) # PayrollRow(employee_id=11, net_pay=0.0)
print(r1 == r2) # True - raw_line is excluded from __eq__go deeper
Know that dataclasses.field() takes more than just a default, and that init, repr and compare are on/off switches deciding whether a field appears in the constructor, the printed form and equality.
Explain each flag's exact effect and the trap that field(init=False) without a default leaves the attribute unset, so the first read — including the one inside the generated repr — raises AttributeError.
Demonstrate the judgement: which values must stay out of logs, and which fields genuinely constitute record identity. Be ready with a concrete case where a volatile field inside eq broke de-duplication.
Own equality as a contract across a codebase: what 'the same record' means for a shared value type, how that interacts with de-duplication and caching, and how a team keeps that decision explicit instead of inherited from defaults.
## One field, three independent memberships A dataclass field appears in three generated methods, and `dataclasses.field()` controls each membership separately: | flag | default | effect when False | |---|---|---| | `init` | True | the field is not a constructor parameter | | `repr` | True | the field is omitted from `__repr__` | | `compare` | True | the field is ignored by `__eq__` | They are orthogonal, and the useful skill is knowing which one a given problem calls for. ## repr=False — what you are willing to print `__repr__` is the reason many teams reach for dataclasses at all: every log line, every traceback frame and every debugger view renders it. That makes it a disclosure surface and a volume problem at the same time. Exclude a field from the repr when it holds a secret (an API token, a password hash, a bank account number), when it holds bulk data (a decoded file, a large payload) that would drown the log, or when it is noise that never helps a reader. Concretely, when importing a payroll CSV, the record can keep the raw source line for error reporting — but that line contains everything the row knows about a person, and it has no business appearing in a log. `field(repr=False)` is a one-word fix that keeps it available on the object and out of the output. Be honest about what this is: it removes the value from the *generated* repr, not from the object. Anything that walks the instance's attributes still sees it. It reduces accidental disclosure; it is not a security boundary. ## compare=False — what makes two records the same The generated `__eq__` compares the tuple of every field with `compare=True`. That is a design decision disguised as a default. Ask, per field: *if two records differ only in this value, are they the same record?* - An import timestamp: two rows parsed a second apart are the same row — `compare=False`. - A raw source line: `"11,Ada,3900"` and `"0011,Ada,3900"` parse to the same employee — `compare=False`. - A cached or derived value: it is a function of the other fields, so including it adds nothing and risks a stale value making two equal records look different — `compare=False`. - A monotonically increasing id assigned at load time: including it means no two loaded records ever compare equal, which quietly defeats de-duplication. That last case is the one worth telling a story about. De-duplicating an import for an 11-person team by putting the rows in a set, or by testing `row in seen`, is a one-line idea that silently does nothing if a per-row timestamp or sequence number is part of equality — every row looks distinct, every duplicate is processed twice, and the duplicated side effect (a second payment record, a second notification) is discovered downstream rather than at the import. ## init=False — fields the caller must not supply `init=False` removes the field from the constructor signature. Use it for values the object owns rather than accepts: a computed total, a lazily-filled cache, a sequence counter. The sharp edge: `x: int = field(init=False)` with no default means nothing ever assigns the attribute. Construction succeeds, and the first read — including the one inside the generated `__repr__` — raises `AttributeError: 'B' object has no attribute 'x'`. So an `init=False` field needs either a default, a `default_factory`, or code elsewhere in the class that assigns it. Note also that `init=False` does not make the field read-only; the attribute is still assignable after construction. ## Ordering interaction Because `init=False` fields are not constructor parameters, they do not participate in the "non-default argument follows default argument" rule at all — an `init=False` field with a default can sit anywhere in the class body without forcing defaults onto the fields after it. That is occasionally a clean way out of an awkward field order. ## The judgement, stated plainly `__repr__` and `__eq__` generated from *all* fields is a sensible default and a poor contract. The senior move is to read the field list once and ask, per field, "does this belong in the printed form?" and "does this belong in identity?" — then write the answer down with `field(repr=False)` / `field(compare=False)` so the next reader does not have to re-derive it. ## Version notes `init`, `repr` and `compare` have been `field()` parameters since `dataclasses` arrived in Python 3.7, and their behaviour is unchanged on 3.14.
- What happens to a field declared field(init=False) with no default and nothing assigning it?Construction succeeds, because the field is not a constructor parameter, but the attribute is never created. The first read raises `AttributeError`, and that includes the read performed inside the generated `__repr__`, so even printing the object fails. An `init=False` field needs a `default`, a `default_factory`, or an assignment elsewhere in the class.
- Does field(repr=False) make a value secure?No. It only removes the value from the generated `__repr__`, so it stops the field appearing in log lines, tracebacks and debugger summaries that render the object. The attribute still exists and anything that walks the instance's attributes or serialises it will find it. Treat it as a way to reduce accidental disclosure, never as a security control.
- Why can a field(init=False) field with a default sit before fields that have none?The 'non-default argument follows default argument' rule is about the generated `__init__` signature, and an `init=False` field is not part of that signature at all. It is therefore invisible to the ordering check and can appear anywhere in the class body, which is sometimes the tidiest way to place a derived value next to the fields it is derived from.
saying these in an interview costs you the question
- Believing repr=False removes the attribute from the object
- Treating repr=False as a security control
- Assuming init=False makes a field read-only
- Leaving volatile timestamps inside generated equality
- Thinking compare=False also hides the field from the repr
- Expecting init=False to default the attribute to None