Why does a serializer built on vars(obj) drop fields that dir(obj) lists?
answer
- Two builtins, two different questions
- One returns values, one returns names
- Instance storage versus the whole hierarchy
- Assignment order versus alphabetical order
- Slots leave nothing for vars to return
basics
~20 svars(obj) returns only the instance's own dict, so class attributes, computed properties and slot-based fields never appear. dir(obj) merges the whole class hierarchy with the instance and returns a sorted list of names, not values.
solid answer
~40 s`vars(obj)` is a thin wrapper over `obj.__dict__`: the live, mutable, insertion-ordered mapping of attributes *this instance* stored. It cannot see class-level defaults, computed attributes, or fields declared through `__slots__` — and for a slotted instance it raises `TypeError`, because there is no `__dict__` at all. `dir(obj)` answers a different question: it collects names from the instance and every class in the method resolution order, deduplicates them, **sorts them alphabetically**, and returns names only. It is also advisory — a type may supply its own list — so it is built for interactive discovery, not for a data contract. That ordering difference bites hardest on serializers: a record built from `vars()` follows assignment order, the same record built from `dir()` follows the alphabet.
code
python · 15 linesclass Row:
kind = "message"
def __init__(self):
self.ts = 1
self.user = "al"
self.body = "hi"
@property
def day(self):
return self.ts // 86400
r = Row()
print(list(vars(r)))
print([n for n in dir(r) if not n.startswith("_")])go deeper
Know that vars(obj) hands you the instance's own attribute dictionary of names to values, while dir(obj) gives back a sorted list of names drawn from the object and its classes.
Explain precisely what each one omits: no class attributes or properties from vars, no values from dir, and a TypeError from vars on a slotted instance. Know that vars returns the live dict, not a copy.
Demonstrate the production judgement: reflection-driven serialization encodes an accidental field set and an accidental field order into an output format, so you argue for a declared schema and show what breaks downstream when the two builtins are swapped.
Own the boundary: decide where reflection is legitimate in your systems — debugging dumps and object browsers — versus where an explicit, reviewable schema must own a persisted or transmitted format, and make that rule survive contact with a large codebase.
## Three things that look like the same question `obj.__dict__`, `vars(obj)` and `dir(obj)` are routinely treated as three ways to ask "what attributes does this object have". They answer three different questions, and a serializer that picks the wrong one loses data quietly. **`obj.__dict__`** is the instance's own attribute mapping — the names this particular object stored, mapping to their values. It is a real `dict`: mutable, insertion-ordered (a language guarantee since 3.7), and *live*, so writing into it writes to the object. **`vars(obj)`** is the builtin spelling of exactly that; it returns the same object, not a copy. Called with no argument it returns `locals()` instead. Called on a class it returns a `mappingproxy` — read-only, and containing only that class's own namespace, not anything inherited. **`dir(obj)`** is a *name* query, not a value query. It gathers the instance's own names plus every name reachable through the type's method resolution order, deduplicates, **sorts alphabetically**, and returns a `list[str]`. A type may define its own name list, in which case `dir()` returns whatever that says, sorted — which is why the documentation describes the result as an interactive convenience rather than an authoritative inventory. ## What a vars()-driven serializer misses Walk `vars(obj)` and every one of these vanishes from the output: - **class attributes** — a default such as `kind = "message"` that no instance ever assigned; - **computed attributes** — anything backed by a `property` or another descriptor, since the value is produced on access and was never stored in the instance; - **slot fields** — a class with `__slots__` has no instance `__dict__`, so `vars()` raises `TypeError: vars() argument must have __dict__ attribute` rather than returning an empty mapping; - **inherited stored state** written by a base class only when some branch ran. And it *gains* things you may not want: private bookkeeping, caches, and any attribute a library set on the instance behind your back. ## The ordering trap Here is the failure that is easy to reproduce and easy to miss. A chat-transcript archiver walks each message object and emits one record per message; the run takes about six hours nightly and its output is checksummed downstream to detect re-emitted days. ```python class Row: kind = "message" def __init__(self): self.ts = 1 self.user = "al" self.body = "hi" @property def day(self): return self.ts // 86400 r = Row() list(vars(r)) # ['ts', 'user', 'body'] -> assignment order [n for n in dir(r) if not n.startswith("_")] # ['body', 'day', 'kind', 'ts', 'user'] -> alphabetical, plus class attr and property ``` Swapping one builtin for the other to "also capture the computed fields" changes the key order of every record in the archive. Nothing raises. The nightly run completes, and the downstream checksum declares six hours of correctly-archived transcripts to be different from the previous run's. The lesson generalises: **`vars()` order is the object's assignment history and `dir()` order is the alphabet**, and neither is a contract you should be encoding into an output format at all. ## Doing it deliberately The reliable pattern is to stop asking the object what it has and start telling it what to emit. - If the class is a dataclass, `dataclasses.fields` gives the declared field order, and `dataclasses.asdict` produces a record from it. - If it is not, an explicit tuple of field names on the class — the same list you would have written into `__slots__` — is both the schema and the order, and it fails loudly when a field is renamed. - Where reflection genuinely is the point (a generic debugging dump, an object browser), `inspect.getmembers` walks names and *values* together and takes a predicate, which is closer to what a dump wants than `dir()` plus a loop of `getattr` calls. - To include computed attributes you must actually evaluate them, which means running user code per field, which means deciding what happens when one of them raises. That decision is the real work; choosing between `vars()` and `dir()` is not. ## Two mechanical details worth having ready First, `vars(obj)` hands back the object's live mapping, so mutating the returned dict mutates the object — convenient, and a genuine hazard if you pass it to code that normalises in place. On CPython 3.11 and later instance attributes are stored inline and the `__dict__` object is materialised on demand, so touching `vars(obj)` in a hot loop is not the free read it looks like. Second, `dir()` deliberately omits nothing on grounds of privacy — dunder and underscore-prefixed names come back too, which is why every real use filters them, and why filtering by a leading underscore is a convention rather than a guarantee that the remainder is your public surface.
- What does vars() return for a class rather than an instance?A `mappingproxy` over that class's own namespace: read-only, and containing only what this class defines — not anything inherited from its bases. That is why you can assign a class attribute with `setattr(SomeClass, name, value)` but cannot write into the mapping `vars(SomeClass)` returns. Called with no argument at all, `vars()` returns the current `locals()`.
- If you must serialize computed attributes too, what do you have to decide first?What happens when one of them raises. Including computed fields means executing user code per field, per object, so you need a policy: propagate, skip the field, or record an error marker. You also inherit their cost and any side effects. That policy is the substance of the design — picking a reflection builtin is not.
- Why is dir() described as a convenience rather than an authoritative inventory?Because a type can supply its own name list, so `dir()` reports whatever the type chooses to advertise, sorted; and even by default it merges names from the whole method resolution order without telling you where each came from or whether reading it will succeed. It is built for interactive exploration and completion, which is why production code should read a declared schema instead.
saying these in an interview costs you the question
- Thinks vars(obj) includes class attributes and properties
- Expects vars() to work on an instance with __slots__
- Believes dir() returns values rather than names
- Assumes dir() preserves the order attributes were assigned
- Treats dir() output as an authoritative attribute inventory
- Mutates the dict from vars(obj) without realising it is live