skip to content

How do data and non-data descriptors differ in Python's attribute lookup order?

level: middleimportance: must knowfreq 55%

answer

  1. Two kinds, split by one method
  2. The type is searched first, always
  3. __set__ or __delete__ means it wins
  4. Methods can be shadowed per instance
  5. Data beats instance dict beats non-data

basics

~20 s

A data descriptor defines set or delete and outranks the instance dict; a non-data descriptor defines only get and loses to it. So instance data can shadow a non-data descriptor but never a data one.

solid answer

~50 s

Both live as class attributes, and the split is decided purely by which methods the descriptor's class defines. A **data descriptor** defines `__set__` or `__delete__` (with or without `__get__`); a **non-data descriptor** defines only `__get__`. `object.__getattribute__` resolves `obj.attr` in this order: look the name up on `type(obj)` and its MRO; if that found a **data** descriptor, call its `__get__` and stop; otherwise return `obj.__dict__['attr']` if present; otherwise, if the class attribute is a **non-data** descriptor, call its `__get__`; otherwise return the class attribute as-is. The practical consequences: a method is a non-data descriptor, so `obj.method = something` shadows it per instance, whereas a data descriptor keeps control no matter what is written into the instance dictionary — and a non-data descriptor that writes its result into `obj.__dict__` is thereafter bypassed entirely, which is how caching descriptors get their speed.

code

python · 21 lines
python
class Plain:
    def __get__(self, obj, objtype=None):
        return "from the descriptor"


class Guarded(Plain):
    def __set__(self, obj, value):
        raise AttributeError("read-only")


class Reading:
    plain = Plain()      # non-data descriptor
    guarded = Guarded()  # data descriptor


r = Reading()
r.__dict__["plain"] = "from the instance dict"
r.__dict__["guarded"] = "from the instance dict"

print(r.plain)    # from the instance dict
print(r.guarded)  # from the descriptor

go deeper

for a junior

Recall that both kinds sit on the class, and that defining set or delete makes a descriptor win over anything stored on the instance. Knowing the two names and the direction of the rule is enough here.

for a middle

Walk the full lookup order out loud: type first, data descriptor, instance dict, non-data descriptor, plain class attribute, then AttributeError. Explain why methods must be non-data for per-instance patching to work.

for a senior

Use the rule to diagnose. Be able to explain an attribute that reads fine but is missing from vars(obj), a monkey-patch that mysteriously has no effect, and a caching descriptor that stopped caching after someone added a setter.

for a principal

Own the API consequences: making an attribute a data descriptor removes callers' ability to override it per instance, which affects testability and patchability across the codebase. Decide deliberately which attributes are sealed and which stay open.

### The classification rule Every descriptor is one of two kinds, and nothing about how you *use* it decides which: * **data descriptor** — its class defines `__set__` **or** `__delete__` (usually alongside `__get__`); * **non-data descriptor** — its class defines `__get__` and neither of the other two. A descriptor that defines `__set__` but no `__get__` is still a data descriptor; reads then fall through to the instance dictionary while writes are intercepted, which is a legitimate if unusual write-audit shape. ### The resolution order, in full `obj.attr` on an ordinary object runs `object.__getattribute__`, which does roughly this: 1. Search `type(obj).__mro__` for the name `attr`. Note whether anything was found, and what kind it is. 2. If it was found **and it is a data descriptor**, call its `__get__(obj, type(obj))` and return that. 3. Otherwise, if `attr` is a key in `obj.__dict__`, return that value untouched. 4. Otherwise, if the class attribute found in step 1 is a **non-data descriptor**, call its `__get__`. 5. Otherwise return the plain class attribute found in step 1. 6. If nothing was found anywhere, raise `AttributeError` — which is what invites the class's `__getattr__` fallback to run. Two things about that list are load-bearing. The **type is searched first, always** — the instance dictionary is never consulted before the class, it is merely preferred at step 3 over a *non-data* result. And the whole ordering exists to give a class the power to make an attribute non-overridable per instance, while keeping ordinary attributes cheap. ### Why it is arranged that way Methods are the reason. A plain function is a non-data descriptor, so a method must lose to the instance dictionary — otherwise you could never monkey-patch a single object's behaviour by assigning a callable to it, and every instance attribute would have to be checked against the class's method table first. Meanwhile, a computed or validated attribute must win, or every write would silently drop a plain value into the instance dictionary and permanently shadow the logic you installed. Making the split depend on `__set__` gives you exactly that: an attribute that only reads can be overridden, an attribute that manages writes cannot. ### What it looks like in practice ```python class Plain: # non-data: only __get__ def __get__(self, obj, objtype=None): return "from the descriptor" class Guarded(Plain): # data: adds __set__ def __set__(self, obj, value): raise AttributeError("read-only") class Reading: plain = Plain() guarded = Guarded() r = Reading() r.__dict__["plain"] = "from the instance dict" r.__dict__["guarded"] = "from the instance dict" r.plain # 'from the instance dict' -> step 3 wins r.guarded # 'from the descriptor' -> step 2 wins ``` Both instance-dictionary entries were written by hand, because the *ordinary* assignment `r.guarded = value` would have gone to `__set__`. That is the second half of the story: writes are handled by `object.__setattr__`, which calls the class attribute's `__set__` if it has one and otherwise writes into `obj.__dict__`. A data descriptor therefore controls both directions. ### The consequences interviewers are looking for **Shadowing.** `obj.method = lambda: ...` works and affects only that object, because methods are non-data. The same trick against a validated, `__set__`-bearing attribute either raises or is routed through the setter — the instance dictionary never gets a chance. **The caching pattern.** A non-data descriptor whose `__get__` stores the computed value under the same name in `obj.__dict__` is *skipped* on every subsequent access: the second lookup stops at step 3 and never calls `__get__` again. That is precisely why lazy-attribute descriptors are non-data by design; adding a `__set__` to one would silently destroy the optimisation and make every read pay full price. **Class-level access is a different path.** `Reading.guarded` goes through `type.__getattribute__`, which applies the same rules one level up: it searches the metaclass for a data descriptor, then the class's own MRO, and a descriptor found there is invoked with `obj=None`. **Debugging.** Because step 2 runs real code, `getattr` can raise or block; `inspect.getattr_static` walks the same structures without invoking anything, which is the right tool when you are trying to see what is actually installed. `vars(obj)` shows step 3's dictionary and nothing else, so an attribute that is visibly readable but absent from `vars(obj)` is a strong hint that a descriptor is producing it.

  • Is a class that defines __set__ but not __get__ a data descriptor?
    Yes. The classification asks only whether __set__ or __delete__ is present. With no __get__, writes are intercepted while reads fall through to the instance dictionary, because step 2 needs a __get__ to call and finds none. It is an unusual but valid shape for auditing or validating writes while leaving reads at plain-attribute speed.
  • Why must a lazily-computed attribute descriptor be non-data if it wants to cache?
    Because the cache works by writing the computed value into the instance __dict__ under the same name. On later reads the lookup stops at the instance dictionary and the descriptor is never consulted again, which is the whole speed win. Adding a __set__ would promote it to a data descriptor, so it would win every lookup and the stored value would be ignored on every access.
  • How can you inspect what is really installed on an attribute without triggering it?
    Use inspect.getattr_static, which walks the type's MRO and the instance dictionary and returns the raw object it finds without invoking __get__ or __getattr__. vars(obj) complements it by showing only the instance dictionary. Both are safe on objects whose attribute access has side effects or can block.

saying these in an interview costs you the question

  • Says the instance __dict__ is always checked before the class
  • Thinks a descriptor with only __get__ overrides instance attributes
  • Classifies descriptors by how they are used rather than which methods exist
  • Claims a method cannot be shadowed on a single instance
  • Believes __delete__ has no effect on lookup precedence

context