skip to content

Name Mangling and Privacy

A leading double underscore is rewritten by the compiler to carry the class name, so a subclass cannot collide with the base's attribute. Interviewers ask whether that counts as privacy - it does not.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

In Python, what does `_name` mean versus `__name` on a class attribute?

level: juniorimportance: must knowfreq 70%

answer

  1. Two different kinds of "private"
  2. One is a promise, one is a rewrite
  3. The compiler edits the identifier
  4. Class name is prepended
  5. _ClassName__name is readable from anywhere

basics

~20 s

A single leading underscore is pure convention: it marks an attribute as internal and the interpreter ignores it. A double leading underscore inside a class body is rewritten by the compiler to _ClassName__name, which avoids subclass collisions rather than blocking access.

solid answer

~50 s

`_name` is a naming convention and nothing more. Nothing in CPython restricts access to it; the only place the language reacts at all is `from module import *`, which skips module-level names starting with one underscore when the module defines no `__all__`. Reviewers and linters read it as "internal, may change". `__name` written inside a class body is a different mechanism: private name mangling, a compile-time rewrite. An identifier that starts with two or more underscores and does not end in two or more underscores becomes `_ClassName__name`. So `self.__token = "x"` in `class Order` stores the key `_Order__token` in the instance `__dict__`. Its purpose is collision avoidance, so a subclass or a mixin can reuse the same spelling without stomping on the base's attribute. It is not access control: `obj._Order__token` reads it from anywhere. Default to one underscore; reach for two only when you genuinely need the per-class namespace.

code

python · 11 lines
python
class Order:
    def __init__(self):
        self._internal = "convention only"
        self.__token = "mangled"


o = Order()
print(o._internal)          # convention only
print(vars(o))              # {'_internal': ..., '_Order__token': 'mangled'}
print(o._Order__token)      # mangled
print(hasattr(o, "__token"))  # False

go deeper

for a junior

Be ready to state the difference in one breath: one underscore is a convention the interpreter ignores, two underscores are rewritten by the compiler to include the class name. Know that neither one prevents access.

for a middle

Explain the mechanics: the rewrite happens at compile time, produces _ClassName__name, and lands that spelling in the instance __dict__. Say what it is for — subclass collision avoidance — and show you can still read it.

for a senior

Show the judgement of when to spend it. Discuss what mangled names cost you in serialization, debugging and reflective helpers, and why a single underscore is the right default for almost every internal attribute you write.

for a principal

Own the API-boundary argument: what a codebase's underscore conventions promise callers, how they are enforced in review and lint rather than by the language, and when a widely subclassed base class genuinely earns a mangled attribute.

Python has no `private` keyword. What it has instead are two spellings that look alike and operate on completely different levels: a leading **single** underscore, which is a message to human readers, and a leading **double** underscore, which is an instruction to the compiler. ## `_name` — the convention A single leading underscore on a module-level name, a class attribute, a method or a function means "this is internal; I may change it without warning". The interpreter enforces nothing. `obj._buffer` reads fine from anywhere, `dir(obj)` lists it, and `vars(obj)` shows it spelled exactly as written. The one place the language notices the prefix is star-import: `from module import *` skips module-level names that begin with a single underscore — unless the module defines `__all__`, in which case `__all__` alone decides what is exported. Everything else about the convention is enforced socially, by code review and by linters that flag reaching into another object's underscore names. ## `__name` — the rewrite An identifier written textually inside a class body that begins with two or more underscores and does **not** end in two or more underscores is transformed by the compiler into `_ClassName__name`, with the class name's own leading underscores stripped. This is *private name mangling*. It happens while the class body is compiled, so nothing resolves it at run time — by the time the code executes, the identifier simply **is** `_Order__token`: ```python class Order: def __init__(self): self.__token = "x" print(vars(Order())) # {'_Order__token': 'x'} ``` The rewrite applies to every identifier in the class body, not only to attributes: method names, class-level constants, and even bare name references inside a method. A method that reads a module-level global written `__g` compiles to a lookup of `_ClassName__g` and raises `NameError` — a genuine surprise the first time you hit it. ## What mangling is actually for It exists to keep a base class's internal state from colliding with a subclass's. Imagine a base class that stores `self.__state` and a method that reads it. A subclass author who happens to pick the name `__state` for something unrelated does not corrupt the base: the base's code compiled to `_Base__state`, the subclass's to `_Sub__state`, and both keys sit side by side in the instance `__dict__`. The same property is what makes mangled names safe inside a mixin intended to be combined with classes it has never seen. ## What mangling is not It is not privacy, and describing it as "private" in an interview is the classic wrong answer. `obj._Order__token` works from any module. `vars(obj)` and `dir(obj)` display the mangled name. Debuggers, `pickle`, serializers and anything that walks `__dict__` see it. The mangled spelling is also a nuisance in exactly those places: `dataclasses`, ORM-style attribute mapping, and any helper that resolves attributes from strings must know the prefix, because the rewrite only touches identifiers in source code and never a string. ## Which to use Default to one underscore. It communicates intent, keeps the name stable for tests, debuggers and serialization, and costs nothing. Use two only where the collision risk is real: a widely subclassed base class or a mixin whose bookkeeping must not be silently overwritten. Using `__name` everywhere as a Java-style `private` produces mangled keys throughout your dumps and tracebacks while providing no protection whatsoever. ## The neighbouring spellings `__dunder__` names such as `__init__` end in two underscores, so the rule deliberately exempts them — mangling them would break every protocol method. A name with a *single* trailing underscore is not exempt: `__x_` inside `class Trail` still mangles to `_Trail__x_`. A **trailing** single underscore, as in `class_` or `id_`, is a separate convention for avoiding a clash with a keyword or a builtin, and involves no rewriting at all. Finally, if the enclosing class's name consists only of underscores, there is nothing left after stripping them and no mangling occurs. ## Why the language settled here Enforced privacy needs the runtime to check who is asking on every attribute access, and Python's object model deliberately does not have that machinery: attributes are dictionary entries, and any code holding the object can read the dictionary. Rather than fake a guarantee it could not keep, the language shipped the cheap half of the problem — making sure two unrelated authors do not accidentally choose the same attribute name in the same object — and left intent to convention. That is why the honest one-sentence summary of `__name` is "namespaced by class", not "private". ## How it shows up in review The practical signal in a code review is a caller reaching past an underscore. `other._parser.reset()` is a design smell, not an error, and the fix is usually to add a method rather than to add underscores. Conversely, a class whose every attribute is spelled with two underscores is a class whose tracebacks, `repr` output, pickles and test fixtures all carry the class name inside attribute keys for no benefit — that is the pattern to push back on.

  • Does a single leading underscore ever change what Python actually does?
    In one place only. `from module import *` skips module-level names that begin with a single underscore, when the module has no `__all__`; if `__all__` is defined it decides on its own. Explicit imports, attribute access, `dir()` and `getattr` are all unaffected. Everywhere else the prefix is a message to readers and to linters, not to the interpreter.
  • How would you read a mangled attribute from outside its class if you had to?
    Spell the mangled name: `obj._Order__token`, or read `vars(obj)['_Order__token']`. Nothing raises. Do it in a debugger or a one-off script if you must, but not in shipped code — the class author chose `__token` precisely to say the name is theirs to change, and a rename breaks you with no deprecation.
  • Why are `__init__` and `__len__` not mangled?
    The rule only fires on identifiers that start with two or more underscores **and** do not end in two or more underscores. Dunder names end in two, so they are exempt — otherwise every protocol method defined in a class body would be renamed and the interpreter would never find it. A single trailing underscore does not exempt a name: `__x_` inside `class Trail` still becomes `_Trail__x_`.

One underscore is a sign on a door saying "staff only"; two underscores is renaming the door so nobody else's floor plan happens to label a door the same way. Neither one locks it.

saying these in an interview costs you the question

  • Says a double underscore makes an attribute private and unreachable
  • Claims the interpreter blocks access to `_name` from outside the class
  • Thinks mangling happens at attribute-lookup time rather than at compile time
  • Uses `__name` everywhere as a Java-style private modifier
  • Believes `_name` and `__name` differ only in visual emphasis
  • Expects `getattr(obj, '__name')` to find the mangled attribute

context

open as a page

Which class name does Python's `__attr` mangling use, and when is it applied?

level: middleimportance: should knowfreq 45%

basics

~20 s

It uses the name of the class whose body lexically encloses the code, with that name's leading underscores stripped, and applies it while compiling the class body. The runtime type of the object is irrelevant.

open as a page

Why does `getattr(obj, "__cache")` fail when `self.__cache` works inside the class?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Mangling rewrites identifiers in source code, not strings. self.__cache in the class body compiled to self._Worker__cache, while the string handed to getattr is used verbatim, so the lookup asks for a literal __cache that does not exist.

open as a page

Why does `class Cls: __slots__ = ('__buf',)` create a slot named `_Cls__buf`?

level: middleimportance: nice to knowfreq 15%

basics

~10 s

Class creation applies the same private-name mangling to __slots__ entries that the compiler applies to identifiers, so the descriptor is built as _Cls__buf — exactly the name self.__buf compiles to inside the class body.

open as a page