Why can `a is b` for two equal string literals differ between the Python REPL and a script?
answer
- Same source, different way of running it
- How much source is compiled at once
- Each statement typed is its own unit
- Constants are stored once per code object
- Interning is global and cuts across it
basics
~20 sThe compiler stores each distinct constant once per compiled code object, so two equal literals in one script share an object. The interactive interpreter compiles each statement separately, so each line gets its own constant and identity fails.
solid answer
~50 sIdentity of literals is decided by the unit of compilation. A script's module body compiles to one code object, and the compiler stores each distinct constant once in it, so `a = "hello world"` and `b = "hello world"` on separate lines both load the same object and `is` is True. In the interactive interpreter each statement you enter is compiled and executed on its own, so the two lines produce two constants and `is` is False. The same split explains functions: a constant inside a function body lives in that function's own code object. Auto-interned strings escape the effect entirely, because the intern table is global — `"total"` is the same object everywhere, while `"hello world"` is not. The lesson is that identity of equal immutable values depends on how the code was compiled and entered, so never encode behaviour in it.
code
console · 1 lineprintf 'a = "hello world"\nb = "hello world"\nprint(a is b)\n' > /tmp/t.py && python3 /tmp/t.pygo deeper
Remember the headline: whether two equal literals are the same object depends on how the code was compiled and run, not on the values, so you cannot test it at the prompt and rely on it in a file. Compare text with ==.
Explain the mechanism: the module body compiles to one code object whose distinct constants are stored once, while the interactive interpreter compiles each statement separately, and functions have their own code objects too.
Show where it burns a team: tests that assert identity pass on in-module fixtures and fail on run-time data or after a refactor moves the literal into a helper, and shell experiments that appear to prove something about production code paths.
Own the standard: identity of immutable values is unspecified and compiler-dependent, so it never belongs in program logic or in assertions, and the guard is a lint rule plus warnings-as-errors rather than a review convention.
### The unit of compilation is the thing that varies Python compiles source into code objects before executing anything. A module run as a script becomes one code object for the module body, plus a nested code object per function, class and comprehension. Within a single code object the compiler stores the constants it needs once each: if the same literal value appears twice in that body, both loads fetch the identical object. That is a compile-time deduplication of constants, and it is why a script can print True for values that have nothing to do with the small-integer cache or the intern table. The interactive interpreter changes exactly one thing: it compiles and runs each statement you type as its own unit. Two assignments typed on two lines are two compilations, so each gets its own copy of the constant, and `is` reports False. Type them on one line separated by a semicolon and it is one statement, one code object, one shared constant — True again. ### The same effect in three other places - `python -c` with several statements is a single compilation unit, so it behaves like a script. - `exec(source)` and `compile(source, ...)` compile whatever you hand them as one unit; two separate `exec` calls do not share constants. - A module body and a function body are different code objects, so an equal literal in each is not necessarily the same object. That is four different ways for the same two lines of source to give different `is` answers, none of which is about the values themselves. ### Where interning cuts across it Auto-interning is global and therefore survives compilation-unit boundaries. Identifier-like string constants — ASCII letters, digits and underscores — go through the intern table, so `"total"` typed on two separate REPL lines is the same object, while `"hello world"` is not. Small integers behave the same way for the opposite reason: 256 is a cached object regardless of how many code objects are involved, while 257 is not. So the observable behaviour is the *combination* of three mechanisms — the small-int cache, the string intern table, and per-code-object constant storage — and predicting the answer requires knowing all three. ```console $ printf 'a = "hello world"\nb = "hello world"\nprint(a is b)\n' > /tmp/t.py && python3 /tmp/t.py True ``` Run the identical three lines by typing them into the interactive interpreter and the last one prints False. ### Why this matters beyond trivia First, it is the reason "it works when I try it in the shell" and "it works when I run the file" are not the same evidence. A developer who checks a hypothesis about identity at the prompt and then writes code relying on it in a module has tested a different mechanism than the one that will run. Second, it defeats reasoning about tests. A test that asserts identity of parsed values may pass because the fixture literals sit in one compiled test module and fail once the values arrive from a file at run time, or once the fixture is moved into a helper function with its own code object. The assertion did not change; the compilation layout did. Third, it is a portability trap. Constant deduplication is a CPython compiler behaviour, not a language rule, and the compiler is free to change how aggressively it merges constants between versions. Nothing in the language reference promises that two occurrences of an equal literal are one object. ### The correct mental model Say it as a rule: **for immutable values, equal does not imply identical, and identical is an artefact of caching, interning and compilation.** Use `==` for values. Use `is` only where identity is the actual question — `None`, `True`, `False`, and unique sentinel objects you created precisely so that identity would be meaningful. If you find yourself explaining why an identity check is safe here, you have already found the bug.
- Why does `"total" is "total"` stay True on two separate REPL lines when `"hello world"` does not?Because the compiler interns identifier-like constants, and the intern table is global rather than per code object. `"total"` is ASCII letters only, so both lines fetch the one canonical object. `"hello world"` contains a space, so it is stored as an ordinary constant in whatever code object compiled it.
- Does a constant inside a function share the module-level constant of the same value?Not necessarily. A function body is compiled into its own code object with its own constant storage, so an equal literal there can be a different object from the one in the module body — unless the value is auto-interned or a cached small integer, which are shared globally.
- How would you actually check whether two names refer to the same object?Compare with `is`, or compare `id()` values, but only as a diagnostic. Both answer a question about the current run under the current interpreter, and neither is a fact you may encode in program logic, because caching, interning and compilation layout can all change the answer.
A printer setting a whole book at once reuses a single block for a repeated word; a printer handed one line at a time cannot know the word appeared before, and cuts a new block each time.
saying these in an interview costs you the question
- Claims the interactive interpreter uses different string types
- Says the script result is the correct, guaranteed one
- Explains it purely as interning, ignoring code objects
- Believes equal immutable values are always one object
- Treats shell experiments as proof of module behaviour