skip to content

Why does `a = 1000` then `b = 1000` make `a is b` True in a script but False in the REPL?

level: juniorimportance: nice to knowfreq 24%

answer

  1. Depends on how much source compiled at once
  2. Constants stored once per compilation unit
  3. The shell compiles line by line
  4. A separate cache covers the small integers
  5. Range ends at 256

basics

~20 s

A whole file compiles as one unit, and the compiler stores each distinct constant once, so both names get the same 1000 object. The REPL compiles each entered statement separately, so each line allocates its own.

solid answer

~40 s

Two different mechanisms produce shared objects in CPython. **Constant sharing** happens per compilation unit: a script is compiled in one go and each distinct constant is stored once, so `a` and `b` are bound to the same `1000`. In the REPL every statement is compiled on its own, so the second line has no knowledge of the first and allocates a fresh `1000`. The **small-integer cache** is separate: CPython preallocates one object for every integer from -5 through 256, so `256` typed line by line in the REPL still gives one shared object while `257` does not. Since 3.12 those cached integers are immortal under PEP 683. All of this is a CPython implementation detail and no program should depend on it.

code

pycon · 8 lines
pycon
>>> a = 1000
>>> b = 1000
>>> a is b
False
>>> a = 256
>>> b = 256
>>> a is b
True

go deeper

for a junior

Recall two facts and you have the answer: a file is compiled all at once so a repeated constant is stored once, and small integers up to 256 come from a preallocated cache. Saying the shell compiles each line separately is the key sentence.

for a middle

Explain the two mechanisms separately rather than blurring them, including that constant sharing covers strings too and that identifier-like literals are additionally interned at compile time across modules. Say clearly that none of it is promised by the language.

for a senior

Use it as a debugging pattern: when a quick check in the shell disagrees with the same code in a file, look at the compilation unit before you doubt the values. Be able to explain why this makes identity a useless basis for program logic.

for a principal

The angle to own is portability. Anything that behaves differently on another implementation, another build or another version is a hazard in shared libraries, and you should be able to say where a codebase is allowed to observe such details at all - benchmarks and diagnostics, never behaviour.

### Two separate mechanisms, one confusing symptom Pasting `a = 1000` and `b = 1000` into a file and asking whether the two names point at the same object gives `True`; typing the same two lines one at a time into the REPL gives `False`. Nothing about integers changed between the two runs — what changed is **how much source was compiled at once**. CPython compiles a whole file as a single compilation unit. While compiling, it collects the constants it sees and stores each distinct one exactly once, then hands out references to that single object wherever the literal appears — including inside functions defined in the same file. So in a script, both `a` and `b` are bound to the one `1000` object the compiler created, and the identity check reports `True`. The same is true of a whole line passed to `python -c`, of a single `exec()` of one source string, and of two statements typed on one REPL line separated by a semicolon. The REPL is different in exactly one respect: each statement you enter is compiled and executed on its own. The first line's compilation produces one `1000`; the second line's compilation has no knowledge of the first and produces another. Two objects, equal in value, distinct in identity. ### Why 256 behaves differently from 1000 Now try the same experiment with `256` typed line by line in the REPL: the answer is `True`, and with `257` it is `False`. This is the second mechanism, and it has nothing to do with compilation units. At start-up CPython preallocates one object for every integer from **-5 through 256**, and every operation that would produce one of those values — a literal, arithmetic, a parsed number — returns the preallocated object instead of allocating. The range exists because those values dominate real programs: loop counters, small lengths, flags, indices. `257` falls outside it, so each occurrence is a fresh allocation unless constant sharing happens to catch it. Since PEP 683 in **3.12**, these small integers (along with `None`, `True`, `False` and a set of statically allocated short strings) are *immortal*: their reference counts are pinned to a sentinel value and never move, which removes refcount traffic on the hottest objects in the runtime. That is an implementation detail you can observe but should not depend on. Strings show the same pair of effects. Two identical string literals inside one compilation unit are one object because of constant sharing. On top of that, a literal that looks like an identifier — only ASCII letters, digits and underscores — is put in the interpreter's intern table at compile time, so `"customer_id"` written in two different modules is one object, while `"customer id!"` written in two different modules is two. And a string built at runtime, say by joining pieces, is always a fresh object no matter how many equal strings already exist; `sys.intern()` exists precisely to collapse those. ### What to take from it The practical rule is that *equal is a language guarantee and same-object is not*. Which values are shared depends on the CPython version, on whether the code was compiled as one unit or many, on whether a value came from a literal or from computation, and on which implementation is running. Every one of the results above is reproducible on CPython 3.14 and none of them is promised by the language reference. Two habits follow. First, when a quick experiment in the REPL disagrees with what the same code does in a file, suspect the compilation unit before you suspect the values — this is the single most common cause of "it works in the shell but not in my script" identity confusion. Second, when you genuinely want one shared object — a sentinel, a canonical key, a deduplicated field name — arrange it explicitly with a module-level object or with `sys.intern()`, so the sharing is something your code guarantees rather than something the compiler happened to do for you today.

  • Which integers does CPython preallocate, and why that range?
    Every integer from -5 through 256 gets one object created at start-up and reused everywhere, whatever produced the value. The range covers what real programs use constantly - loop counters, small lengths, indices, flags and a few negative sentinels - so the cache removes an enormous number of allocations. Values outside it are allocated on demand, which is why 257 behaves differently from 256.
  • Do two equal string literals in the same file always end up as one object?
    Within one compilation unit, yes - constant sharing covers strings as well as numbers. Across different modules it depends: a literal that looks like an identifier, meaning only ASCII letters, digits and underscores, is interned at compile time and shared, while something like `"customer id!"` is not. A string built at runtime is never shared automatically.
  • How would you get a guaranteed shared object when you actually need one?
    Create it explicitly. A module-level object that everything imports gives you one instance by construction, and for strings `sys.intern()` gives you a canonical object on purpose. Never lean on constant sharing or the integer cache for that, because which values happen to be shared changes between versions, builds and implementations.

A script is one print run of a book where the typesetter reuses a single plate for a repeated page; the REPL is a new print run per page, so the same page comes out on a fresh plate every time.

saying these in an interview costs you the question

  • Says small integers are cached up to 1000
  • Claims the REPL and a script must always agree
  • Thinks the shared object depends on the value's size in bytes
  • Explains it as a garbage collection effect
  • Treats the sharing as a language guarantee

context