Why is module-level mutable state rebound with `global` risky in a long-running collector service?
answer
- Ask who owns it and for how long
- It is created at import time
- It survives between tests in one process
- Rebinding leaves earlier aliases behind
- Scope state to a run, not a module
basics
~20 sBecause its lifetime is the whole process, its owner is nobody, and its creation is tied to import order. A global rebinding hides who writes the state, leaks it between tests, and lets it grow for as long as the process lives.
solid answer
~40 sThree concrete costs. **Import-time coupling**: a module-level container is created when the module is first imported, so consumers acquire an import order, a circular import at startup can reach a name that is not bound yet, and `from module import READINGS` copies the reference, so a later `global` rebinding leaves every importer pointing at the old object. **Test isolation**: the value survives from one test to the next inside a process, so suites become order-dependent unless every test patches or resets the attribute. **Lifetime and ownership**: nothing scopes the state to a unit of work, so across a six-hour nightly collection run an accumulator only grows, and no function signature reveals who mutated it. The fix is an object created at the entry point and passed down.
code
python · 11 linesREADINGS = {}
alias = READINGS # what "from module import READINGS" hands you
def reset():
global READINGS
READINGS = {} # rebinds the module attribute only
reset()
print(alias is READINGS) # False
alias["t1"] = 21.5
print(READINGS) # {} - the write went to the old objectgo deeper
Recall that a module-level variable lives as long as the process and is visible to every function in the file. Prefer taking values in as parameters and returning results over writing to a module-level name.
Explain the mechanics behind the advice: the state is created at import time, from module import name copies the reference so a later rebinding is invisible to the importer, and the value persists between tests running in one process.
Diagnose it in a real service. Order-dependent test failures, an accumulator that grows for the length of a long run, and an import cycle that leaves a name unbound at startup all trace back to the same cause; propose the refactor to an object created at the entry point.
Set the boundary for the codebase: which state may live at module scope -- immutable configuration, at most one lazily-built resource behind an accessor -- and which must be owned by an object created at the process entry point. A stated rule is cheaper to enforce than case-by-case review.
### What the phrase means "Module-level mutable state" is a name in a module namespace bound to something mutable -- a dict, a list, a counter -- that functions in the file read and, with a `global` declaration, rebind. It is the path of least resistance: no parameter to thread, no object to construct, importable from anywhere. Take a sensor-telemetry collector as the running example: one module owns `READINGS = {}`, every ingest function writes into it, and a `reset()` function does `global READINGS; READINGS = {}` between runs. It works on the first day and it costs you on every day after. ### Cost 1: the state is created at import time A module-level binding is created the first time the module is imported, and its creation is therefore entangled with import order. Two things go wrong. **Import cycles.** If the collector module imports the sensor-driver module and the driver imports the collector back to reach `READINGS`, one of them runs against a **partially initialised** module. The name may not be bound yet, and a circular import at startup then surfaces as an `ImportError` or an `AttributeError` from a line that looks unrelated -- and only in whichever import order the entry point happens to produce. State that is created inside a function called from `main()` cannot have this problem, because nothing about it depends on which module was imported first. **Aliasing.** `from collector import READINGS` copies the *reference* into the importing module's namespace. A later `global READINGS; READINGS = {}` rebinds the attribute on the collector module only -- every importer still holds the old dict, and now two halves of the system are writing to different objects with no error anywhere. This is the single most common concrete bug in this area, and it is why `import collector` plus `collector.READINGS` is the safer form when module-level mutable state exists at all. ### Cost 2: the state outlives every test Module state is created once per process and then persists. Tests run in one process, so whatever the first test wrote is still there for the second: suites start passing or failing depending on order, and a test that passes alone fails in the suite. The workaround is to patch the module attribute per test, for example with `unittest.mock.patch`, or to call a `reset()` in setup -- which means every test must remember, and a new one that forgets breaks a different test rather than itself. Treat that as a stopgap; the durable fix is state a test can simply construct. ### Cost 3: nothing scopes the lifetime to a unit of work During a six-hour nightly collection run, an accumulator keyed by sensor and timestamp only grows. There is no natural point at which it is discarded, because its lifetime is the process, not the run. You end up adding manual pruning, and manual pruning is a thing to get wrong. When the same state belongs to a run object created at the start of the run, the memory is reclaimed when the run ends, for free, and a second concurrent run is simply a second object rather than a design problem. ### Cost 4: nobody owns it, and no signature says so A function that writes to a module-level dict has a signature that lies: it looks like a pure transformation and it is not. You cannot tell from a call site that state changed, you cannot enumerate the writers without grepping for the `global` statement and every mutating method call, and you cannot reason about a function in isolation. That is a review and onboarding cost that compounds with the size of the module. ### Cost 5: "global" is not as global as it sounds Two mechanisms give you a second copy of the state whether you wanted one or not. `importlib.reload` creates fresh module-level bindings while old references to the previous objects survive. And in Python 3.14 the standard library's multiple-interpreters support (PEP 734) gives each interpreter its own imported copy of a module, so a module-level name is per-interpreter, not per-process. Code that assumes "there is exactly one of these in the process" is making an assumption the language does not guarantee. ### What is fine, and what the fix looks like Module scope itself is not the problem; *mutability plus rebinding* is. Constants -- a timeout, a URL template, a tuple of field names -- are bound once at import, read without any declaration, and never rebound, so none of the costs above apply. The one defensible `global` is a lazily-built resource behind a single accessor function that owns both the read and the write, and even that should expose a way to reset it for tests. Everything else becomes an object created at the entry point and passed down: ```python class Collector: def __init__(self): self._readings = [] def record(self, value): self._readings.append(value) def average(self): return sum(self._readings) / len(self._readings) ``` `main()` builds one, hands it to the ingest functions, and drops it when the run ends. Tests construct their own and need no patching. Import order becomes irrelevant, because nothing happens at import. Two runs can coexist. And the signature of every function that touches the collector now says so.
- Is a module-level constant also a problem?No. A name bound once at import to an immutable value -- a timeout, a URL template, a tuple of field names -- is read with no declaration and never rebound, so none of these costs apply. The problem is mutability plus rebinding, not module scope itself. The giveaway is a `global` statement appearing in the file at all.
- How would you test code that already depends on a module-level dictionary?Patch the module attribute for the duration of each test, with `unittest.mock.patch` or an equivalent setup that saves and restores it, so state cannot leak between tests. Treat that as a stopgap: the durable fix is to pass the container in as a parameter so each test simply constructs its own and no patching is needed.
- What use of `global` would you still accept in review?A lazily-built resource behind a single accessor -- a connection or a cache created on first use and stored in a module-level name -- where the alternative is threading the object through every call site. Keep the read and the write in that one function, and expose a way to reset it so tests are not stuck with whatever the first one built.
saying these in an interview costs you the question
- Says it is fine because there is only one process
- Confuses a module-level constant with mutable shared state
- Assumes `from module import name` sees later rebindings
- Treats order-dependent test failures as a framework problem
- Claims a module-level name is unique across the whole process
- Reaches for `global` purely to avoid passing a parameter