How do you identify code that Python's site startup imports before your script's first line?
answer
- Ask what the interpreter did, not your code
- Compare two environments, not two commits
- One command prints the assembled state
- Small text files that can execute
- Named modules the startup imports if present
basics
~20 sEnumerate what startup added: run python -m site to see the assembled path, the user site directory and site.ENABLE_USER_SITE, then read every .pth file in each site directory and check whether a sitecustomize or usercustomize module is importable there.
solid answer
~40 sWork from the interpreter outward, not from your code inward. First confirm which executable is running (`sys.executable`) and which prefixes it computed (`sys.prefix`, `sys.base_prefix`). Then run `python -m site`, which prints the assembled `sys.path`, `site.USER_BASE`, `site.USER_SITE` and `site.ENABLE_USER_SITE` — that tells you whether a per-user directory is contributing. Next enumerate the `.pth` files in `site.getsitepackages()` and the user site directory and read them: a line beginning with `import` runs before your program. Finally check for a `sitecustomize` or `usercustomize` module on the path, since `site` imports those by name when they exist. Diff that whole picture against the environment where the behaviour does not occur; the difference is the culprit, and it is almost always an environment built differently, not a code change.
code
console · 1 linepython -m sitego deeper
Know that code can run before your script starts, and that printing sys.path and sys.executable is the first thing to do when imports behave unexpectedly on one machine.
Be able to list the startup hooks — an executable line in a .pth, a sitecustomize or usercustomize module, the per-user site directory — and enumerate them with site.getsitepackages and python -m site.
Demonstrate the discipline: run the checks as the failing process runs, diff the two environments rather than the two commits, and follow the diagnosis with idempotent registration and explicit wiring instead of deleting a file.
Own the systemic answer: environments built from a lockfile by a controlled process, startup side effects treated as a supply-chain surface, and a startup assertion that fails loudly when the assembled site directories differ from the expected set.
### The shape of the bug The symptom that sends people here is behaviour that no line of your code explains. An inventory sync between two systems starts writing every record twice in the deployed environment — a duplicated side effect — while the 27-minute test suite stays green and reproduces nothing. A handler is registered twice; the second registration is nowhere in your source. When code appears to run that you did not call, the question stops being "what does my program do" and becomes "what did this interpreter do before my program started". ### What can run before your first line Startup gives four openings, and all of them belong to `site`: 1. **A `.pth` file with an executable line.** In any site directory, a line beginning with `import` is executed at startup. This is the most common source of invisible setup, because a distribution can install one without the application ever mentioning it. 2. **A `sitecustomize` module.** After processing site directories, `site` tries to import a module by that name; if it is importable anywhere on the path, it runs. It is intended for site-wide administrator setup. 3. **A `usercustomize` module**, imported the same way, and only when the per-user site directory is enabled. 4. **Import side effects of anything a `.pth` added**, since appending a directory can change which copy of a module wins. Anything in that list runs with your process's privileges, in every process started by that interpreter, with no mention in your source tree. ### The diagnostic ladder Run each step *with the interpreter that misbehaves*, invoked exactly as production invokes it — an absolute path, the same user, the same environment. Half of these investigations end here, because the answer is that the shell and the service are running different interpreters. 1. **Identify the interpreter.** `sys.executable`, `sys.prefix`, `sys.base_prefix`. Different prefixes mean a different search list before anything else is considered. 2. **Print the assembled state.** `python -m site` gives `sys.path`, the user base and user site directories (annotated when they do not exist) and `ENABLE_USER_SITE`. In a virtual environment created with defaults that flag is `False`; if you see `True`, a per-user directory is in play and is a prime suspect. 3. **Read the `.pth` files.** Glob `*.pth` in each of `site.getsitepackages()` plus `site.getusersitepackages()` and print their contents. They are tiny; read them all rather than guessing. Any line starting with `import` is code that ran. 4. **Look for the customize hooks.** Attempt to import `sitecustomize` and `usercustomize` in a throwaway process and print the resulting module's `__file__` when the import succeeds. Their presence is invisible otherwise. 5. **Check the environment variables that reshape the picture.** `PYTHONPATH` adds directories ahead of the site directories; `PYTHONNOUSERSITE` suppresses the per-user directory. Read the actual environment of the running service, not your shell's. 6. **Diff against the green environment.** Two lists side by side — the CI environment's and production's — reduce the problem to a set difference. The entry present in only one is the answer. ### Why the suite stayed green This class of defect survives testing precisely because the test environment is built differently. A test environment is usually created fresh from a lockfile, in a container, with the per-user directory disabled by the environment's defaults. The deployed one may have been built incrementally, may run as a user with a populated per-user site directory, or may carry an extra distribution installed for an unrelated reason that ships a startup hook. The code is identical; the interpreter's pre-main state is not. A suite can only catch this if it asserts on the environment — for example, a startup check that records the assembled site directories and fails when they differ from the expected set. ### Fixing it properly Finding the `.pth` is the diagnosis, not the fix. Three things follow. Make the registration **idempotent**, so a double invocation cannot double the side effect — a guarded registry keyed by name is enough, and it is worth having regardless. Make the wiring **explicit** in your entry point, so the setup that must happen is visible in code rather than inherited from the environment. And make the environment **reproducible**: build it from a lockfile with a process you control, so the set of installed distributions — and therefore the set of startup hooks — is the same everywhere. The security corollary is worth stating once: anything that can write into a site directory executes code in every process that interpreter starts. That is a supply-chain surface, not merely a debugging curiosity, and it is a good reason to keep environment construction out of hands that do not need it. All of this behaves the same on 3.14 as on earlier supported releases.
- Why would the same code register the handler once in CI and twice in production?Because the two interpreters assembled different search lists. A production environment built incrementally, or running as a user with a populated per-user site directory, can carry a distribution that ships a startup hook the CI environment never installed. The source is identical; the pre-main state is not, which is why a diff of environments beats a diff of commits.
- What is the durable fix once you have found the offending startup hook?Make the registration idempotent so a repeat cannot duplicate the side effect, move the wiring into your entry point so it is visible and ordered, and rebuild the environment from a lockfile so the installed set — and therefore the startup hooks — is identical everywhere. Deleting the file alone leaves the environment free to drift back.
- How do you tell whether the per-user site directory contributed anything?`site.ENABLE_USER_SITE` reports whether it was added, and `site.getusersitepackages()` gives the directory to inspect. In a virtual environment created with the default settings the flag is `False`. Setting `PYTHONNOUSERSITE` in the service's environment suppresses the directory outright when you need to prove it is the cause.
- Why is an executable line in a .pth file a security concern and not just a debugging one?It runs before any of your code, with the process's privileges, in every process that interpreter starts. Write access to a site directory is therefore equivalent to code execution in all of them. That makes environment construction a supply-chain boundary: build it from a controlled process, and keep site directories out of reach of anything that does not need to write there.
saying these in an interview costs you the question
- Blames application code without checking the interpreter
- Assumes the shell and the service share one interpreter
- Ignores the per-user site directory entirely
- Never reads the .pth files actually present
- Deletes the offending file without making registration idempotent
- Treats a green suite as proof the environments match