Why can an empty sys.path[0] plus os.chdir hijack a log-ingest worker's imports?
answer
- The first entry is not pinned at startup
- The empty string follows the process around
- Only names not yet in sys.modules are reachable
- The rollback path imports late
- Absolute paths, safe path mode, eager imports
basics
~20 sThe empty first entry means the working directory at import time, not where the process started. After an os.chdir into a writable spool directory, any module not already in sys.modules can be supplied by a file an attacker dropped there.
solid answer
~50 sA worker started with `python -c` gets `''` as `sys.path[0]`, and `''` is re-resolved against the process's current working directory on every import rather than pinned at startup. The worker chdirs into the spool directory holding a 6,800-row batch, so from that moment the spool is the first place every import looks — ahead of the standard library. Lazy imports make that exploitable: the helper on the partial-failure rollback path has never been imported, so it is not in `sys.modules`, and whoever can write into the spool wins the search with a file of that name. A module body runs on import, so this is arbitrary code inside the worker at exactly the moment error handling starts. Fix it by not chdir'ing into untrusted directories, starting with `-P` or `PYTHONSAFEPATH=1`, and importing eagerly at module scope.
code
python · 13 linesimport os
import pathlib
import sys
sys.path.insert(0, "") # what starting with -c gives the worker
spool = pathlib.Path("spool")
spool.mkdir(exist_ok=True)
(spool / "rollback_helper.py").write_text("print('code from the spool directory ran')\n")
os.chdir(spool)
import rollback_helper # resolved against the NEW working directory
print(rollback_helper.__file__)go deeper
The takeaway to carry: Python searches the current directory first in some launch modes, and importing a module runs its code. Never treat a directory that receives uploads as a place code can be found.
Explain the mechanics: the empty first entry is re-resolved against the working directory at each import, and only modules absent from sys.modules can be captured, which is why a lazy import inside an except block is the reachable one.
Show production judgement: name the vector, order the fixes so the strongest comes first — stop chdir'ing into untrusted directories, start with -P, import eagerly — and describe how you would audit running services for it rather than reasoning from the source alone.
Frame it as a platform default: whether workers ever inherit a working directory they do not own, whether safe path mode is enforced at the launcher rather than per service, and how you prove across a fleet that no writable directory sits on any process's import path.
### The setup A log-ingest pipeline runs worker processes launched from a supervisor as `python -c "from ingest import main; main()"`. Because the interpreter was started with `-c`, `sys.path[0]` is the empty string. Each worker takes a spool directory holding one uploaded batch — say a 6,800-row batch — `os.chdir`s into it so that the parser can use short relative filenames, and processes the rows. When a row fails mid-batch the worker takes a partial-failure rollback path, which begins with `import rollback_helper` inside the exception handler because that module is expensive and rarely needed. Every one of those choices is individually reasonable. Together they are remote code execution. ### Why it is exploitable Three facts compose: 1. **The empty entry is lazy.** `''` on `sys.path` is not resolved once at startup. The path importer interprets it as the process's current working directory at the instant of each import. After the `chdir`, the spool directory is the first place every subsequent import looks — ahead of `PYTHONPATH`, ahead of the standard library, ahead of `site-packages`. 2. **The spool directory is attacker-influenced.** Whoever produces batches can put files in it. It was never meant to be a code directory, so nobody reviewed its write permissions as if it were one. 3. **The import happens later.** `rollback_helper` has never been imported, so it is not in `sys.modules` and no cache short-circuits the search. A file called `rollback_helper.py` dropped into the spool wins, and a module body executes on import — inside the worker, with the worker's privileges, at exactly the moment error handling runs. The attacker does not even need to know your helper's name. Any standard-library module the process has not yet imported is equally available: `csv`, `json`, `secrets`, `traceback` and most of the library are absent from `sys.modules` at startup, so a `json.py` in the spool directory is loaded the first time anything in the process imports `json`. ### What is *not* reachable The attack surface is exactly "names not yet in `sys.modules`", and it is worth being precise about that, because overstating it is a red flag of its own: * Modules already imported — everything loaded at startup (`os`, `codecs`, `site`, `abc`, `stat`, `encodings`) plus everything the worker imported during normal processing — are served from the cache without touching `sys.path`. * Built-in modules listed in `sys.builtin_module_names` are compiled into the interpreter and never reach a path search. Which is why **lazy imports on rare error paths are the sharp edge**. They are the names still resolvable at the moment the process is standing in a directory it does not control. ### Fixing it, in priority order 1. **Do not `chdir` into untrusted directories.** Keep the process's working directory fixed and pass absolute paths to the parser. This removes the vector rather than mitigating it, and it also makes the worker's logs and relative-path handling deterministic. 2. **Start the process without an implicit first entry.** `python -P`, or `PYTHONSAFEPATH=1` in the supervisor's environment for the child, means there is no `''` to follow the `chdir` at all; `sys.flags.safe_path` reads `True` and you can assert on it at startup. Both landed in Python 3.11. 3. **Import eagerly.** Move `import rollback_helper` to module scope so the name is in `sys.modules` before any untrusted directory is in play. This is worth doing regardless: an error path that performs its first import while handling an error is fragile even without an attacker. 4. **Harden the directory.** The spool should be writable only by the uploading identity, and ideally should reject `.py` uploads outright. Treat this as defence in depth, not the fix. ### Auditing an existing fleet Log `sys.path` and `sys.flags.safe_path` once at startup; that log is the only real evidence of what a deployed process searched. Then grep the codebase for `os.chdir` and for `import` statements indented inside functions and `except` blocks, and check whether any directory on the path is writable by a uid other than the deploying one. A group- or world-writable path entry is the finding. The empty entry is the same finding with a moving target, which is why it deserves the flag rather than a permissions audit. ### The shape of a good answer State the mechanism first — the empty first entry is re-resolved against the working directory on every import — then connect it to the lazy import that makes it reachable, then give the fixes in the order that removes the vector rather than papering over it. A candidate who says "so don't chdir, and start it with `-P`, and import the helper at module scope" has demonstrated they have actually operated something like this.
- Which modules are actually reachable by a file dropped into that directory?Only names not already in `sys.modules`. Everything imported during interpreter startup — `os`, `codecs`, `site`, `abc`, `stat`, `encodings` — and everything the worker imported during normal processing is served from the cache with no path search, and built-in modules in `sys.builtin_module_names` never reach the path at all. Standard-library modules not yet loaded, such as `csv`, `json`, `secrets` or `traceback`, are as reachable as your own helper.
- Does starting the worker with -P alone make it safe?It removes the implicit first entry, which is the vector here, so it closes this specific hole. But it does not touch `PYTHONPATH`, whose entries still sit above the standard library, so a process that inherits an environment you do not control is still steerable. Use `-P` together with an explicitly constructed environment, or isolated startup, and assert on `sys.flags.safe_path` at boot so a misconfigured launcher fails loudly.
- How would you audit an existing fleet of workers for this pattern?Log `sys.path` and `sys.flags.safe_path` once at startup — that log is the only proof of what a process actually searched. Then grep for `os.chdir` and for import statements indented inside functions or `except` blocks, and check whether any directory on the path is writable by a uid other than the deploying one. A group- or world-writable path entry is the finding; the empty entry is the same finding with a moving target.
The empty entry is a search location written in pencil that says "wherever I am standing" — walk into someone else's room and you start picking up their tools.
saying these in an interview costs you the question
- Says the empty entry is resolved once at process start
- Claims a dropped file can shadow already-imported modules
- Thinks -P also neutralises PYTHONPATH
- Treats the spool directory as trusted because the app owns it
- Fixes it by renaming the helper rather than the path entry
- Deletes sys.path[0] mid-run and calls that the fix