skip to content

Why can an empty sys.path[0] plus os.chdir hijack a log-ingest worker's imports?

level: seniorimportance: should knowfreq 35%

answer

  1. The first entry is not pinned at startup
  2. The empty string follows the process around
  3. Only names not yet in sys.modules are reachable
  4. The rollback path imports late
  5. Absolute paths, safe path mode, eager imports

basics

~20 s

The empty first entry means the working directory at import time, not where the process started. After an os.chdir into a writable spool directory, any module not already in sys.modules can be supplied by a file an attacker dropped there.

solid answer

~50 s

A worker started with `python -c` gets `''` as `sys.path[0]`, and `''` is re-resolved against the process's current working directory on every import rather than pinned at startup. The worker chdirs into the spool directory holding a 6,800-row batch, so from that moment the spool is the first place every import looks — ahead of the standard library. Lazy imports make that exploitable: the helper on the partial-failure rollback path has never been imported, so it is not in `sys.modules`, and whoever can write into the spool wins the search with a file of that name. A module body runs on import, so this is arbitrary code inside the worker at exactly the moment error handling starts. Fix it by not chdir'ing into untrusted directories, starting with `-P` or `PYTHONSAFEPATH=1`, and importing eagerly at module scope.

code

python · 13 lines
python
import os
import pathlib
import sys

sys.path.insert(0, "")  # what starting with -c gives the worker

spool = pathlib.Path("spool")
spool.mkdir(exist_ok=True)
(spool / "rollback_helper.py").write_text("print('code from the spool directory ran')\n")

os.chdir(spool)
import rollback_helper  # resolved against the NEW working directory
print(rollback_helper.__file__)

go deeper

for a junior

The takeaway to carry: Python searches the current directory first in some launch modes, and importing a module runs its code. Never treat a directory that receives uploads as a place code can be found.

for a middle

Explain the mechanics: the empty first entry is re-resolved against the working directory at each import, and only modules absent from sys.modules can be captured, which is why a lazy import inside an except block is the reachable one.

for a senior

Show production judgement: name the vector, order the fixes so the strongest comes first — stop chdir'ing into untrusted directories, start with -P, import eagerly — and describe how you would audit running services for it rather than reasoning from the source alone.

for a principal

Frame it as a platform default: whether workers ever inherit a working directory they do not own, whether safe path mode is enforced at the launcher rather than per service, and how you prove across a fleet that no writable directory sits on any process's import path.

### The setup A log-ingest pipeline runs worker processes launched from a supervisor as `python -c "from ingest import main; main()"`. Because the interpreter was started with `-c`, `sys.path[0]` is the empty string. Each worker takes a spool directory holding one uploaded batch — say a 6,800-row batch — `os.chdir`s into it so that the parser can use short relative filenames, and processes the rows. When a row fails mid-batch the worker takes a partial-failure rollback path, which begins with `import rollback_helper` inside the exception handler because that module is expensive and rarely needed. Every one of those choices is individually reasonable. Together they are remote code execution. ### Why it is exploitable Three facts compose: 1. **The empty entry is lazy.** `''` on `sys.path` is not resolved once at startup. The path importer interprets it as the process's current working directory at the instant of each import. After the `chdir`, the spool directory is the first place every subsequent import looks — ahead of `PYTHONPATH`, ahead of the standard library, ahead of `site-packages`. 2. **The spool directory is attacker-influenced.** Whoever produces batches can put files in it. It was never meant to be a code directory, so nobody reviewed its write permissions as if it were one. 3. **The import happens later.** `rollback_helper` has never been imported, so it is not in `sys.modules` and no cache short-circuits the search. A file called `rollback_helper.py` dropped into the spool wins, and a module body executes on import — inside the worker, with the worker's privileges, at exactly the moment error handling runs. The attacker does not even need to know your helper's name. Any standard-library module the process has not yet imported is equally available: `csv`, `json`, `secrets`, `traceback` and most of the library are absent from `sys.modules` at startup, so a `json.py` in the spool directory is loaded the first time anything in the process imports `json`. ### What is *not* reachable The attack surface is exactly "names not yet in `sys.modules`", and it is worth being precise about that, because overstating it is a red flag of its own: * Modules already imported — everything loaded at startup (`os`, `codecs`, `site`, `abc`, `stat`, `encodings`) plus everything the worker imported during normal processing — are served from the cache without touching `sys.path`. * Built-in modules listed in `sys.builtin_module_names` are compiled into the interpreter and never reach a path search. Which is why **lazy imports on rare error paths are the sharp edge**. They are the names still resolvable at the moment the process is standing in a directory it does not control. ### Fixing it, in priority order 1. **Do not `chdir` into untrusted directories.** Keep the process's working directory fixed and pass absolute paths to the parser. This removes the vector rather than mitigating it, and it also makes the worker's logs and relative-path handling deterministic. 2. **Start the process without an implicit first entry.** `python -P`, or `PYTHONSAFEPATH=1` in the supervisor's environment for the child, means there is no `''` to follow the `chdir` at all; `sys.flags.safe_path` reads `True` and you can assert on it at startup. Both landed in Python 3.11. 3. **Import eagerly.** Move `import rollback_helper` to module scope so the name is in `sys.modules` before any untrusted directory is in play. This is worth doing regardless: an error path that performs its first import while handling an error is fragile even without an attacker. 4. **Harden the directory.** The spool should be writable only by the uploading identity, and ideally should reject `.py` uploads outright. Treat this as defence in depth, not the fix. ### Auditing an existing fleet Log `sys.path` and `sys.flags.safe_path` once at startup; that log is the only real evidence of what a deployed process searched. Then grep the codebase for `os.chdir` and for `import` statements indented inside functions and `except` blocks, and check whether any directory on the path is writable by a uid other than the deploying one. A group- or world-writable path entry is the finding. The empty entry is the same finding with a moving target, which is why it deserves the flag rather than a permissions audit. ### The shape of a good answer State the mechanism first — the empty first entry is re-resolved against the working directory on every import — then connect it to the lazy import that makes it reachable, then give the fixes in the order that removes the vector rather than papering over it. A candidate who says "so don't chdir, and start it with `-P`, and import the helper at module scope" has demonstrated they have actually operated something like this.

  • Which modules are actually reachable by a file dropped into that directory?
    Only names not already in `sys.modules`. Everything imported during interpreter startup — `os`, `codecs`, `site`, `abc`, `stat`, `encodings` — and everything the worker imported during normal processing is served from the cache with no path search, and built-in modules in `sys.builtin_module_names` never reach the path at all. Standard-library modules not yet loaded, such as `csv`, `json`, `secrets` or `traceback`, are as reachable as your own helper.
  • Does starting the worker with -P alone make it safe?
    It removes the implicit first entry, which is the vector here, so it closes this specific hole. But it does not touch `PYTHONPATH`, whose entries still sit above the standard library, so a process that inherits an environment you do not control is still steerable. Use `-P` together with an explicitly constructed environment, or isolated startup, and assert on `sys.flags.safe_path` at boot so a misconfigured launcher fails loudly.
  • How would you audit an existing fleet of workers for this pattern?
    Log `sys.path` and `sys.flags.safe_path` once at startup — that log is the only proof of what a process actually searched. Then grep for `os.chdir` and for import statements indented inside functions or `except` blocks, and check whether any directory on the path is writable by a uid other than the deploying one. A group- or world-writable path entry is the finding; the empty entry is the same finding with a moving target.

The empty entry is a search location written in pencil that says "wherever I am standing" — walk into someone else's room and you start picking up their tools.

saying these in an interview costs you the question

  • Says the empty entry is resolved once at process start
  • Claims a dropped file can shadow already-imported modules
  • Thinks -P also neutralises PYTHONPATH
  • Treats the spool directory as trusted because the app owns it
  • Fixes it by renaming the helper rather than the path entry
  • Deletes sys.path[0] mid-run and calls that the fix

context