skip to content

Why is os.chdir the wrong way to point each subprocess.run child at its own directory in a threaded worker?

level: seniorimportance: should knowfreq 35%

answer

  1. It is not per-thread state
  2. One directory for the whole process
  3. Threads race between chdir and exec
  4. Let the child do the change, per spawn

basics

~20 s

The working directory is process-global, not per thread, so concurrent workers overwrite each other's value and a child can start in another job's directory. Pass cwd= to subprocess.run instead: that change of directory happens inside the child, once per spawn.

solid answer

~50 s

There is exactly one working directory per process. `os.chdir` mutates it for every thread at once, so a worker that chdirs and then spawns is racing: between its `os.chdir` and its exec, another thread can chdir somewhere else, and the child inherits whichever value was current at the instant of the exec. In a warehouse pick-list builder that runs one helper per order from a thread pool, that shows up as a duplicated side effect - two pick lists written into one order's directory and none into another's. The fix is to stop treating a process-global as job state: pass `cwd=` to `subprocess.run`, which changes directory in the child after the fork and before the exec, so it is private to that child and the parent's own `os.getcwd()` never moves. Make the paths you hand the helper absolute, and chdir the parent at most once, at startup, before threads exist.

code

python · 10 lines
python
import os
import subprocess
import sys
import tempfile

with tempfile.TemporaryDirectory() as order_dir:
    out = subprocess.run([sys.executable, "-c", "import os; print(os.getcwd())"],
                         cwd=order_dir, capture_output=True, text=True, check=True)
    print("child ran in:", out.stdout.strip())
    print("parent still in:", os.getcwd())

go deeper

for a junior

Know that subprocess.run takes a cwd argument that decides where the child starts, and that os.chdir moves your own program instead. Reach for cwd whenever a helper needs to run somewhere specific.

for a middle

Explain the difference in mechanism: cwd is applied on the child's side of the spawn, so it is private to that child, while os.chdir mutates one value shared by the whole process. Note that a relative program path is resolved against cwd, but a bare name still goes through PATH.

for a senior

Diagnose the concurrency failure it causes - interleaved chdirs producing a child in the wrong directory, seen as duplicated or missing output under load - and show the discipline that prevents it: absolute paths resolved when the job is accepted, per-spawn cwd and env, no process-global mutation once threads are running.

for a principal

Generalise the rule and hold the line on it: process-global attributes - directory, environment, umask, signal dispositions, locale - are never per-job state in a concurrent service, and the spawn boundary is the only place a private copy is genuinely made. Decide where that rule is enforced so every service inherits it.

**One directory per process.** The working directory is a property of the process, kept by the kernel, and threads do not get their own. Calling `os.chdir` from a worker thread relocates the whole program, including every other thread's relative `open`, every relative path a library resolves, and every child that any thread execs from then on. It is exactly as global as the environment, and it fails in the same way. **The race, concretely.** A pick-list builder handles orders from a thread pool. Each worker does the obvious thing: ```python os.chdir(order_dir) # process-global subprocess.run(['picklist-helper', order_id]) ``` Two workers interleave: A chdirs to order 4181, B chdirs to order 4182, A execs. A's helper now runs in 4182's directory and writes its pick list there. When B's helper follows, 4182 holds two pick lists and 4181 holds none - a duplicated side effect whose cause is invisible in either helper's code, because neither helper did anything wrong. The same run showed the helper's local cache hit rate fall from a steady 83% to almost nothing, since its cache path was relative too and every invocation looked in whatever directory it happened to land in. The window is microseconds wide, which is why it appears under load and never in a test. **Why `cwd=` is a different mechanism.** `subprocess.run(argv, cwd=order_dir)` does not touch the parent at all. The directory change is performed on the child's side of the fork, before the exec (or handed to the platform's spawn call), so it applies to that one child and to nothing else. The parent's `os.getcwd()` is unchanged, other threads are unaffected, and two spawns with two different `cwd` values cannot interfere no matter how they interleave. The argument accepts a string, bytes or a path-like object. **What `cwd` does and does not affect.** It sets the directory the child starts in, so the child's relative paths resolve there. It also determines how a *relative program path* is resolved: `['./picklist-helper']` is looked for under `cwd`, not under the parent's directory. It does **not** change the lookup of a bare command name, which still goes through `PATH`; the only interaction is that a relative entry in `PATH` - including the empty entry that `os.defpath` starts with - would be interpreted relative to `cwd`. And it is not a sandbox: the child can chdir wherever it likes afterwards, and an absolute path in its arguments ignores `cwd` entirely. **Diagnosing it rather than guessing.** Have the child tell you where it is. Spawning the interpreter with `-c 'import os; print(os.getcwd())'` under the same `cwd` argument as the real helper takes ten seconds and settles the question. In the failing system, log the job id together with the `cwd` value you passed, and the duplicate becomes obvious: two consecutive spawns with different job ids and the same directory. A reproduction is a thread pool of two workers where the parent chdirs between submissions. **The general rule.** Process-global state is not job state. The working directory, the environment, the umask, signal dispositions and the locale are all attributes of the process; using any of them to carry per-job information in a concurrent program creates the same class of race. Every one of them has a per-child equivalent at the spawn boundary - `cwd` and `env` are two arguments of the same call - and that boundary is where per-job configuration belongs, because a fork-and-exec is the one moment where a private copy is genuinely created. **When the parent really must move.** Some legacy libraries insist on relative paths and will not take a directory. Do the chdir once at startup, before any thread is created, and never again; or push that library into its own single-purpose process so its global state is its own problem. If you only need a directory-relative file operation rather than a directory-relative *process*, several `os` functions accept a directory descriptor, which scopes the operation without moving anything. **Absolute paths as the quiet fix.** Most of these bugs disappear before they start if the parent resolves everything to absolute paths at the point it accepts a job - the order directory, the output file, the helper binary - and passes those. `cwd=` then serves the child's own convenience rather than being load-bearing, and a mistake becomes a wrong path in a log line instead of a silent write into the wrong directory.

  • A legacy helper insists on being run from its own directory. What do you do?
    Give it what it wants per call: pass `cwd=` on each spawn, which is private to that child, and keep every other path you hand it absolute. Do not chdir the parent to satisfy it. If a library rather than a subprocess forces the parent to move, do that chdir once during startup before threads exist, or isolate the library in its own process so its global state cannot reach anyone else.
  • How would you prove the race instead of guessing at it?
    Make the child report its own directory - spawn the interpreter with a one-liner printing `os.getcwd()` under the same argument - and log the job id next to the directory you intended for it. Two consecutive spawns with different job ids and the same directory is the proof. To reproduce deliberately, drive two workers through a thread pool with a chdir between submissions and watch the outputs land in one place.
  • Does passing cwd also fix a command that cannot be found?
    No. A bare command name is still resolved through `PATH`, not through `cwd`, so a `FileNotFoundError` there is an environment problem rather than a directory one. `cwd` matters for the program path only when that path is relative, such as `./picklist-helper`, which is looked for under the directory you passed. Resolving the program to an absolute path keeps the two concerns separate.

Using os.chdir before a spawn is moving the whole company to another floor because one employee needs a file there; cwd= sends that one employee.

saying these in an interview costs you the question

  • Treats the working directory as thread-local state
  • Calls os.chdir before each spawn in a threaded worker
  • Thinks cwd= also moves the parent process
  • Uses relative output paths across concurrent jobs
  • Assumes cwd= fixes a bare command name lookup
  • Guards os.chdir with a lock and calls the design fine

context