skip to content

Why does a data-file path built from __file__ break under a zip import?

level: middleimportance: must knowfreq 50%

answer

  1. A loader sets that attribute, not the disk
  2. Not every import comes from a directory
  3. A path that walks through a file
  4. Ask the loader, not the filesystem
  5. files() returns a traversable, not a path

basics

~20 s

Under a zip import, file points inside the archive, not at a real file, so opening a path derived from it fails with an OS error. Read the resource through importlib.resources.files(), which asks the module's loader instead of the filesystem.

solid answer

~40 s

`__file__` is set by whichever loader imported the module, and it is only a usable filesystem path when that loader read from an ordinary directory. When the package is imported from a zip archive -- a zipapp bundle, a zipped path entry, an embedded interpreter -- `__file__` looks like `/opt/app.pyz/converter/__init__.py`, a path that walks *through* a file, so `open()` raises an `OSError` such as `NotADirectoryError`. Some loaders omit `__file__` entirely, and then the same code raises `AttributeError`. The portable answer is `importlib.resources.files("converter")`, which returns a traversable object obtained from the loader: you compose with `.joinpath()` or `/`, and read with `.read_text()`, `.read_bytes()` or `.open()`. Identical code then works from a directory, a zip, or any loader that implements the resource-reading protocol, without the caller knowing which it is.

code

python · 7 lines
python
import json
import os
from importlib.resources import files

fragile = os.path.join(os.path.dirname(json.__file__), "decoder.py")
robust = files("json").joinpath("decoder.py")
print(os.path.exists(fragile), robust.is_file())

go deeper

for a junior

Recall that file tells you where a module came from, and that this need not be a real directory. If you need a packaged data file, reach for the resource API rather than building a path by hand.

for a middle

Explain the mechanism: the loader sets file, a zip import sets it to a location inside the archive, and opening it raises an OS error. Show the replacement with files() and name what a traversable can and cannot do.

for a senior

Bring the deployment view -- why a bundle, a frozen executable or a vendored runtime turns this latent habit into an outage, how you would grep a codebase for it, and what test actually catches it rather than passing against the source checkout.

for a principal

Frame it as an interface decision: application code should depend on a resource abstraction rather than on a filesystem layout, which is what keeps deployment format a reversible choice instead of a constraint baked into every module.

### What `__file__` actually is `__file__` is not a promise of a real file. It is an attribute the loader sets on the module object to describe where it got the code. For the ordinary filesystem loader it is an absolute path you can open, which is why the habit of computing sibling paths from it forms so easily: ```python import os SCHEMA = os.path.join(os.path.dirname(__file__), "schemas", "job.json") ``` This is correct exactly as long as the assumption behind it holds: that the module came from a directory on a mounted filesystem. ### Where the assumption breaks Python can import from places that are not directories. `sys.path` entries may be zip archives, handled by the standard library's zip importer; a `zipapp` bundle is one zip containing the whole application; interpreters embedded in another process may serve modules from a frozen table or a custom loader. In the zip case `__file__` is still set, and that is precisely the trap -- the string exists and looks plausible: ``` /opt/converter.pyz/converter/__init__.py ``` There is no directory named `converter.pyz`; it is a file. Asking the operating system to look inside it fails, typically with `NotADirectoryError` on Linux and macOS and a `FileNotFoundError` on Windows -- both subclasses of `OSError`. The code did not crash at import time and did not crash in your tests. It crashes the first time a request needs the schema, in whatever deployment happens to use an archive. With some loaders `__file__` is absent altogether, and the failure mode changes to `AttributeError` at import time. Neither variant can be fixed by defensive path arithmetic, because there is no path to arrive at. ### The mechanism that replaces it `importlib.resources` inverts the question. Instead of asking the filesystem where the module lives, it asks the *loader* for the resource. `files("converter")` returns a traversable object -- an abstraction with a small path-like surface, whose protocol lives in `importlib.resources.abc`: ```python from importlib.resources import files schema = files("converter").joinpath("schemas/job.json").read_text(encoding="utf-8") ``` The traversable supports `.joinpath()` and the `/` operator, `.is_file()`, `.is_dir()`, `.iterdir()`, `.name`, and the three readers `.read_text()`, `.read_bytes()` and `.open()`. Crucially it is **not** a filesystem path object and does not pretend to be one: there is no `.resolve()`, no `os.path` interoperability, and it may correspond to no path at all. That is the honesty that makes it portable. A zip-backed traversable reads bytes out of the archive; a directory-backed one reads from disk; the calling code is unchanged. The anchor argument is the *importable* package name, not the distribution name -- another place where Python's two senses of "package" cause confusion. A project installed as `document-converter` whose package is `converter` is anchored by `files("converter")`. An older loader-based helper, `pkgutil.get_data(package, resource)`, solves the same problem by returning bytes and is fine for a single read, but it lacks the traversal, directory iteration and text-decoding surface of the newer API. ### Why this is not merely theoretical A document-conversion worker that runs perfectly from a virtual environment can be repackaged as a single-file bundle for deployment, or vendored into a runtime that serves modules from an archive, and nothing about that decision touches your source. The `__file__` habit converts a deployment-format change into an application bug, discovered late and blamed on the packaging step. The same class of failure appears when code is frozen into a single-file executable, when a build tool prunes source directories, or when a resource is read from a namespace package spread across several directories -- a case where "the directory of the module" is not even well defined. ### Migration and testing The change is usually mechanical: replace each path computation with a `files()` anchor plus a read, and delete the module-level constant that held the path. Do the reads lazily rather than at import time, so a missing resource surfaces with a clear error at a defined point rather than during interpreter start-up. Test it by running the test suite against the package installed from a built wheel, and -- if archives are a real deployment target -- by building the bundle in CI and running one smoke command from it. A test that patches `__file__` or reaches into the source tree is testing your checkout, not your artefact. ### The remaining need for a real path Some consumers cannot take bytes: a subprocess that accepts a filename, or a library that wants the operating system to open the file itself. The resource API has a separate mechanism for materialising a traversable as a real path when that is genuinely required -- and it is a context manager precisely because the path may only exist for the duration of the block.

  • What does files() return, and how does it differ from a filesystem path object?
    It returns a traversable: an object with a deliberately small path-like surface -- `.joinpath()` and `/`, `.is_file()`, `.iterdir()`, `.read_text()`, `.read_bytes()`, `.open()`. It is not a `pathlib.Path`, has no filesystem resolution, and may correspond to no path at all, because it might be backed by a zip entry. That narrowness is what lets the same code work from a directory or an archive.
  • Which name do you pass to files() -- the installed distribution name or the import name?
    The import name of the package. A project distributed as `document-converter` whose importable package is `converter` is anchored with `files("converter")`. The two names are routinely different in Python, and passing the distribution name raises a module-not-found error rather than quietly returning nothing.
  • Would a namespace package spread across two directories break the __file__ approach too?
    Yes, and for a related reason: there is no single directory that constitutes the package, so "the directory of the module" is not well defined and the package's `__init__` may not exist as a file at all. Going through the resource API is the only approach that stays coherent, because each portion is resolved by its own loader.

A page number printed in a book is useless once the text is read aloud; you have to ask the reader for the passage rather than reach for a page that no longer physically exists.

saying these in an interview costs you the question

  • Says __file__ is always a real, openable path
  • Fixes zip failures by adding path-normalisation calls
  • Cannot name a situation where an import is not from a directory
  • Treats files() as returning a pathlib.Path
  • Passes the distribution name instead of the import name
  • Extracts the whole archive at start-up to make paths work

context