skip to content

Package Data and Resources

Non-code files that must travel inside the wheel — templates, schemas, certificates — and reading them back with importlib.resources rather than a path built from __file__. The zip-import trap lives here.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

How do you ship a JSON schema file inside your Python package's wheel?

level: juniorimportance: must knowfreq 55%

answer

  1. The backend packs only what it is told
  2. Location inside the package matters
  3. Setuptools declares; hatchling assumes
  4. A package-data glob, or include-package-data
  5. Unzip the wheel and look

basics

~20 s

Put the file inside the importable package directory and declare it as package data in your build backend's configuration. Setuptools needs a package-data entry, or include-package-data together with MANIFEST.in; hatchling and flit ship non-code files under the package by default.

solid answer

~40 s

A wheel contains exactly what the build backend was told to collect, and setuptools collects `.py` files unless told otherwise. First, put the schema **inside** the importable package -- `src/converter/schemas/job.json`, not a sibling top-level `data/` directory -- because only files under a package can be read back through an import-based API. Then declare it: with setuptools, list a glob under the `[tool.setuptools.package-data]` table in `pyproject.toml`, or leave `include-package-data` enabled (its default for pyproject-configured setuptools) so files captured by `MANIFEST.in` or a version-control plugin also land in the wheel. Hatchling and flit take the opposite default and include everything under the package that is not ignored by version control. Finally, verify rather than trust: build the wheel, list its entries, and confirm the file is there.

code

python · 5 lines
python
from importlib.resources import files

# any installed package behaves the same way as your own
resource = files("json").joinpath("decoder.py")
print(resource.is_file(), resource.name)

go deeper

for a junior

Be ready to say that non-code files must live inside the package directory and be declared to the build backend, and that a wheel is a zip you can list to check. Knowing the file does not travel automatically is the point.

for a middle

Explain the mechanics: which backend you use, the package-data declaration versus include-package-data, and why the source tree passing tests proves nothing about the wheel. Show how you would verify the built artefact.

for a senior

Demonstrate the production habit -- an install-from-wheel smoke test in CI, resources read through the import-based API rather than a computed path, and a story for how a missing resource would surface loudly at startup rather than on one unlucky request.

for a principal

Own the convention across many repositories: one project template with the packaging rules pre-wired, a shared CI job that installs the built wheel and exercises resources, and a rule about what belongs in a distribution at all versus what belongs in configuration.

### The wheel is a manifest, not a snapshot of your checkout A wheel is a zip archive whose contents were *chosen* by a build backend. Nothing gets in because it happened to be in your repository. Your test suite passes locally because it imports from the source tree, where every file on disk is reachable; the installed wheel is a different filesystem, and a non-code file only exists in it if some rule put it there. This is why "works on my machine, missing in production" is the classic shape of a package-data bug. ### Rule one: the file must live inside the importable package There are two different senses of "package" in Python, and this is where they meet. The *importable* package is the directory with the module code -- `converter/`, containing `__init__.py`. The *distribution* is the thing you install by name. Package data is data attached to the importable package, so the file has to sit under that directory: ``` src/ converter/ __init__.py schemas/ job.json ``` A top-level `data/job.json` next to `pyproject.toml` has no package to belong to. It can be added to a source distribution, and legacy `data_files` can scatter it somewhere under the installation prefix, but there is no reliable, relocatable way to find it again at runtime -- it may land in a different place under a virtual environment, a system install, or a zip. Keep resources under the package. ### Rule two: declare it, unless your backend assumes it Backends differ, and knowing which one you use is the whole answer: * **setuptools** ships only Python modules by default. Declare the data explicitly: ```toml [tool.setuptools.package-data] converter = ["schemas/*.json", "templates/**/*.html"] ``` Alternatively, leave `include-package-data` at its default of true for pyproject-based configuration and let `MANIFEST.in` (or a version-control plugin that knows which files are tracked) decide which files inside packages get included. * **hatchling** and **flit** invert the default: everything inside the package directory that is not ignored by version control ships, and you subtract with exclusions rather than adding with inclusions. * **poetry-core** sits in between, with an include list in its own table. Because the defaults are opposite, "it worked in my last project" is not a reason -- read the backend you actually declared in `[build-system]`. ### Rule three: verify the artefact, not the intention Build the distributions and look inside. A wheel is a zip, so the standard library can list it: ```console python -m zipfile --list dist/converter-1.0.0-py3-none-any.whl ``` A source distribution is a gzipped tar, and the two need checking separately -- a file can be present in one and missing from the other. The strongest form of this check is an install test: create a clean virtual environment, install *from the built wheel* rather than from the source tree, and run a smoke test that imports the package and reads the resource. Running that in CI is what turns a silent packaging mistake into a red build instead of a production incident, and it is cheap: one job, a few seconds. ### Anti-patterns worth naming Computing a path from `__file__` and opening it works on a normal disk install and fails the moment the code runs from an archive, so read resources through the import-based resource API instead. Reaching for `data_files` to place configuration under `/etc` breaks virtual environments. Committing a generated file and assuming the build regenerates it is another version of the same trust problem. And listing a data file only in a documentation table, rather than in the build configuration, ships a wheel that imports fine and then raises a file-not-found error on the first request that needs the schema. ### What this buys you With the file inside the package and declared, one wheel carries code and resources together, versioned as a unit. There is no second deployment step to copy templates or schemas, no environment variable pointing at a data directory, and no divergence between the schema the code expects and the schema the host happens to have.

  • What happens to a file you put in a top-level data/ directory instead of inside the package?
    It can be captured into the source distribution, but it has no importable package to attach to, so the import-based resource API cannot find it. Legacy `data_files` can install it somewhere under the environment prefix, but that location varies between a virtual environment, a system install and a relocated tree, so runtime code cannot reliably compute it. Move the file under the package directory instead.
  • How would you catch a missing data file in CI before it reaches users?
    Build the wheel, then install *that wheel* into a clean virtual environment -- not the source tree, and not an editable install -- and run a smoke test that imports the package and reads each resource. Listing the archive entries is a useful second check, but the install test is the one that catches path assumptions as well as missing files.
  • Does adding a data file require a version bump before republishing?
    Yes. Published versions are immutable, so a wheel missing a resource cannot be replaced in place; you release a new patch version and, if the broken one is actively harmful, yank the old one so resolvers stop selecting it for new installs while existing pins still work.

Packing a suitcase from a list, not by tipping the wardrobe in: whatever is not on the list stays home, however obviously you meant to bring it.

saying these in an interview costs you the question

  • Assumes every file in the repository lands in the wheel
  • Puts resources outside the package and expects them to install
  • Believes MANIFEST.in alone controls the wheel's contents
  • Only ever tests from the source checkout, never an installed wheel
  • Reaches for data_files to ship package resources
  • Opens the file via a path built from __file__

context

open as a page

Why does a data-file path built from __file__ break under a zip import?

level: middleimportance: must knowfreq 50%

basics

~20 s

Under a zip import, file points inside the archive, not at a real file, so opening a path derived from it fails with an OS error. Read the resource through importlib.resources.files(), which asks the module's loader instead of the filesystem.

open as a page

Why does a file listed in MANIFEST.in appear in the sdist but not the wheel?

level: middleimportance: should knowfreq 35%

basics

~20 s

MANIFEST.in tells setuptools which extra files to add to the source distribution. The wheel is built from package-data rules instead, so unless include-package-data is enabled and the file sits inside a package directory, the wheel leaves it out.

open as a page

When do you need importlib.resources.as_file, and what does it cost a worker calling it 1,200 times a minute?

level: seniorimportance: should knowfreq 25%

basics

~20 s

Use as_file when something outside Python needs a real filesystem path -- a subprocess argument, or a library that opens the file itself. It is a context manager: an archived resource is extracted per entry, so per-request use copies per request.

open as a page