skip to content

An image-thumbnail worker ships as a self-contained bundle carrying the interpreter; it runs on the build machine but dies on the deploy host. How do you diagnose it?

level: seniorimportance: should knowfreq 32%

answer

  1. Same artifact, different machine
  2. Read the error before theorising
  3. What the build scan never saw
  4. Compare the platform baseline of both
  5. One unpack directory, several processes

basics

~20 s

Run the artifact in the foreground on the failing host and read the real error first. The usual causes are a module the build's static scan never saw, uncollected data or metadata, a C library or architecture mismatch, and processes racing over one unpack directory.

solid answer

~50 s

Reproduce it in the foreground on the host and capture the real traceback — everything after that is classification. Four families cover almost all of it. First, imports the build's static scan could not see: anything reached through `importlib.import_module` or a plugin lookup is absent from the artifact and fails only on the path that uses it. Second, files that are not code: package data and `.dist-info` metadata never collected, so a resource read or an `importlib.metadata` version lookup raises. Third, platform: an interpreter and native libraries built against a newer C library or another architecture than the host has, so it dies before Python starts. Fourth, concurrency: a one-file bundle unpacks itself, and a fixed shared unpack path lets simultaneous starts race, one importing a half-written tree. Compare `sys.version` and `sysconfig.get_platform()` on both machines, then fix the build, not the host.

code

python · 9 lines
python
import platform
import sys
import sysconfig

print("version:", sys.version.replace("\n", " "))
print("executable:", sys.executable)
print("platform:", platform.platform())
print("build platform:", sysconfig.get_platform())
print("sys.path head:", sys.path[:3])

go deeper

for a junior

Recall that a bundled artifact contains only what the build put into it, and that the first step on a failing host is to run it in the foreground and read the error rather than guessing.

for a middle

Explain why dynamic imports and data files go missing from a scan-built artifact, and how to tell a Python import failure from a dynamic-loader failure that happens before the interpreter starts.

for a senior

Demonstrate the full triage: classify the failure, compare interpreter and platform fingerprints across machines, spot the concurrent-unpack race behind intermittent corruption, and fix the build rather than patching the host.

for a principal

Own the prevention. Argue for a pinned build environment matched to the oldest supported production baseline, CI that smoke-tests the built artifact on a clean host, and self-describing artifacts so a bad deploy is diagnosed by comparison, not by archaeology.

## Start by refusing to guess A thumbnail worker that starts on the build machine and dies on the deploy host is not a mystery, it is an unread error message. Bundled artefacts are usually launched by a supervisor that swallows stderr, so the first move is to run the artifact by hand on the failing host, in the foreground, with stderr attached, and to keep the exact text. `python -v` style import tracing is available if the bundle exposes its interpreter, and even without it the traceback distinguishes the four families below within seconds. Everything else in this answer is classification; the diagnosis is the transcript. ## Family one: imports the build never saw A bundler that decides what to include by scanning source for `import` statements sees only static imports. A worker that selects a codec or a storage backend at runtime with `importlib.import_module(name)`, or that discovers plugins through entry points, has dependencies that appear nowhere in the source as an import statement — so they are not in the artifact. The symptom is characteristic: the process starts fine, serves some work, and then raises `ModuleNotFoundError` on the first job that takes the uncovered branch. The fix is to make the dependency visible to the scan — a module that imports every plugin explicitly, or the bundler's own include list — and then to prove it with a smoke test that runs the *built artifact*, not the source tree, exercising each branch. ## Family two: everything that is not code The same scan-based collection misses non-code files. Package data (a template, a default config, an ICC profile for the thumbnailer) is not an import and is not collected unless declared. So is `.dist-info` metadata: strip it and `importlib.metadata.version("...")` raises at runtime, often inside a library doing its own version check. A related trap is code that builds a path from `__file__` and opens it; inside a bundle `__file__` points into an unpacked or virtual location that may not hold what the developer's checkout held. Reading resources through `importlib.resources` rather than constructing paths removes a whole class of these. ## Family three: the machine underneath A bundle that carries the interpreter also carries native code, and native code is built against something. If the build ran on a distribution with a newer C library than the deploy host, the executable fails before Python is ever reached — the error comes from the dynamic loader, not from Python, which is a strong signal in itself. CPU architecture mismatches look similar. So does a subtler one on 3.14: the free-threaded build is a distinct ABI with its own tag, so extension modules built for the default build are not interchangeable with a bundle built around the free-threaded interpreter. The structural fix is that the build environment must be at or below the oldest production baseline you support, and it must be pinned rather than 'whatever the build image had this week'. ## Family four: the race This one is specific to single-file bundles, and it is the failure that arrives late and looks haunted. A one-file artifact is an archive with a launcher that unpacks the interpreter and libraries into a temporary directory and then executes from it. If that directory is a *fixed* path rather than a fresh private one, several worker processes starting at the same moment — say a supervisor spinning up eight thumbnail workers at once, or an old and a new release overlapping during a rolling restart — will unpack over each other. One process imports a file another is still writing, and you get a truncated module, an `ImportError` in a package that plainly exists, or a checksum error. It is intermittent, it never reproduces on the build machine where you start one copy by hand, and it correlates with restarts rather than with load. The fix is a per-process unpack directory, or a bundle shape that unpacks once at install time under a version-stamped path so that two releases can never share one. ## Making it not happen again On a three-week release train the artifact is built rarely enough that a broken build is discovered by the deploy, which is exactly the wrong place. Three habits close that gap. Build the artifact in an environment of the same class as production, pinned, so the C-library and architecture baseline is a decision rather than an accident. Test the artifact, not the repository: CI should launch the built bundle on a clean host image and exercise the real code paths, because every failure above is invisible to a test suite run against the source tree. And make the artifact self-describing — have it print its interpreter version, platform tag, and dependency versions on demand, so a failing host is a two-command comparison against the build record instead of an afternoon. The judgement being tested is that you fix the build, never the host. Installing a missing library onto the deploy machine makes the symptom go away and makes the artifact a lie, because the next machine will not have it either.

  • How do you get a module that is only imported dynamically into a bundle whose contents come from a static scan?
    Make it visible to the scan or declare it. Either add a small module that imports every plugin explicitly and is itself imported from the entry point, or list the module in the bundler's include configuration. Then prove it: a smoke test that starts the built artifact and exercises each dynamic branch, run in CI on a clean host image, so a missing module fails the build rather than the deploy.
  • The worker fails only when several copies start at once after a restart. What do you suspect first?
    A shared unpack path. A single-file bundle extracts itself before running, and if that destination is a fixed directory rather than a per-process or version-stamped one, concurrent starts write over each other and one process imports a partially written file. It looks like random corruption, correlates with restarts rather than load, and never reproduces when you launch one copy by hand. Give each process a private extraction directory, or unpack once at install time.
  • How would you tell a C-library mismatch apart from a missing Python module?
    By where the error comes from. A missing module produces a Python traceback ending in `ModuleNotFoundError` or `ImportError`, so the interpreter clearly started. A C-library or architecture mismatch produces an error from the dynamic loader before any Python runs — no traceback at all, often a message naming a symbol or a version. Confirm it by comparing the platform reported by the build machine and the host.

saying these in an interview costs you the question

  • Assumes the bundle carries everything because it started locally
  • Installs missing libraries onto the production host as the fix
  • Ignores C library and CPU architecture differences
  • Treats a missing dynamic import as an interpreter bug
  • Unpacks every process into the same fixed directory
  • Rebuilds repeatedly instead of reading the traceback

context