skip to content

Your compiled package must ship wheels for Linux, macOS and Windows on x86-64 and arm64 — how do you build that matrix?

level: seniorimportance: nice to knowfreq 18%

answer

  1. Count the axes before buying machines
  2. One job per platform, versions inside
  3. Build, repair, then install and import
  4. Missing hardware: native, cross, emulate
  5. Collapse an axis with a stable ABI

basics

~20 s

Run one CI job per operating system and architecture, and inside each job loop the interpreter versions with a wheel-building tool that builds, repairs and smoke-tests each wheel. Use native runners where they exist, cross-compilation or emulation where they do not.

solid answer

~50 s

The grid has four axes: operating system, processor architecture, interpreter version, and ABI variant — and on Linux the compatibility baseline as well. You do not want a machine per cell. The standard shape is a CI matrix with one job per operating-system-and-architecture pair, running a wheel-building tool (cibuildwheel is the common one) that iterates the interpreter versions inside the job, builds in a controlled image on Linux so the baseline is fixed, runs the repair step, and imports the repaired wheel in a fresh environment. Missing hardware is covered by cross-compilation where the toolchain supports it and by emulation where it does not — emulation being slow enough that a heavy test step run inside those legs will start timing out. You shrink the grid rather than scale it: a stable-ABI wheel removes the interpreter-version axis, and an sdist covers the platforms you choose not to build for.

go deeper

for a junior

Know that a compiled package needs a separate wheel per operating system, architecture and interpreter version, which is why maintainers publish many files for one release rather than a single archive.

for a middle

Be able to sketch the pipeline: a job per platform, the interpreter versions looped inside it by a wheel-building tool, and each wheel built, repaired and then imported from a clean environment before it is published.

for a senior

Show the operating judgement: choose native runners over emulation, keep heavy suites out of the wheel jobs so slow legs stop timing out, and know which axes you can collapse rather than scale.

for a principal

Set the support policy: which platforms and interpreter versions the project promises, what an sdist-only platform means for users, and whether the maintenance cost of a stable-ABI build is worth the release cadence it buys.

## Count the axes before you build anything A compiled package's release artefacts are a cross product: - **Operating system** — Linux, macOS, Windows. - **Architecture** — x86-64 and arm64 at minimum today. - **Interpreter version** — every minor release you support, because an ordinary extension's ABI changes with each one. - **ABI variant** — the free-threaded build, officially supported since 3.14, is a distinct ABI and therefore a distinct set of wheels. - **Linux userspace** — a glibc-based baseline and, if you support it, a musl-based one. Five supported interpreter versions across three operating systems and two architectures is thirty wheels before you have thought about variants. The engineering question is not "how do I get thirty machines" but "how few axes can I honestly ship". ## The standard pipeline shape One CI job per operating-system-and-architecture pair, and inside each job a wheel-building tool that loops the interpreter versions. That tool — cibuildwheel is the one most projects reach for — does the same four things in every cell: provision the interpreter, build the wheel, run the platform's repair step so the wheel vendors its shared libraries, and install the repaired wheel into a clean environment to import and test it. Keeping that loop inside a single job is what makes the matrix manageable: the CI configuration expresses only the platform axis, and the interpreter axis lives in the tool's configuration next to the code. On Linux the build runs inside a controlled image whose system libraries are deliberately old, because the oldest libraries you link against define the newest baseline you can claim. On macOS the deployment target you set decides which OS releases can load the result. On Windows the toolchain is the platform's C++ build tools and the repair step injects the DLL search path the loader has no run-path equivalent for. ## Covering hardware you do not have Three techniques, in descending order of preference: 1. **Native runners.** Hosted arm64 Linux and Apple silicon macOS runners exist now, and a native job is both faster and a truer test than the alternatives. Prefer them wherever your provider offers them. 2. **Cross-compilation.** macOS builds universal or per-architecture binaries from one host easily; other targets vary with how well the project's build system handles a cross toolchain. Cross-compiled wheels build fast but cannot be *tested* on the build host — you have compiled for a machine you are not running on — so the test step must be skipped or deferred to a real device. 3. **Emulation.** A user-space emulator lets an x86-64 runner build and run arm64 artefacts. It works, and it is roughly an order of magnitude slower. That last property is where matrices break. A project whose wheel job runs a 340-case regression pack inside every cell will find the emulated arm64 leg intermittently timing out while every native leg is green — the classic signature being a failure that moves between cases from run to run rather than sticking to one. The fix is not a longer timeout. Either move that leg to a native runner, or split the concern: run the full regression pack once on the native platform, and inside the wheel jobs run only a smoke test that imports the extension and exercises one code path from each compiled module. The wheel job's question is "does this artefact load and work on this platform", not "is the library correct". ## Shrink the grid instead of scaling it - **A stable-ABI wheel collapses the interpreter axis.** Built against the limited C API, one wheel serves the current interpreter release and later ones, which is the difference between rebuilding on every Python release and not. It costs you access to the parts of the C API outside that subset, and some measurable performance in exchange, so it is a real decision rather than a free win. - **Publish an sdist as the long tail.** For architectures and userspaces you decline to build for, a buildable sdist lets determined users compile. Say so in the project's documentation rather than leaving them to discover it from a compiler error. - **Drop interpreter versions on schedule.** Supporting versions past their upstream end of life multiplies every other axis for a shrinking number of users. - **Be explicit about the free-threaded leg.** It is a separate ABI and therefore separate wheels; decide deliberately whether you support it rather than discovering the gap from a user's source-build failure. ## Make the release reproducible One pipeline should produce every artefact for a tag, collect them as a single set, and publish them together, so a release is not half-built when one leg fails. Wheels that trickle out per platform over days leave users on the unlucky platform compiling from source in the meantime, which is precisely the experience the matrix exists to prevent.

  • Why can a cross-compiled wheel be built but not tested in the same job?
    Because the artefact targets a machine the runner is not. The build host can produce the binary but cannot load it, so the usual build-repair-import loop stops after repair. You either skip the test on that leg and accept the reduced signal, or route the artefact to a native runner or a real device for the import check before you publish it.
  • What do you give up by shipping a stable-ABI wheel to collapse the interpreter axis?
    Access to the parts of the C API outside the limited subset, which for some extensions rules the approach out entirely, and typically some performance, since the stable interface hides the concrete structures that a version-specific build can touch directly. In exchange one wheel serves future interpreter releases, so a new Python version does not strand your users on a source build.
  • Should the full test suite run inside every wheel-building job?
    No. The wheel job asks whether this artefact loads and works on this platform; run a smoke test that imports each compiled module and exercises a representative path. The full suite belongs to a normal test job on the platforms where it runs fast. Putting a heavy suite inside every cell is what makes slow legs, especially emulated ones, time out.

saying these in an interview costs you the question

  • Wants a physical machine for every cell of the matrix
  • Forgets that each interpreter version is a separate ABI
  • Treats emulated builds as equivalent in cost to native ones
  • Runs the entire test suite inside every wheel job
  • Ignores the free-threaded build as just another interpreter
  • Publishes wheels per platform as each leg finishes

context