Why do C extensions block PyPy adoption, and which bindings port cleanly?
answer
- The API leaks one runtime's object model
- No reference counts, and the heap moves
- Every crossing builds and syncs a proxy
- Plain C-ABI bindings avoid the problem
- Audit transitive compiled dependencies first
basics
~20 sCPython's C API hands extensions raw pointers to objects and makes them maintain reference counts. PyPy has a moving collector and no reference counts, so it emulates that API — costly at every crossing and not always complete.
solid answer
~50 sA compiled extension built for CPython is written against an API that exposes CPython's own object model: pointers into a heap that never moves, manual reference counting, borrowed references, and direct struct access. PyPy cannot honour that natively — it has no reference counts and a generational, moving collector — so it runs such extensions through an emulation layer that builds and syncs a CPython-shaped shadow object at each boundary crossing. That creates two problems: **cost**, since a fine-grained extension called once per field can end up slower on PyPy than on CPython, and **completeness**, since extensions that poke at internals may not build at all — and binary wheels are tagged for CPython, so you are often compiling from source. Bindings that go through the plain C ABI instead — `ctypes` in the standard library, or a foreign-function library designed to be implementation-neutral — port cleanly, as does pure Python.
code
python · 4 linesimport ctypes
buf = ctypes.create_string_buffer(b"payload", 32)
print(ctypes.sizeof(buf), buf.raw[:7])go deeper
Be ready to say that some Python packages ship compiled C code rather than pure Python, and that such packages are built specifically for CPython, which is why they are the hard part of running anything else.
Explain the mechanism: CPython's C API exposes reference counts and non-moving object pointers, PyPy has neither, so it emulates the API through proxies. Know that ctypes-style bindings sidestep this because they only cross the plain C boundary.
Demonstrate the audit. Enumerate transitive compiled dependencies, check for builds or a source-build path, run their test suites on the target runtime, and benchmark the boundary-heavy ones with recorded payloads — then explain why an extension-heavy service can end up slower.
Own the strategic side: whether the organization keeps its compiled dependencies behind a narrow interface so the runtime stays a reversible choice, and what a hard dependency with no cross-implementation story costs you in optionality.
**What a compiled extension actually is.** A C extension is a shared library that the import system loads like a module. It is compiled against CPython's C API, and that API is not an abstract interface — it is a window onto CPython's internals. Extensions receive raw pointers to objects living in a heap that never relocates them, adjust reference counts by hand with macros, deal in *borrowed* references whose lifetime is someone else's problem, and in places read struct fields directly for speed. Every one of those is a promise about how CPython represents objects. **Why PyPy cannot just implement it.** PyPy's object model is deliberately different, and that difference is the whole source of its speed. There are no per-object reference counts; memory is reclaimed by a generational collector that *moves* objects to compact the heap. Handing a moving object's address to C code and asking that code to keep it alive by incrementing a counter is not something PyPy can do directly. So PyPy provides an emulation layer: when an object crosses into C, the runtime materializes a CPython-shaped proxy for it, pins it so it will not move, tracks the counter the extension manipulates, and synchronizes state back when control returns. That works, for well-behaved extensions. What it is not is free. On CPython a call into an extension is close to a function call. On PyPy each crossing carries proxy construction, pinning and state synchronization. The consequence is counter-intuitive and worth stating plainly in an interview: **a workload dominated by compiled extensions can be slower on PyPy than on CPython**, not faster — you pay boundary costs, and the JIT gets nothing to optimize because the hot code was already machine code. A coarse-grained extension called once with a large batch of work amortizes the crossing; a fine-grained one called once per field multiplies it. **Completeness is the second problem.** Extensions that use private APIs, reach into concrete struct layouts, or depend on subinterpreter- or refcount-specific behaviour may fail to build or misbehave. Distribution compounds it: binary wheels are tagged for a specific CPython ABI, so on another implementation you frequently fall back to building from source, which means shipping a compiler and headers in your image and hoping the project's build works there. Many projects do not test on any implementation but CPython, so you are the one who finds out. **What ports cleanly.** - **Pure Python.** Nothing to port, and it is the code a tracing JIT can actually speed up. - **`ctypes`.** In the standard library, it calls into a shared library through the platform's ordinary C calling convention. The foreign function never touches Python object internals — the runtime marshals arguments at the boundary — so any implementation that can make a foreign call can support it. It is slower per call than a hand-written extension on CPython, but it is portable across implementations. - **Foreign-function libraries designed to be implementation-neutral.** The same principle as `ctypes`, with a nicer declaration story: you describe the C signatures and the library handles the boundary, with no CPython object model in sight. - **Handle-based, implementation-agnostic extension APIs.** Newer designs replace raw object pointers with opaque handles that the runtime can move behind your back, which is precisely what a moving collector needs. An extension written against such an API can be built for several implementations from one source. Note that CPython's own stable ABI is a different axis. It lets one binary work across CPython minor releases, which is valuable, but it is still CPython's object model — it does not by itself make an extension work on a different implementation. **How to run the audit before a trial.** Take a real example: a webhook receiver that verifies signatures with a compiled crypto library and then parses payloads in pure Python. The pure-Python half is a genuine candidate for a tracing JIT; the compiled half is not. So enumerate every dependency that ships a platform-specific binary artifact — transitively, because the blocker is usually three levels down and nobody remembers it is there — and for each ask: does a build exist for the target implementation, or must it compile from source in my image? Does the project run its tests there? Is it called coarsely or once per field? Then run the project's own test suite under the target runtime, and benchmark the boundary-heavy dependencies with **real recorded payloads** rather than synthetic ones. Real traffic is what exposes the paths synthetic input never takes — the malformed body, the unexpected content type, the encoding mismatch that sends a parser down a slower fallback branch on one runtime and not the other. The conclusion candidates should reach: a service whose hot path is pure Python and whose dependencies are pure Python is a real candidate for an alternative runtime. A service whose hot path lives inside compiled libraries gains nothing from the JIT and pays at every crossing — and that, far more often than raw speed, is what decides the question.
- How would you audit a service's dependencies for compatibility before trialling another runtime?List every dependency that ships a platform-specific binary artifact, transitively — the blocker is usually a package nobody remembers depending on. For each, check whether a build exists for the target implementation or whether it must compile from source in your image, and whether the project tests there at all. Then run those projects' own test suites under the target runtime and benchmark the boundary-heavy ones with recorded production payloads, because per-call crossings are where the emulation cost lands.
- Why can an extension-heavy workload run slower on PyPy than on CPython?Two effects compound. The tracing JIT cannot optimize anything inside a compiled library, so the part you hoped to speed up is exactly the part it cannot touch. And each crossing into the emulated C API costs proxy construction, pinning and state synchronization, where on CPython it is close to a plain function call. A fine-grained extension called once per field multiplies that cost until it dominates.
- Why does a ctypes binding port more easily than a hand-written C-API extension?`ctypes` calls a shared library through the platform's ordinary C calling convention, so the foreign function never sees a Python object, a reference count or a struct layout — the runtime marshals arguments at the boundary. Any implementation that can make a foreign call can implement it. The trade is speed: per call it is slower than a hand-written extension on CPython, so it suits coarse-grained calls rather than tight inner loops.
saying these in an interview costs you the question
- Assumes any binary wheel installs on any implementation
- Thinks the JIT accelerates code inside compiled libraries
- Believes reference counting is part of the language specification
- Ignores transitive compiled dependencies in the audit
- Benchmarks with synthetic input instead of recorded payloads
- Confuses CPython's stable ABI with cross-implementation portability