skip to content

When should a Python library ship types inline rather than as a stubs distribution?

level: principalimportance: should knowfreq 18%

answer

  1. Two places types can live
  2. Only one of them is verified
  3. Ownership and release cadence decide
  4. Independent versioning cuts both ways
  5. A marker is hard to withdraw

basics

~20 s

Ship inline with a py.typed marker whenever you own source that can carry annotations: one artefact, no drift, and the types are checked against the implementation. Reserve a -stubs distribution for code you do not own or compiled modules.

solid answer

~50 s

Inline annotations plus `py.typed` should be the default for code you control: there is one source of truth, the annotations are validated against the implementation on every check, and consumers get types by installing the package they already depend on. A stub-only `-stubs` distribution buys three things inline cannot — you can type a package you do not own, you can use typing syntax newer than the runtime the package supports, and you can version and roll back the types independently of the code. It costs a second artefact that nothing keeps honest. The decision therefore turns on ownership and release coupling, not on taste: if a bad annotation must be revocable without re-releasing the runtime, or the types come from a different team than the code, stubs earn their keep; otherwise the drift risk dominates and inline wins.

go deeper

for a junior

Take away the default: if you own the source, annotate it in place and ship the marker. Separate stub packages are for code you cannot edit, not a general style choice.

for a middle

Be able to list what stubs buy — typing code you do not own, compiled modules, newer syntax than the runtime supports — and the price, which is a second artefact nothing validates against the implementation.

for a senior

Argue the release-coupling side: how stub versions are pinned to library versions, how drift is caught in CI, and why a wrong stub is more dangerous than no types at all.

for a principal

Own the estate-wide policy — when a library graduates to advertising types, who maintains external stubs, how a type-surface change is sequenced across consumers, and how you avoid the split where some are pinned to old stubs and some to new.

### Framing the decision There are two ways for a Python distribution to hand types to its consumers. **Inline**: annotate the source and ship a `py.typed` marker inside the package. **Stub-only**: publish a separate installable whose importable directory is named `<package>-stubs` and contains nothing but `.pyi` files. Both are PEP 561 mechanisms, and a checker prefers the `-stubs` distribution over the inline types when both exist. Choosing between them is a release-engineering question wearing a typing costume. ### The case for inline Inline annotations have one decisive property: they are checked against the code they annotate. Write `-> list[str]` on a function that returns `None` and the checker says so on the next run. A stub cannot be wrong in a way anything notices. Everything else follows from that — one artefact to install, one version to reason about, no possibility of a consumer having types from one release and behaviour from another, and no second review surface where a change can be forgotten. The historical arguments against inline have mostly expired. Annotations used to cost import time and could raise `NameError` on forward references, which pushed some projects to `from __future__ import annotations`. On Python 3.14, PEP 649/749 makes annotations lazily evaluated by default: they are computed only when something asks for them, through `annotationlib`, so an annotation naming a class defined later no longer needs quoting and the import-time cost is largely gone. That removes most of the remaining reason to keep types out of the source. The real cost of inline is that it makes annotations part of your public contract at the same moment as your code. Ship the marker and every consumer type-checks against you; tightening a parameter type later is a breaking change for people whose code still runs perfectly. That is an argument for annotating carefully before adding the marker, not for hiding the types in a second package. ### The case for stub-only Four situations genuinely call for stubs. First, **you do not own the code**: a dependency that ships no types can be typed by anyone from the outside, which is how much of the ecosystem got typed at all. Second, **there is no source to annotate**: a compiled extension module has nowhere to put an inline annotation. Third, **syntax skew**: a stub is parsed only by the checker, so it can use unions written with `|` or PEP 695 type parameters while the package itself still imports on an older Python; inline annotations do not get that freedom unless they are strings. Fourth, **independent release cadence**: types that ship separately can be corrected — or reverted — without cutting a release of the runtime code. That fourth point is the one worth thinking hardest about, because it cuts both ways. Consider a flight-schedule differ published to a dozen internal consumers. A tightened annotation goes out, several teams rebuild, half of them break, and the fix must be withdrawn. If the types were inline, the withdrawal is a runtime release that everyone must adopt; if they were in a `-stubs` distribution, you yank one artefact and the code keeps running untouched. But the same decoupling is what makes a *partial-failure rollback* possible: after the withdrawal some consumers are pinned to the old stubs and some to the new, the runtime is identical for all of them, and the two groups now disagree about what the API is while every build stays green. Independent versioning is a genuine capability and a genuine hazard, and if you take it you owe your consumers a compatibility statement about which stub versions go with which library versions. ### Containing drift when you do choose stubs Stubs need the discipline inline gets for free. Pin the stub distribution's version range to the library's. Run an import-and-diff consistency check in CI that compares the stub's declared surface against the imported module's real attributes. Exercise the typed surface with real calls — a regression pack of, say, 340 recorded differ cases routed through the annotated API is both a behaviour test and a signature test, because those calls are checked and executed. And name an owner: a stub repository with no owner is a stale stub repository within two releases. ### A workable policy Inline by default for anything you own that has Python source. Stubs for compiled modules, shipped in the same distribution so they cannot skew. Stub-only distributions for dependencies you do not control, and for the rare case where the types must be revocable on their own. Add the marker only when the public surface is annotated and checked in CI, because withdrawing it later turns every downstream symbol back into `Any`. Whatever you choose, say it once in the project's contributing guide, since the failure mode is not picking wrong — it is having both, half-maintained, and no statement about which one is authoritative.

  • You inherit a library with both inline annotations and an external stubs distribution. What do you do?
    Establish which one consumers actually get — a `-stubs` distribution outranks the inline types — then pick one as authoritative and delete the other. Usually the stubs go: the inline annotations are checked against the implementation and the stubs are not, so keeping both means shipping a contract nobody validates. Announce the change, since consumers pinned to the stub distribution will see their types shift the moment it stops being installed.
  • How does tightening an annotation become a breaking change even though nothing at runtime changes?
    Because downstream builds check against your annotations. Narrowing a parameter from `object` to `str`, or a return from `Any` to a concrete type, turns previously valid calling code into build failures, in a release whose runtime behaviour is identical. Treat type-surface changes with the same care as signature changes: land them on a version boundary, note them in the changelog, and prefer widening what you accept over narrowing it.
  • When is it right to leave a library untyped rather than ship types you are unsure about?
    When the annotations would be guesses. An untyped dependency gives consumers `Any` and no false confidence, and they can write their own local stubs; a wrong advertised type produces green builds around code that fails. If coverage is genuinely partial, a stub distribution whose `py.typed` contains `partial` is the honest middle, since checkers then keep consulting the runtime package for what you have not described.

Inline annotations are a promise written into the contract you are already signing; a stubs distribution is a side letter — easier to amend, and easier for the two documents to end up disagreeing.

saying these in an interview costs you the question

  • Treats stubs as the modern default for owned code
  • Ignores that stubs are never checked against the implementation
  • Thinks annotations are free to tighten after release
  • Ships both inline types and stubs with no authority stated
  • Assumes independent stub versioning has no downside

context