skip to content

How would you stage adding type hints to a large untyped Python codebase?

level: middleimportance: must knowfreq 55%

answer

  1. Types arrive in jumps, not smoothly
  2. Which functions do other modules import?
  3. Signatures before locals; return types first
  4. Name the data crossing the boundary
  5. Annotation-only commits, then a ratchet

basics

~20 s

Start at the public boundary: annotate the function and method signatures other modules import, then work inward. Type one module per change, keep annotation-only diffs separate from behaviour changes, and let the checker infer locals.

solid answer

~50 s

I treat it as a migration, not a formatting pass. First I annotate the **public boundary** of a module: the parameters and return types of the functions other modules import, plus the shapes of the data crossing that boundary (a `TypedDict` or a dataclass instead of a bare `dict`). Callers immediately get real types even though the bodies are still loose. Then I work inward, module by module, roughly in import order so a module I annotate already sees typed dependencies. Two habits keep it sane: annotation-only commits, never mixed with refactors, so review is cheap and a revert is safe; and a ratchet so a module that is typed stays typed. Since Python 3.14 annotations are evaluated lazily, so adding them to legacy modules costs nothing at import time and rarely needs quoted forward references.

code

python · 14 lines
python
from typing import TypedDict


class Record(TypedDict):
    accession: str
    year: int


def ingest(records: list[Record]) -> int:
    """The module's public boundary: annotated first."""
    return sum(1 for record in records if record["year"] > 1900)


print(ingest([{"accession": "A-1", "year": 1912}]))

go deeper

for a junior

Be ready to say what an annotation does and does not do: it records an intended type for tools and introspection, and it does not check or convert anything when the function runs. Knowing that is enough to see why adding types to old code is safe.

for a middle

Explain the mechanics: a function with no annotations is opaque to the checker and its body is usually skipped, so signatures at a module's public surface pay off first, return types most of all. Mention keeping annotation-only diffs separate.

for a senior

Show the migration judgment: ordering modules by the import graph and by where shape bugs actually occur, replacing bare dicts at the boundary with named shapes, and installing a ratchet so typed modules cannot silently regress.

for a principal

Own the strategy question: what the rollout is buying, which modules deserve the effort, and how progress is measured. Argue for 'new code is typed by default' over a coverage percentage, and be explicit about where you deliberately stop.

## The shape of the problem An untyped codebase does not become useful to a checker gradually in proportion to how many annotations you add. Value arrives in jumps, and where you put the first annotations decides how big those jumps are. The reason is how gradual typing works: an unannotated function is, to a checker, a black box. Its parameters and its return are unknown, and most checkers do not even analyse the body of a function that has no annotations at all. So a hundred annotations sprinkled on local variables inside private helpers buy almost nothing, while ten annotations on the signatures at a module's public surface make every caller in the codebase checkable. ## Boundary first, then inward The *public boundary* of a module is the set of names other modules import: the functions, methods and classes that appear in someone else's `from x import y`. Annotate those signatures first — parameters and, above all, return types. A missing return type is the single most expensive omission, because the value flows outward into caller code and takes the checking blackout with it. At the same time, name the data crossing that boundary. Legacy Python moves dictionaries around; a boundary annotated `dict` is barely better than none. Introduce a `TypedDict` for a record that really is a JSON-shaped mapping, or a dataclass where you can afford to construct objects. That one act converts a stringly-typed interface into something the checker can verify at every call site. Only then work inward: private helpers, then locals that inference cannot resolve. Most locals never need an annotation — inference handles them, and annotating them adds diff noise without adding checking. ## Ordering across modules Within a codebase, prefer to annotate in dependency order: modules with few internal imports first, then their consumers. If you annotate a consumer before its dependency, everything it receives is still unknown and you end up writing speculative annotations you later have to correct. Working from the leaves up means each module you touch already sees typed inputs, so the checker's complaints are real findings rather than noise about unknown types. There is a tension between "leaves first" and "the module with the most callers first". Resolve it by value: the module whose shapes are most often misunderstood — the one at the middle of the import graph that everyone passes records through — is worth doing early even if its dependencies are untyped, because that is where the shape bugs live. ## Keep the diff boring Annotation commits should change no runtime behaviour. Annotations are inert: they are stored for introspection and, since Python 3.14 (PEP 649/749), not even evaluated until something asks for them. That property is what makes the migration safe, and it is worth protecting. Mixing an annotation pass with a refactor destroys it — a reviewer can no longer tell whether the checker's silence means "correct" or "unchanged". So: one module, annotations only, and a separate commit for any behaviour fix the annotations exposed. It is common for the first typing pass over a legacy module to *find* a bug; write it down and fix it separately. ## Ratchet, do not sprint A typed module drifts back to untyped the moment somebody adds an unannotated function to it. The cheap defence is a list of modules that are considered done, enforced wherever your checker runs, plus a review habit: in a file that is already typed, a new function without annotations is a review comment, the same as a missing test. Progress is then monotonic, and the metric that matters is not "percent annotated" but "is new code typed by default". ## What it looks like in practice On a museum-catalogue importer, the first change is not the parser internals. It is `def ingest(records: list[Record]) -> IngestReport:` at the top of the package, with `Record` given a real shape. Every scheduler, every test and every downstream report immediately gets checked against that contract, and the untyped parsing internals underneath can be annotated over the following weeks without blocking anyone.

  • Why is a missing return annotation more costly than a missing parameter annotation?
    A parameter's unknown type stays inside the function. An unknown return value flows outward into every caller, and each expression derived from it becomes unchecked too, so one missing return type can blank out checking across a large slice of caller code. Annotating returns first is the highest-leverage move in a gradual migration.
  • How do you stop a module you have just typed from drifting back to untyped?
    Keep an explicit list of modules considered done and have the checker run against it wherever the team already runs checks, so a regression fails loudly rather than silently. Pair it with a review habit: in a file that is already annotated, a new function without annotations gets the same review comment a missing test would. The point is monotonic progress, not a coverage percentage.
  • The first typing pass over a legacy module surfaces a real bug. What do you do with it?
    Record it, and fix it in a separate commit from the annotations. An annotation-only change is provably behaviour-preserving and can be reviewed quickly or reverted safely; the moment it also changes logic, that property is gone and the reviewer cannot tell which half caused a regression. Ship the annotations, then ship the fix with a test.

saying these in an interview costs you the question

  • Annotate the entire codebase in one giant pull request
  • Start with local variables inside private helper functions
  • Add Any everywhere until the checker goes quiet
  • Claims annotations enforce types and slow the code down
  • Mixes annotation passes with refactors in one commit

context