skip to content

How do you scope a typing rollout for a legacy Python importer, and decide when it has gone far enough?

level: principalimportance: should knowfreq 30%

answer

  1. Coverage is measurable and a poor goal
  2. Cost flat, value uneven across modules
  3. Evidence: incident history, import graph, change rate
  4. Name where the cost curve turns
  5. Types catch shapes, not resource bugs

basics

~20 s

Scope by where shape bugs actually occur and where interfaces cross team lines, not by coverage percentage. Stop when new code is typed by default and the remaining untyped code is stable, low-traffic and cheap to leave alone.

solid answer

~50 s

I scope it as an investment with a sharply diminishing curve. The first pass — public boundaries and the data shapes crossing them — is cheap and catches most of the real defects. The last stretch is expensive: generics, overloads, dynamic factories, untyped dependencies, code nobody understands. So I pick targets by evidence: modules that appear in incident write-ups as shape or `None` bugs, interfaces owned by one team and consumed by several, and anything under active change. Stable code that has not been edited in two years buys almost nothing from annotations. The success metric is not percent annotated; it is *new and changed code is typed by default*, with the untyped remainder shrinking as a by-product. I also state plainly what the rollout does not buy: types catch shape and contract errors, not resource or concurrency bugs, so I would not sell it as a cure for an unbounded memory growth incident.

go deeper

for a junior

You will not be asked to scope a rollout, but know the outcome it aims at: annotations that describe real contracts. Adding types to code you are already changing is the part of this you own day to day.

for a middle

Be able to argue which module you would annotate next and why — callers, change rate, and past shape bugs are better reasons than file size or alphabetical order.

for a senior

Show that you can defend a stopping point and describe the ratchet that keeps progress monotonic, including how escape-hatch counts and review habits are what preserve the value already bought.

for a principal

Own the economics and the pitch: uneven value against flat cost, where the curve turns for this codebase, what types do not buy, and why a coverage mandate produces Any at scale rather than safety.

## Scope is a value question, not a coverage question The failure mode of a typing rollout is treating it as a percentage to drive to 100. Coverage is measurable, which is why it becomes the goal, and it is a poor proxy for value. The cost of annotating code is roughly flat per function; the value is wildly uneven. A module with three callers and no data shapes returns nearly nothing. A module through which every record passes on its way from parsing to persistence returns a great deal. So the scoping question is: where do shape errors actually cost us? Three evidence sources answer it without guessing: - **Incident and bug history.** Search closed defects for the signatures of a type error: `AttributeError` on `None`, a `KeyError` on a record field, a string where a date object was expected. The modules that recur are the rollout's first targets. - **The import graph.** Interfaces owned by one team and imported by several are where a contract exists but is unwritten. Typing those makes the contract explicit and moves the failure from runtime to review. - **Change rate.** Annotations pay off when code changes, because that is when the checker is consulted. Code untouched for two years is not going to break in a new way; annotating it is archaeology, and archaeology should be optional. ## The cost curve, and where it turns The first tranche — boundary signatures and named record shapes — is mechanical. The second is judgment: generic containers, callables, and the places where the code is genuinely polymorphic in ways older annotations struggle to express. The third is a swamp: dynamically constructed classes, attribute injection, deep inheritance chains, and dependencies that ship no type information. On modern versions the middle tranche is cheaper than it used to be — PEP 695's type-parameter syntax landed in 3.12 and deferred annotation evaluation in 3.14 removes the import-cycle and forward-reference friction that used to make typing legacy modules unpleasant — but the third tranche is still where teams burn months for little return. A principal's job is to name where the curve turns for *this* codebase and to make stopping there a decision rather than a failure. "These eleven modules are typed and ratcheted; the legacy report generators are deliberately not, and here is why" is a healthy end state. ## What the rollout must not be sold as Static types catch shape and contract errors. They do not catch resource exhaustion, races, ordering bugs or performance regressions. It is worth being explicit about this when a rollout is proposed in the aftermath of an incident. If a museum-catalogue importer's six-hour nightly run died of unbounded memory growth because a cache was never evicted, no annotation would have prevented it; the fix is a bounded structure and a test that watches it grow. Overselling types after an incident buys short-term funding and long-term cynicism when the next outage has nothing to do with them. The honest pitch is narrower and more durable: types reduce the class of defect that reaches production from interface misunderstandings, they make refactoring safe enough to do at all, and they replace a lot of documentation and defensive `isinstance` code with something checkable. ## Organizational shape Two structures fail. A dedicated "typing team" annotating other people's modules produces annotations nobody trusts and a backlog that never converges; the owners must do their own modules. A blanket mandate with a deadline produces exactly what a mandate produces — `Any`, casts and suppressions in volume, which is measurable progress and negative value. What works is a ratchet plus ownership: new and changed code is typed, module by module a team marks its own code done, and regressions are caught wherever checks already run. Escape-hatch counts are watched for direction rather than absolute value, because a rollout that adds a suppression for every annotation has bought silence. ## Reviewing typed diffs One habit deserves naming because it is where the value is preserved: in review, a typed diff needs attention on the *annotations* as claims, not just on the logic. A parameter widened to `Any`, a cast with no justification, a return type loosened to make a caller compile — each is a small withdrawal from the account the rollout was funding. Reviewers who read annotations as promises keep the rollout worth its cost; reviewers who skip them let it decay into decoration within a year.

  • What metric would you report to leadership instead of percent-annotated?
    The share of new and changed code that ships typed, plus the trend in escape hatches — casts and suppressions — over the same period. Those two say whether the discipline is holding. I would pair them with a qualitative line: which interfaces now have explicit contracts, and which modules are deliberately out of scope with the reason recorded.
  • A team proposes a company-wide deadline for full annotation coverage. What is your objection?
    A deadline against a coverage number is satisfied most cheaply by `Any`, casts and suppressions, which is measurable progress with negative value. I would replace it with a ratchet — new and changed code is typed, teams mark their own modules done — and let coverage rise as a by-product. Deadlines also push annotation onto people who do not own the code, and annotations nobody trusts get ignored.
  • How do you keep a completed rollout from decaying?
    Ownership plus review habits. Typed modules stay on a list the checker enforces wherever checks already run, so a regression fails rather than drifts. In review, annotations are read as claims: a widened parameter, an unjustified cast or a loosened return type gets the same scrutiny as a swallowed exception. Without that, a typed codebase becomes decorative within about a year.

saying these in an interview costs you the question

  • Sets a percent-annotated target as the goal
  • Claims static types would have prevented a memory or race incident
  • Staffs a dedicated typing team to annotate other teams' code
  • Mandates full coverage with a deadline and no ratchet
  • Never plans to stop; treats untyped legacy code as always worth annotating

context