As a lead, how do you decide whether adding a static type checker to an untyped codebase pays off?
answer
- Classify your own escaped defects first
- No industry constant to quote here
- Lifetime, churn, fan-out, turnover
- Coverage of paths, not of files
- Bounded pilot with a stopping rule
basics
~20 sDecide from your own defect record and code profile, not from taste: classify past escaped defects by whether types would have caught them, weigh annotation cost against code lifetime, churn and interface fan-out, and define which regions must be covered for the benefit to be real.
solid answer
~50 sStart with evidence you already own: sample the escaped-defect record and classify how many were type-shaped. There is no reliable industry constant for that share, so measure your own corpus rather than quoting one. Then weigh value drivers — long-lived code, high churn, wide interface fan-out, frequent refactors, team turnover, few other safety nets — against costs: annotation effort concentrated in the worst-understood code, review friction, build time, tooling upkeep, and the false-confidence risk of a half-covered codebase. Decide the target state before you start: which regions must be fully covered for the guarantee to mean anything, usually the shared data model and the boundaries into the core. Pilot on the highest-churn module for a bounded window, measure what it caught and what it cost, and be willing to stop. A permanent partial coverage that protects the money paths is a legitimate end state, not a failed migration.
go deeper
You are unlikely to own this call, but know that adopting a checker is a cost as well as a benefit, and that the cost lands hardest in the oldest, least-understood code.
Be ready to describe the mechanics you would use: pick a high-churn module, cover the shared data model first, and raise strictness one setting at a time rather than all at once.
Show that you would gather evidence before advocating: classify past escaped defects and incidents, run a bounded pilot, and report real person-time cost alongside what it caught.
Own the framing as a scoped investment: define the target state in covered paths, set a stopping rule, decide which regions stay untyped forever, and refuse vanity metrics when reporting upward.
## Frame it as an investment with a measurable return The question is not whether types are good. It is whether *this* codebase, with *this* team, at *this* moment, gets more back than the migration costs. Two teams can rationally reach opposite answers, and a lead who cannot articulate their own inputs is arguing aesthetics. ## Start with your own defect record The cheapest evidence is already sitting in the issue tracker. Pull the escaped defects from the last few quarters and classify each: would a checker have caught this before merge, could it have caught it if the domain quantities were modelled as distinct types, or is it purely behavioural? Do not quote an industry percentage for the share of defects types catch — the published studies are small, context-bound, and disagree; treating any single figure as a constant is exactly the overreach an interviewer is listening for. Your own corpus is both more defensible and more persuasive to the people funding the work. Run the same classification on incidents, not just bugs, because severity is what pays for the migration. Twenty cosmetic type-shaped defects matter less than one that took the payroll run down. ## The value drivers - **Lifetime.** Annotations amortise over years. Code with a known decommission date next quarter should not be annotated. - **Churn.** The benefit is largest where code changes often, because each change re-tests the interfaces mechanically. - **Interface fan-out.** A module with many callers converts a signature change into a list of exact edit sites. A leaf with one caller gains little. - **Refactor appetite.** If the team is avoiding a restructuring because nobody can find all the call sites, that is the strongest argument available. - **Team turnover and size.** Annotations are the cheapest form of always-true documentation for a newcomer, and their value rises as the people who remember the code leave. - **Absence of other nets.** Thin test coverage on a wide surface raises the return; a fast, thorough suite on a narrow surface lowers it. ## The costs, stated honestly Annotation effort is not evenly spread — it concentrates exactly in the oldest, least-understood, most dynamically written code, which is also where it takes longest and where you are most likely to encode a wrong assumption. Add review friction while the team argues about how to express things, longer feedback cycles, ongoing config and dependency-declaration upkeep, and skill ramp. The subtlest cost is **false confidence**: a half-covered codebase where people believe the checker is protecting a path it never analysed. ## Define the target state before the first annotation Decide what "done enough" means, in terms of coverage of paths rather than of files. For a payroll engine that usually reads: the shared money and calendar model, every function that other modules can enter, and everything on the posting and rollback paths. Coverage that stops short of the boundaries buys cost with no guarantee, which is the most common way these migrations waste a year. ## Pilot, measure, and keep the exit Run a bounded pilot on one high-churn module — a few weeks, not a quarter — and record concrete numbers: errors surfaced during annotation, feedback-cycle change, review-time change, and how many of the recorded past defects in that module would have been caught. A 4-person team maintaining a payroll engine of 1,163 files did exactly this on the 38-file money core: eleven weeks of part-time work surfaced 19 latent defects, 6 of them on a partial-failure rollback path that no test exercised, at a cost of roughly a fifth of one person's time. That result justified continuing into the boundaries and explicitly *not* continuing into the reporting and formatting code, which stayed unannotated permanently. Be willing to conclude the opposite. If the pilot surfaces two trivial defects and burns a third of the team's capacity, stop and say so. A lead who cannot end an initiative they started will not be trusted to start the next one. ## Choose the strictness ladder deliberately Adopting at the strictest setting on day one produces a wall of findings that the team will answer with escape hatches, and the escape hatches will outlive the enthusiasm. Pick a level the current code can nearly satisfy, then raise one dial at a time with a specific owner and a date. Watch escape-hatch density as the health metric — rising density means the ladder moved faster than the codebase. ## The anti-goals Do not annotate dead or deprecated code. Do not annotate generated code by hand. Do not make annotated-file percentage the reported KPI, because it rewards the easy files and hides the boundaries. And do not present this as a moral argument about typed versus untyped languages; present it as a scoped investment with a measured return, an explicit target state, and a stopping rule.
- What would make you recommend against adopting a checker at all?Code with a near decommission date, low churn, or a single caller per module — the annotations never amortise. Also a team already saturated by delivery pressure, since a half-finished migration leaves the worst outcome: cost paid, boundaries uncovered, and false confidence in place. And a defect record where almost nothing was type-shaped: if the escaped defects were domain-rule and ordering errors, the same effort spent on domain modelling or boundary validation returns more.
- How do you keep a partially annotated codebase from creating false confidence?Make the covered region explicit and visible: name which modules and which call paths are fully covered, and mark every function that can be entered from uncovered code as a validating boundary. Report coverage as covered critical paths, never as annotated-file percentage. Then treat escape-hatch density inside the covered region as a health metric, because a covered region full of unchecked assertions is uncovered in everything but name.
- What would you measure during a bounded pilot to justify continuing?Latent defects surfaced during annotation and where they sat — a defect found on an untested rollback path is worth more than one in a well-covered helper. The share of that module's recorded past defects the checker would have caught. Cost in real person-time, not story points. Change in feedback-cycle and review time. And a qualitative read from the people who did the work on whether the second module will be cheaper than the first.
saying these in an interview costs you the question
- Quotes an industry percentage of defects types catch
- Argues from language preference rather than defect data
- Reports progress as annotated-file percentage
- Starts at the strictest setting across the whole codebase
- Has no stopping rule or defined target state
- Plans to annotate deprecated or generated code