How do you stop test-only wiring from quietly becoming a second application configuration that no longer matches production?
answer
- each replacement removes evidence
- budget, do not ban
- keep one unmodified boot
- startup faults need the real graph
- a long list is a design signal
basics
~20 sTreat each replacement as a piece of the real graph the suite stops exercising. Budget them, review additions like production configuration, and keep at least one tier that boots the unmodified wiring so startup faults still have somewhere to surface.
solid answer
~50 sEvery replacement buys speed or determinism by removing a part of the real arrangement from the evidence, and the cost is invisible because the suite goes greener as it grows. I manage it as a budget rather than a ban: a few named wirings owned like production configuration, additions reviewed with the question "what does this stop us proving?", and a visible count so growth is a decision rather than drift. Crucially, at least one tier boots the production wiring unmodified — against disposable instances where the dependency allows — because startup faults such as an unsatisfied dependency, an ambiguous candidate, a misread setting or a component whose construction fails only exist in a graph nobody replaced. I also read a growing override list as a design signal: when a test cannot boot without replacing five collaborators, the coupling is the defect, and fixing the application beats fixing the test configuration.
go deeper
Notice that every collaborator your test replaces is one the run no longer checks. Prefer replacing the awkward dependency only, not everything around it.
Be able to list what an amended graph can no longer fail for: unsatisfied dependencies, ambiguous candidates, construction-time work and the settings a deployment supplies.
Keep an unmodified boot somewhere in the pipeline and be able to justify which paths get it. Spot the replacements that exist because of a design defect and fix the application instead.
Run it as an explicit budget with owners, a visible count and a stated cost per wiring, and defend where you sit between fidelity and suite duration for this system rather than in general.
## The failure this prevents Replacements accumulate one reasonable decision at a time. A dependency is slow, so it gets a stand-in. A component is awkward at startup, so it is bound to a no-op. A setting is inconvenient, so the test mode overrides it. Each step is defensible, and after two years the application the suite boots shares a name with production and not much else. Nothing fails; the suite is green and fast, which is exactly why nobody revisits it. The cost is specific and worth naming precisely: **the evidence removed is everything that only exists when the real arrangement is assembled**. Namely - whether every dependency in the graph can actually be satisfied; - whether two candidates for one boundary would be ambiguous; - whether construction-time work succeeds against a real target — a pool opening, a value read and cached, a listener subscribing; - whether the settings the deployment supplies are the ones the components read; - whether the real implementation of a boundary honours the contract the stand-in was written to imitate. A suite that replaces liberally cannot fail for any of those reasons, so it stops reporting them. ## Running it as a budget | lever | what it does | what it costs | |---|---|---| | a small set of named wirings, owned | keeps the arrangements countable and reviewable | some tests carry replacements they do not need | | review each addition as configuration | makes the lost evidence explicit at the moment it is lost | slows down adding a test that needs new wiring | | a tier that boots unmodified | restores the startup-fault evidence | the slowest tier, and needs disposable dependencies | | a visible count of replacements | turns drift into a number someone owns | only works if someone actually looks | The judgment call is where to sit on that table, and it differs by system: a service with two out-of-process dependencies can afford to boot unmodified often, while one with a dozen cannot, and should concentrate its unmodified boots on a handful of paths that carry the most risk. ## The design signal The most useful reframing to bring to this discussion is that **a long replacement list is usually a statement about the application, not about the tests**. When a graph cannot start without a stand-in for five collaborators, the cause is generally one of: 1. components doing remote work during construction rather than on first use; 2. a boundary defined in the vendor's shape rather than the shape the application needs, so nothing cheap can stand in for it; 3. a module that starts everything or nothing, with no seam for a narrower slice; 4. configuration that is read at construction time from a source only a deployment provides. Each of those is fixable in the application, and each fix removes replacements from every test at once. A team that only ever fixes these in test configuration is paying rent forever on a defect it could buy out. ## What the policy looks like in practice - **Name and own the wirings.** Not "whatever this test class happens to declare" but a short list with a stated purpose each, reviewed like production configuration. - **Keep an unmodified boot in the pipeline.** Even a handful of paths, run per merge, catches unsatisfied and ambiguous dependencies before a deployment does. - **Make the replacement count visible.** A number in review beats an annual audit that never happens. - **Re-examine long-lived replacements periodically.** A stand-in added because a dependency was slow three years ago may now be replacing something cheap. - **Say what each wiring stops proving.** One sentence per wiring is enough, and it is the artefact that makes the tradeoff reviewable by someone who did not write it. ## What interviewers listen for This is a judgment question with no single right answer, and the failure modes are the two extremes. "Replace everything, tests must be fast" gives up the evidence the suite exists to produce. "Never replace anything" produces a suite nobody runs, which produces no evidence at all. What distinguishes a strong answer is naming the specific evidence each replacement removes, holding at least one place in the pipeline where the real arrangement is assembled, and treating a growing list as a prompt to change the application rather than the test configuration.
- Which failures can only be caught by booting the wiring the deployment actually uses?Unsatisfied dependencies, ambiguous candidates for a boundary, construction-time work that fails against a real target, and settings the deployment supplies being read differently than expected. All of them live in the assembly step, so a graph where that step was amended cannot report them.
- How do you decide which tests keep replacements and which move to an unmodified boot?By risk and cost. Paths where an assembly or settings fault would be expensive, and dependencies cheap to stand up disposably, earn an unmodified boot. Paths dominated by branching logic over a slow or unavailable dependency keep the replacement. State the split explicitly rather than letting it settle by accident.
- What would make you change the application rather than adding another replacement?A replacement that exists only because a component does remote work during construction, because a boundary is shaped around a vendor type, or because a module starts all-or-nothing. Those are design defects charging rent on every test; fixing them removes replacements across the whole suite at once.
saying these in an interview costs you the question
- Argues every out-of-process dependency should always be replaced for speed
- Never boots the unmodified wiring anywhere in the pipeline
- Treats a growing replacement list as inevitable rather than a design signal
- Assumes a green suite proves the real graph can be assembled
- Lets each team add test-only wiring with no review or owner