skip to content

Fail fast is about detecting problems as early as possible. Rank the points at which a defect can be caught - compile/build, CI, service startup, first request, and never - and explain what you gain by moving detection earlier.

level: principalimportance: should knowfreq 35%

answer

  1. compile > CI > startup > first request > never
  2. earlier = smaller audience, more context
  3. lazy init defers failure to the worst moment
  4. startup checks are fleet-correlated
  5. reconciliation catches the 'never' rung

basics

~20 s

Best is catching it when the code is compiled or built, then in CI, then when the service starts, then on the first request, and worst is never noticing. The earlier the catch, the fewer people are affected and the cheaper the fix.

solid answer

~50 s

Treat fail fast as a ladder of detection points. Compile or build time is best: the defect never reaches a running system, feedback is seconds, and the blast radius is one developer - achieved with types that make invalid states unrepresentable, schema-generated clients, and static analysis. CI is next: contract tests, config-schema validation and migration checks catch what types cannot, still before deployment. Startup validation - required settings present, secrets resolvable, schema version compatible - fails a bad instance before it takes traffic, but is fleet-correlated, so it needs canaries and staged rollout. First-request detection means a real user meets the defect; it is acceptable only for genuinely per-request conditions. Never detecting is the true worst case: corrupt data spreads into backups and downstream systems. Each rung earlier reduces the number of people affected, the amount of context lost, and the reversibility cost - which is why an architect invests in pushing checks down the ladder rather than adding more runtime guards.

go deeper

for a junior

State the ordering - compile, CI, startup, first request, never - and that earlier catching means fewer people affected and an easier fix.

for a middle

Give a concrete technique per rung (types and non-nullability, contract and config tests, startup validation of required settings, guards at the boundary) and note that lazy loading pushes detection later.

for a senior

Quantify the gains per rung (audience, context, reversibility, latency), discuss eager versus lazy initialisation, generated contracts, and the correlated-failure risk of fleet-wide startup checks.

for a principal

Present it as an investment policy: pay for early detection where late failure is irreversible or wide-blast-radius (money, permissions, persisted data, cross-team contracts), accept later detection elsewhere, and back the bottom rung with reconciliation and impossible-state alerting.

## Fail fast as an axis, not a switch Fail fast is usually taught as "throw early in the function". At system scale the more useful framing is: **how early in the lifecycle of a change can this class of defect be detected?** Every defect has a point of first possible detection, and design decisions move that point. ### The ladder, best to worst **1. Compile / build time.** The defect cannot exist in a running system. Feedback is seconds, the audience is one developer, the fix cost is trivial, and no data is at risk. Achieved by: types that make invalid states unrepresentable (a non-empty-list type, a `PositiveAmount`, a sum type covering every case with exhaustive matching); non-nullable types; generating API clients from a schema so a contract change breaks the build; static analysis and dependency rules that reject illegal architectural references. **2. Build pipeline / CI.** Not expressible in the type system, but still pre-deployment: contract tests against a shared schema, config files validated against a schema, database migration dry-runs, linting for banned patterns, and tests that assert invariants. Feedback is minutes, audience is one team, nothing is deployed. **3. Service startup.** The instance validates what it needs before accepting traffic: required settings present and parseable, secrets resolvable, database schema version compatible, feature flag defaults sane. A bad instance never serves a request. This is genuinely valuable - the alternative is discovering a missing setting at 3 a.m. on the first request that happens to hit that code path. The caveat is **correlated failure**: a deterministic check runs identically on every instance, so a bad artifact or bad shared config fails the whole fleet at once and can even prevent the replacement from ever becoming healthy. Mitigations: canary and staged rollout, health/readiness signalling instead of immediate exit where the missing thing is non-essential, validating the same config in CI so startup is a backstop rather than the primary gate, and keeping the ability to roll back the *configuration* independently of the *code*. **4. First request / runtime.** Precondition guards and boundary validation live here. This rung is *unavoidable* for anything that depends on the request itself - and it is exactly where guards belong. It becomes a smell when it is the *first* opportunity to catch something that was knowable earlier: a lazily initialised client that only fails when first used, a config value read on demand deep inside a rare branch, a code path only exercised at month end. **5. Never.** Nothing throws; a wrong value is written. This is the failure the whole principle exists to prevent, and it is the most expensive because corruption spreads: into caches, backups, exports, analytics, and partner systems. Remediation may require reconstructing history and coordinating with external parties, and some of it may be irreversible. ## Why each rung is worth roughly an order of magnitude Moving one rung earlier improves four things simultaneously: - **Audience** - one developer, then one team, then one deploy, then all users. - **Context** - at compile time the author is looking at the code; a month later nobody remembers it. - **Reversibility** - a build failure costs a rebuild; corrupt persisted data may be unrecoverable. - **Feedback latency** - seconds, minutes, minutes, hours-to-months, never. ## Practical architectural moves - **Read all configuration once at startup into a validated, typed object** rather than fetching strings on demand. This converts a class of "first request" defects into "startup" defects. - **Eagerly initialise critical clients and connections at boot** so an unusable credential fails a healthy-instance check, not a user request. Lazy initialisation moves detection *later* - convenient, but it defers failure to the worst moment. - **Generate rather than hand-write cross-service contracts** so incompatibility is a compile error, not a production 500. - **Make illegal architectural dependencies fail the build** rather than relying on review. - **Add a startup self-check for things the code assumes but cannot type**: schema version, required capability of a dependency, presence of required secrets. - **Instrument the "never" rung**: reconciliation jobs, invariant checks over stored data, and alerting on impossible states. If something can only be detected after the fact, at least detect it deliberately rather than by customer report. ## The trade-off to state explicitly Pushing detection earlier costs design effort and can reduce flexibility: rich types are more work to introduce, generated contracts couple release processes, eager initialisation slows startup, and strict startup validation risks fleet-wide correlated failure. The judgement is to pay that cost for anything whose late failure would be *irreversible or wide-blast-radius* (money, permissions, persisted data, cross-team contracts) and to accept later detection for cheaply reversible, narrow-impact concerns.

  • You want strict startup validation but fear a fleet-wide crash on a bad config. How do you get both?
    Validate the same config schema in CI so startup is a backstop rather than the first gate; roll out via canary and staged deployment so a bad artifact fails a small percentage first; keep configuration rollback independent of code rollback; and split checks into fatal (cannot be correct without it) versus non-fatal (start, report not-ready or degraded).
  • Why is lazy initialisation often at odds with fail fast?
    Lazy initialisation defers the moment of truth to whenever the resource is first used, which may be a rare code path under production load. Eager initialisation at boot converts the same defect from a user-visible runtime failure into an instance that never becomes healthy - detected by the deploy, not by a customer.
  • What can you do about defects that genuinely cannot be detected before they occur?
    Detect them deliberately after the fact rather than by customer report: reconciliation jobs that re-check invariants over stored data, alerting on states the model says are impossible, and audit trails that make reconstruction possible. That converts the "never" rung into a bounded, monitored delay.

Catching a manufacturing defect on the drawing board, on the assembly line, at final inspection, at the customer's door, or in a recall two years later. The same defect, five wildly different costs - and only the last one involves lawyers.

saying these in an interview costs you the question

  • Treating fail fast as purely a runtime concern and never asking whether the type system or the build could have caught it.
  • Reading configuration on demand deep inside code paths, so a missing value surfaces on a rare production request instead of at startup.
  • Adding fleet-wide fatal startup checks with no canary, staged rollout, or independent config rollback.
  • Lazily initialising critical clients and calling it an optimisation, when it simply moves failure to the least convenient moment.
  • Assuming that if nothing has thrown, nothing is wrong - the 'never detected' rung needs reconciliation and invariant checks, not silence.

context