skip to content

When should a system deliberately NOT fail fast, and how do you decide between failing fast and degrading gracefully?

level: seniorimportance: must knowfreq 60%

answer

  1. can I still be correct without it?
  2. degrade features, fail requests
  3. degradation must be visible, never silent
  4. identical check on every node = fleet outage
  5. timeouts + circuit breaker = fail fast at the edge

basics

~20 s

Fail fast when continuing would give wrong answers or corrupt data. Degrade instead when the failure only affects a non-essential part and the core result is still correct - for example hiding a recommendations panel rather than failing the whole page.

solid answer

~50 s

The decision hinges on whether continuing can produce an incorrect result. If an invariant is broken or the data needed for a correct answer is missing, fail fast: a visible error is cheaper than silently wrong output or corrupt persisted state. If the failure is confined to an optional capability and the primary answer is still correct, degrade - omit the panel, serve a stale cached value clearly marked as such, or fall back to a simpler algorithm. Two extra dimensions matter. First, correlated failure: a fail-fast check that every instance evaluates identically, such as a bad config value crashing on boot, can take a whole fleet down at once, so such checks need canaries and staged rollout. Second, at service boundaries fail fast changes meaning: rejecting work immediately when a dependency is known unhealthy (open circuit breaker, full bounded queue, exceeded deadline) is how you avoid queue buildup and cascading failure. Fail fast on the *request*, degrade on the *feature*, and never trade correctness for uptime silently.

go deeper

for a junior

Give the core test: if the result would be wrong, stop; if only an extra feature is missing, show the rest without it. One concrete example is enough.

for a middle

Add the requirement that degradation be visible (stale markers, metrics, alerts) and mention timeouts and fallbacks as the practical mechanisms.

for a senior

Discuss cascading failure, circuit breakers, bounded queues, load shedding, deadline propagation, retry budgets, and the fail-closed-versus-fail-open choice for authorisation.

for a principal

Frame it as a per-capability correctness-versus-availability policy tied to business impact, address correlated failure from identical fail-fast checks across a fleet, rollout and rollback strategy, crash-loop amplification, and how degraded modes are made observable and time-bounded rather than permanent.

## Two questions that get confused "Should I fail fast?" is really two questions: 1. **In-process correctness:** an invariant is broken or an argument is invalid. Here fail fast is nearly always right - continuing yields wrong results. 2. **System behaviour under an environmental fault:** a dependency is slow, a cache is cold, a third-party API is down. Here the answer is a design decision, not a principle. Conflating them produces both bad extremes: services that crash on any hiccup, and services that swallow everything and serve nonsense. ## The deciding test **Can the system still produce a correct primary result without the missing piece?** - **No** -> fail fast. A payment service that cannot reach the ledger must not "assume success". A pricing service without the current rate must not guess. - **Yes** -> degrade. A product page whose recommendation service is down should render the product. A dashboard whose analytics panel times out should render the rest with an explicit "unavailable" marker. The crucial refinement: degrading must be **visible**. A stale value served silently is indistinguishable from a fresh one and re-introduces silent corruption at the system level. Mark it stale, expose it in metrics, alert if the degraded mode persists. ## Vocabulary - **Graceful degradation** - continuing with reduced functionality when a component is unavailable. - **Fail-safe / fail-secure** - on failure, move to the state that is safest. Note these can point in opposite directions: a fail-safe door unlocks in a fire (life safety), a fail-secure door locks (asset protection). For authorisation, failing *closed* (deny) is the fail-fast-flavoured choice; failing *open* (allow) trades correctness for availability and is usually unacceptable. - **Load shedding** - deliberately rejecting a fraction of work immediately when overloaded, so accepted work still meets its deadline. - **Backpressure** - signalling upstream to slow down rather than accepting unbounded work. - **Circuit breaker** - after repeated failures against a dependency, stop calling it and fail immediately for a cool-down period. - **Bulkhead** - isolating resources per dependency so one slow dependency cannot consume all threads or connections. ## Fail fast at service boundaries means something specific In distributed systems, "fail fast" usually means *reject immediately rather than wait or queue*. The failure mode it prevents is **cascading failure**: a slow dependency causes callers to hold threads and connections; queues grow; latency crosses the caller's own deadline so the work is wasted anyway; the caller's callers then saturate. Rejecting early keeps the system alive. Concrete mechanisms: aggressive **timeouts** and **deadline propagation** (never an unbounded wait), **bounded queues** that reject when full instead of growing, **circuit breakers**, **concurrency limits**, and **retry budgets with jitter and exponential backoff** so retries do not become a self-inflicted denial of service. ## Where fail fast actively hurts 1. **Correlated crash on boot.** A validation that every instance evaluates identically - a malformed config value, an unreachable dependency at startup - crashes the entire fleet simultaneously and can prevent rollback if the new instances never become healthy. Mitigation: canary and staged rollout, validate config in CI, and distinguish "cannot start correctly" (crash) from "a non-essential dependency is unavailable" (start, report unready or degraded). 2. **Crash loops as an amplifier.** Restarting on failure plus a dependency that is down equals a restart storm hammering the dependency. Mitigation: backoff on restart, and readiness signalling rather than exiting. 3. **User-facing input.** Aborting on the first bad field is technically fail fast but poor UX; collect and report all errors at once. Nothing invalid is constructed either way. 4. **Best-effort data paths.** In log ingestion, metrics or a media pipeline, dropping or skipping a malformed record and counting it is usually right - but the counter must be monitored, otherwise it is silent loss. ## Putting it together A well-designed service typically does all of the following at once: fails fast on invalid arguments and broken invariants inside the process; validates and rejects malformed input at its boundary; fails fast on *requests* to unhealthy dependencies via timeouts and circuit breakers; and degrades *features* that are not essential to the primary result, loudly and observably. The unacceptable combination is silent tolerance of incorrectness - and the second-worst is crashing everything because one optional feature is unavailable.

  • A service starts up but its non-essential recommendation dependency is unreachable. Should it exit?
    No. Start, mark itself ready for its core function, and report the degraded capability through health and metrics. Exiting turns a partial outage into a total one - and if the dependency is down for everyone, every instance crashes simultaneously. Reserve boot-time crashes for things without which the service cannot be correct at all.
  • How do timeouts relate to fail fast?
    An unbounded wait is the ultimate fail-slow: the thread, connection and the caller's deadline are all consumed for work that will be discarded anyway. A timeout converts an indefinite hang into a prompt, explicit failure that the caller can act on, which is what makes circuit breakers, retry budgets and load shedding possible.
  • Is failing open ever the right choice for an authorisation check?
    Very rarely, and only as an explicit, documented, time-bounded business decision with monitoring - because it trades a correctness guarantee for uptime. The default should be fail closed: if you cannot verify permission, deny. Silently granting access on infrastructure failure is a security incident waiting to happen.

An aircraft: a failed entertainment system does not ground the flight (degrade the feature), but a failed altimeter reading is not guessed at - the crew switches to a verified source or diverts (fail fast on correctness). What is never acceptable is displaying a made-up altitude that looks real.

saying these in an interview costs you the question

  • Treating fail fast as an absolute rule and crashing the whole request because one optional widget's data is missing.
  • Serving stale or fallback data silently, with no marker and no metric, so degradation is indistinguishable from healthy operation.
  • Adding a strict startup check that every instance evaluates identically without a canary or staged rollout, creating a fleet-wide correlated outage.
  • Retrying a slow dependency without timeouts, backoff, jitter or a retry budget, converting a partial outage into a self-inflicted overload.
  • Failing open on an authorisation check for availability reasons without an explicit, monitored, time-bounded decision.
  • Assuming an unbounded queue is more robust than one that rejects when full - it merely delays and worsens the collapse.

context