When should a system deliberately NOT fail fast, and how do you decide between failing fast and degrading gracefully?
answer
- can I still be correct without it?
- degrade features, fail requests
- degradation must be visible, never silent
- identical check on every node = fleet outage
- timeouts + circuit breaker = fail fast at the edge
basics
~20 sFail fast when continuing would give wrong answers or corrupt data. Degrade instead when the failure only affects a non-essential part and the core result is still correct - for example hiding a recommendations panel rather than failing the whole page.
solid answer
~50 sThe decision hinges on whether continuing can produce an incorrect result. If an invariant is broken or the data needed for a correct answer is missing, fail fast: a visible error is cheaper than silently wrong output or corrupt persisted state. If the failure is confined to an optional capability and the primary answer is still correct, degrade - omit the panel, serve a stale cached value clearly marked as such, or fall back to a simpler algorithm. Two extra dimensions matter. First, correlated failure: a fail-fast check that every instance evaluates identically, such as a bad config value crashing on boot, can take a whole fleet down at once, so such checks need canaries and staged rollout. Second, at service boundaries fail fast changes meaning: rejecting work immediately when a dependency is known unhealthy (open circuit breaker, full bounded queue, exceeded deadline) is how you avoid queue buildup and cascading failure. Fail fast on the *request*, degrade on the *feature*, and never trade correctness for uptime silently.
go deeper
Give the core test: if the result would be wrong, stop; if only an extra feature is missing, show the rest without it. One concrete example is enough.
Add the requirement that degradation be visible (stale markers, metrics, alerts) and mention timeouts and fallbacks as the practical mechanisms.
Discuss cascading failure, circuit breakers, bounded queues, load shedding, deadline propagation, retry budgets, and the fail-closed-versus-fail-open choice for authorisation.
Frame it as a per-capability correctness-versus-availability policy tied to business impact, address correlated failure from identical fail-fast checks across a fleet, rollout and rollback strategy, crash-loop amplification, and how degraded modes are made observable and time-bounded rather than permanent.
## Two questions that get confused "Should I fail fast?" is really two questions: 1. **In-process correctness:** an invariant is broken or an argument is invalid. Here fail fast is nearly always right - continuing yields wrong results. 2. **System behaviour under an environmental fault:** a dependency is slow, a cache is cold, a third-party API is down. Here the answer is a design decision, not a principle. Conflating them produces both bad extremes: services that crash on any hiccup, and services that swallow everything and serve nonsense. ## The deciding test **Can the system still produce a correct primary result without the missing piece?** - **No** -> fail fast. A payment service that cannot reach the ledger must not "assume success". A pricing service without the current rate must not guess. - **Yes** -> degrade. A product page whose recommendation service is down should render the product. A dashboard whose analytics panel times out should render the rest with an explicit "unavailable" marker. The crucial refinement: degrading must be **visible**. A stale value served silently is indistinguishable from a fresh one and re-introduces silent corruption at the system level. Mark it stale, expose it in metrics, alert if the degraded mode persists. ## Vocabulary - **Graceful degradation** - continuing with reduced functionality when a component is unavailable. - **Fail-safe / fail-secure** - on failure, move to the state that is safest. Note these can point in opposite directions: a fail-safe door unlocks in a fire (life safety), a fail-secure door locks (asset protection). For authorisation, failing *closed* (deny) is the fail-fast-flavoured choice; failing *open* (allow) trades correctness for availability and is usually unacceptable. - **Load shedding** - deliberately rejecting a fraction of work immediately when overloaded, so accepted work still meets its deadline. - **Backpressure** - signalling upstream to slow down rather than accepting unbounded work. - **Circuit breaker** - after repeated failures against a dependency, stop calling it and fail immediately for a cool-down period. - **Bulkhead** - isolating resources per dependency so one slow dependency cannot consume all threads or connections. ## Fail fast at service boundaries means something specific In distributed systems, "fail fast" usually means *reject immediately rather than wait or queue*. The failure mode it prevents is **cascading failure**: a slow dependency causes callers to hold threads and connections; queues grow; latency crosses the caller's own deadline so the work is wasted anyway; the caller's callers then saturate. Rejecting early keeps the system alive. Concrete mechanisms: aggressive **timeouts** and **deadline propagation** (never an unbounded wait), **bounded queues** that reject when full instead of growing, **circuit breakers**, **concurrency limits**, and **retry budgets with jitter and exponential backoff** so retries do not become a self-inflicted denial of service. ## Where fail fast actively hurts 1. **Correlated crash on boot.** A validation that every instance evaluates identically - a malformed config value, an unreachable dependency at startup - crashes the entire fleet simultaneously and can prevent rollback if the new instances never become healthy. Mitigation: canary and staged rollout, validate config in CI, and distinguish "cannot start correctly" (crash) from "a non-essential dependency is unavailable" (start, report unready or degraded). 2. **Crash loops as an amplifier.** Restarting on failure plus a dependency that is down equals a restart storm hammering the dependency. Mitigation: backoff on restart, and readiness signalling rather than exiting. 3. **User-facing input.** Aborting on the first bad field is technically fail fast but poor UX; collect and report all errors at once. Nothing invalid is constructed either way. 4. **Best-effort data paths.** In log ingestion, metrics or a media pipeline, dropping or skipping a malformed record and counting it is usually right - but the counter must be monitored, otherwise it is silent loss. ## Putting it together A well-designed service typically does all of the following at once: fails fast on invalid arguments and broken invariants inside the process; validates and rejects malformed input at its boundary; fails fast on *requests* to unhealthy dependencies via timeouts and circuit breakers; and degrades *features* that are not essential to the primary result, loudly and observably. The unacceptable combination is silent tolerance of incorrectness - and the second-worst is crashing everything because one optional feature is unavailable.
- A service starts up but its non-essential recommendation dependency is unreachable. Should it exit?No. Start, mark itself ready for its core function, and report the degraded capability through health and metrics. Exiting turns a partial outage into a total one - and if the dependency is down for everyone, every instance crashes simultaneously. Reserve boot-time crashes for things without which the service cannot be correct at all.
- How do timeouts relate to fail fast?An unbounded wait is the ultimate fail-slow: the thread, connection and the caller's deadline are all consumed for work that will be discarded anyway. A timeout converts an indefinite hang into a prompt, explicit failure that the caller can act on, which is what makes circuit breakers, retry budgets and load shedding possible.
- Is failing open ever the right choice for an authorisation check?Very rarely, and only as an explicit, documented, time-bounded business decision with monitoring - because it trades a correctness guarantee for uptime. The default should be fail closed: if you cannot verify permission, deny. Silently granting access on infrastructure failure is a security incident waiting to happen.
An aircraft: a failed entertainment system does not ground the flight (degrade the feature), but a failed altimeter reading is not guessed at - the crew switches to a verified source or diverts (fail fast on correctness). What is never acceptable is displaying a made-up altitude that looks real.
saying these in an interview costs you the question
- Treating fail fast as an absolute rule and crashing the whole request because one optional widget's data is missing.
- Serving stale or fallback data silently, with no marker and no metric, so degradation is indistinguishable from healthy operation.
- Adding a strict startup check that every instance evaluates identically without a canary or staged rollout, creating a fleet-wide correlated outage.
- Retrying a slow dependency without timeouts, backoff, jitter or a retry budget, converting a partial outage into a self-inflicted overload.
- Failing open on an authorisation check for availability reasons without an explicit, monitored, time-bounded decision.
- Assuming an unbounded queue is more robust than one that rejects when full - it merely delays and worsens the collapse.