When designing a system, how do you decide between fail-fast and fail-safe (graceful degradation) behavior, and what are the trade-offs?
answer
- Classify the failure: my-bug → fail-fast; world-uncooperative → fail-safe
- Fail-safe ≠ fail-silent: degrade for users, alert operators
- Resilience kit: retries, circuit breaker, fallback, bulkhead, timeout
- Internal invariants fail-fast; integration edges fail-safe
- Money/safety-critical: correctness over availability
basics
~20 sUse fail-fast for programmer mistakes and clearly invalid input — stop immediately so bugs surface early. Use fail-safe (keep running with a fallback or degraded mode) for expected, recoverable problems in production, like a slow external service, where staying available matters more than perfection.
solid answer
~50 sFail-fast and fail-safe are not enemies; they apply at different layers. **Fail-fast** is right for **programmer errors and invalid internal state** — validate at boundaries, throw immediately, surface the bug during development and testing rather than corrupting data. **Fail-safe / graceful degradation** is right for **expected operational failures in production** — a downstream service times out, a cache is cold, a non-critical feature is unavailable — where availability and partial functionality beat a hard crash. The senior judgment is mapping each failure to its category: internal invariant violation → fail-fast (it's a bug, never recover silently); external dependency or transient condition → fail-safe with retries, circuit breakers, fallbacks, and bulkheads. Critically, even fail-safe behavior should **fail loudly to observability** (logs, metrics, alerts) — degrade for the user, but never hide the failure from operators. The wrong move is swallowing internal bugs (fail-silent) or crashing the whole system on a single non-critical dependency.
go deeper
Understands the basic difference: fail-fast stops on errors; fail-safe keeps running with a fallback. Can give a simple example of each.
Can pick the right strategy for common cases (validate input fail-fast; handle a downstream timeout fail-safe) and name patterns like retry and fallback.
Classifies failures into bug vs transient, applies circuit breakers/bulkheads/timeouts, and insists fail-safe paths still log/alert rather than going silent.
Defines system-wide failure-handling policy across services, decides correctness-vs-availability trade-offs per domain (incl. money/safety-critical), and ensures observability and blast-radius isolation are built into the architecture.
## Two strategies, two purposes - **Fail-fast**: on a problem, **stop immediately** and surface it (throw/abort). Best when continuing would corrupt data or hide a bug. Optimizes for **correctness and fast debugging**. - **Fail-safe** (a.k.a. fail-soft / graceful degradation): on a problem, **keep operating in a reduced or fallback mode** instead of crashing. Optimizes for **availability and resilience**. (Note: 'fail-safe' is overloaded. In Java collections it means a non-throwing iterator. In systems design it means graceful degradation. This question is the systems-design sense.) ## The decision rule: classify the failure The core skill is **classifying each failure mode**: 1. **Programmer error / broken invariant** (a `null` that should never be null, an impossible enum value, a corrupted internal data structure). → **Fail-fast.** This is a bug; recovering silently just spreads corruption and ships a defect. Throw, alert, and fix. 2. **Expected, external, transient condition** (downstream timeout, rate limit, network blip, optional feature down, cache miss). → **Fail-safe.** These *will* happen in production no matter how correct your code is; the system should absorb them and stay useful. A useful heuristic: *if the failure means my code is wrong, fail fast; if it means the world is temporarily uncooperative, fail safe.* ## Fail-safe mechanisms (resilience patterns) - **Retries with backoff** — for transient faults; cap attempts, add jitter. - **Circuit breaker** — after repeated failures, stop calling a sick dependency for a cooldown, returning a fast fallback; prevents cascading failure and resource exhaustion. - **Fallback / default values** — serve cached/stale data or a sane default when the live source is unavailable. - **Bulkheads** — isolate resources (thread pools, connection pools) per dependency so one failing dependency can't sink the whole service. - **Timeouts** — never wait forever; a missing timeout turns a slow dependency into a full outage. - **Graceful degradation** — turn off non-essential features (recommendations, analytics) while keeping the core path alive. ## The non-negotiable: fail loudly to observability The biggest trap is conflating **fail-safe** with **fail-silent**. Degrading for the *user* is fine; hiding the failure from *operators* is not. Every degraded path must emit logs, increment metrics, and (for important ones) trigger alerts. Otherwise you accumulate silent partial failures and discover them only via customer complaints. Fail-safe = *visible* recovery, never invisible. ## Trade-offs to articulate | Dimension | Fail-fast | Fail-safe | |---|---|---| | Optimizes for | correctness, fast debugging | availability, resilience | | Risk if misapplied | crashing on recoverable issues; poor UX | masking real bugs; serving wrong data | | Right for | internal invariants, dev/test, data integrity | external deps, transient faults, production UX | | Cost | downtime on edge cases | complexity, possible stale/partial results | ## Where they combine A mature service does **both at once**: fail-fast internally (strict validation, no swallowing of internal exceptions) and fail-safe at the integration edges (circuit breakers, fallbacks). Example: an e-commerce checkout fails fast if the *order total is internally inconsistent* (a bug), but fails safe if the *recommendations service* is down (just hide recommendations). And money/safety-critical paths often deliberately fail-fast even in production — better to reject a transaction than to process a possibly-corrupt one. ## Money/safety-critical nuance For financial, medical, or safety-critical operations, correctness usually outranks availability: prefer **fail-fast / fail-stop** (refuse rather than guess). 'Fail-safe' there can even mean failing *into a safe state* (e.g. brakes engage on power loss) — a third meaning worth naming if relevant.
- How is fail-safe different from fail-silent?Fail-safe degrades functionality for the user while staying available, but still records and alerts on the failure for operators. Fail-silent swallows the error with no signal at all, so problems accumulate invisibly. Good systems fail safe and loud; fail-silent is an anti-pattern.
- Give an example where the same system uses both strategies.A checkout service: fail-fast if the computed order total is internally inconsistent (that's a bug — throw and alert), but fail-safe if the recommendations microservice times out (catch it behind a circuit breaker, hide that widget, log a metric, and complete the purchase).
saying these in an interview costs you the question
- Treating fail-safe as license to swallow exceptions silently with no logging or alerting.
- Crashing the entire service because one non-critical dependency is unavailable.
- Silently recovering from an internal invariant violation (a real bug) instead of failing fast.
- Calling external dependencies with no timeout, turning a slow service into a full outage.
- Applying fail-safe defaults to money/safety-critical paths where correctness must outrank availability.