What does the "fail fast" design principle mean, and why is stopping at the moment invalid state is detected usually better than letting execution continue?
answer
- cause-to-symptom distance
- no silent corruption
- guard at the boundary, trust inside
- invariant broken -> stop now
- crash beats wrong data
basics
~20 sFail fast means checking for invalid input or state right where it appears and stopping immediately with a clear error, instead of continuing with bad data that causes a confusing failure much later somewhere else.
solid answer
~50 sFail fast is detecting an invalid argument, broken object invariant, or missing dependency at the earliest possible point and immediately raising a loud, specific error instead of tolerating it. Its value is shortening the distance - in code, in time, and in stack frames - between the cause (the defect) and the symptom (what you observe). A null, NaN, negative quantity, or half-initialised object that is silently accepted travels through layers, gets persisted, and later surfaces either as an unrelated crash or as quietly wrong data nobody notices. Typical mechanisms: precondition guards at the top of a function, constructor and parameter validation, schema validation at boundaries, assertions on internal invariants, and types that make invalid values unrepresentable. The cost is more upfront checking and a system that halts rather than limps; the payoff is far cheaper debugging and no silent corruption.
code
pseudocode · 11 lines// fail late: bad quantity is stored, discovered next month
function addItem(cart, sku, qty) {
cart.lines.push({ sku, qty }) // qty = -3 accepted silently
}
// fail fast: rejected at the point of the mistake
function addItem(cart, sku, qty) {
if (sku == null || sku == "") throw new ArgumentError("sku must be non-empty")
if (qty <= 0) throw new ArgumentError("qty must be > 0, got " + qty)
cart.lines.push({ sku, qty })
}go deeper
Define it plainly: check inputs at the top of the function and throw a clear error instead of continuing with bad data. Give one concrete example such as rejecting a negative quantity.
Add the mechanisms - precondition guards, constructor validation, boundary schema validation - and explain the cause-to-symptom distance argument, plus where to validate so checks are not duplicated everywhere.
Frame it as a correctness-versus-availability trade-off, distinguish programming errors from expected environmental faults, and mention making invalid states unrepresentable as the strongest form.
Discuss shifting detection left (compile time > startup > first request > never), the blast radius of silent data corruption versus a visible crash, and the organisational cost of debugging late-surfacing defects.
## The idea **Fail fast** means: the moment the program can prove something is wrong - an argument violates the contract, an object's invariant is broken, a required configuration value is missing - it should **stop and report loudly** rather than continue and hope. Some vocabulary used below: - **Invariant** - a statement that must always be true about an object or system ("a shopping cart's total equals the sum of its lines", "an order always has a non-empty customer id"). - **Precondition** - what a function requires of its *caller* before it will work ("quantity must be greater than zero"). - **Postcondition** - what the function promises in return. - **Contract** - the set of preconditions, postconditions and invariants. This vocabulary comes from *Design by Contract*, introduced by Bertrand Meyer with the Eiffel language. ## The opposite behaviours - **Fail late** - the bad value is accepted, flows through several layers, and blows up far from the real cause. The stack trace points at the *victim*, not the *culprit*. - **Fail silently** - the worst case: nothing crashes at all. A bad price is written to the database, a NaN spreads through a report, a permission check quietly returns a default. Corruption is discovered days later, when the evidence is gone and the bad data has been copied into backups, caches and downstream systems. ## Why early detection is cheap and late detection is expensive Debugging cost is dominated by the *search space* between cause and symptom. If a function rejects a negative quantity on line one, the stack trace names the exact caller that passed it - the bug is essentially already located. If instead the negative quantity is stored and only produces a strange invoice next month, you must reconstruct history across many components. Each layer the corrupt value crosses multiplies the candidates you have to consider. There is also a **blast-radius** argument. An unhandled crash affects one request and is visible in monitoring. Silent corruption affects data that may be irreversible, and it is invisible - you cannot alert on a failure you never detected. ## Mechanisms, from strongest to weakest 1. **Make invalid states unrepresentable** - use a dedicated type (`EmailAddress`, `PositiveQuantity`) or a sum type/enum instead of a raw string or int, so the compiler rejects the mistake before the program runs. Strongest, because the failure happens at build time. 2. **Constructor / factory validation** - an object cannot exist in an invalid state, so no later code has to re-check. 3. **Precondition guards** - explicit checks at the top of a function that throw a specific error naming the parameter and the rule. 4. **Boundary schema validation** - validate untrusted input (HTTP body, message payload, config file) once, at the edge, and convert it to trusted domain types. 5. **Assertions** - checks of internal assumptions, often compiled out in production; useful for developer errors, not for untrusted input. ## Costs and honest limits - More code and more tests for error paths. - A strict system **stops** where a lenient one might have muddled through; that is a deliberate trade of availability for correctness and it is not always the right trade (a video player should not crash because one subtitle line is malformed). - Guards can degenerate into paranoid null checks everywhere. The discipline is to check **once, at the boundary or in the constructor**, and then trust the value inside. - Failing fast in the middle of a multi-step mutation can leave partially applied changes, so it should be paired with validate-everything-first or with transactions. ## The mental rule Fail fast when continuing could produce **wrong results or corrupt data**. Prefer degrading when the failure is confined to a non-essential feature and the correct core result is still available.
- Does failing fast contradict writing robust, fault-tolerant software?No - they answer different questions. Fail fast governs *programming errors and invalid internal state*, where continuing yields wrong results. Fault tolerance governs *expected environmental failures* (a slow dependency, a dropped packet), which should be handled, retried or degraded. A robust system fails fast internally and handles expected faults explicitly at its boundaries.
- Where should the validation live if the same value is passed through five layers?Validate once at the trust boundary (parsing untrusted input) or in the constructor of the type that carries the value, then pass the validated type inward. Repeating identical checks in every layer is noise; the inner layers should be able to trust their types.
- What does a good fail-fast error message contain?What rule was violated, the offending value (redacted if sensitive), the parameter or field name, and enough context to identify the caller. "qty must be > 0, got -3" is actionable; "invalid input" is not.
A smoke detector that beeps the second it senses smoke, versus one that waits to be sure and only alarms once the house is fully alight. The early alarm is annoying and sometimes triggered by toast, but it points straight at the kitchen; the late alarm tells you only that something, somewhere, is on fire.
saying these in an interview costs you the question
- Claiming fail fast means "crash the whole application on any error", including expected user input errors.
- Treating a caught-and-swallowed exception (empty catch block, returning null or a default) as acceptable error handling.
- Adding null checks in every layer instead of validating once at a boundary and trusting the type afterwards.
- Saying fail fast is only about exceptions - it is equally about types, constructors and startup checks.
- Arguing that not crashing is always better for users, ignoring that silently persisting corrupt data is usually worse than a visible error.