When the risk scorer misses its deadline on a login submit, should the login path fail open or fail closed?
answer
- degrade, do not flip a switch
- a third action exists
- cached prior, then rules, then exposure
- the failure hits every login at once
- risk owner signs it before the incident
basics
~20 sNeither as a blanket default. Degrade down a ladder — a cached prior score for that account, then request-local rules, then an exposure-based default — and reserve a hard decline for high-exposure actions only, because a scorer timeout hits every login at once.
solid answer
~50 sThe honest answer is that the login path degrades rather than flips a switch. First reach for a **cached prior score** for that account or device if it is recent enough; if there is none, apply **rules that read only request-local data**, which need no feature fetch; failing that, take an action keyed to the **exposure of what the session can do** — allow and mark an ordinary session, but challenge a submit that immediately enables a payout or a password change. Fail-closed as a universal default turns a scorer outage into a login outage, because a timeout is a correlated failure that hits every request at once, not a per-user accident. Fail-open as a universal default leaves the door open for the whole incident. Which one applies is a risk-owner decision signed off in advance, not an on-call choice.
code
pseudocode · 12 linesfallback(reason):
prior = lookup(scoreCache, accountKey) // score from an earlier attempt
if prior.exists and age(prior) < 15min:
return actionFor(prior.score, degraded = true, reason)
if requestLocalRuleMatches(request): // no feature fetch needed
return CHALLENGE with reason
if exposureOf(request) == HIGH: // payout or credential change
return CHALLENGE with reason
return ALLOW_AND_MARK with reason // session flagged for later reviewgo deeper
Know that scoring can time out and that the login path must still answer, so a pre-decided default action exists for the case where no risk score is available.
Be able to lay out the ladder in order and say what each rung depends on, and why a fallback must not rely on the component that failed.
Argue from the correlated nature of the failure: the outage hits all logins at once, so the degraded action has to be sized for full traffic and tiered by what the session can do.
Own the policy as a business line — what loss the company accepts while unscored, and what login friction it will spend to avoid it — and make sure it is signed and rehearsed before it is needed.
## The question is not binary "Fail open or fail closed" frames a scorer timeout as a single switch, and design rounds use it precisely to see whether the candidate accepts the framing. On a consumer login path there is a third action available that is neither: a **challenge** — a step-up authentication that costs the legitimate user a few seconds and costs an attacker the whole attempt. Most of the design work is deciding when each of the three is right. It also matters **which fallback** is meant. Two different things get called "the fallback": a **rules-only degraded path** that decides from data already in the request, and a **cached prior score** computed for this entity on an earlier request. They fail differently and they belong in a specific order. ## The ladder 1. **Cached prior score.** If this account or device was scored recently and the cached risk score is still within its staleness bound, reuse it. It is a real score from real features and it costs one lookup — but it says nothing about *this* attempt, so it is only usable while it is fresh. 2. **Request-local rules.** Decide from what arrived with the request and needs no feature fetch at all: a known-bad credential-stuffing pattern, an allowlisted corporate context, a malformed or replayed client fingerprint. This path survives a total feature-store outage, which is the case the cached score does not. 3. **Exposure-based default.** With neither of the above, act on what the granted session would be able to do. | the submit leads to | degraded action | what you are accepting | |---|---|---| | an ordinary read-only session | allow, and mark the session for later review | some unscored takeovers get in | | a session that can change credentials or contact details | challenge | friction for every user during the incident | | a session that can immediately move money out | challenge, and hold the sensitive action behind a second check | conversion loss on the highest-value flows | The marking in row one matters: an allowed-unscored session should be flagged so it can be re-examined once scoring recovers, rather than disappearing into normal traffic. ## Why the correlated nature of the failure decides it A scorer timeout is almost never one unlucky user. The causes — a slow low-latency key-value store, a model host rolling, a saturated network path — hit **all** requests in the same seconds. That single fact drives the whole design: - **Fail-closed at 100%** is an authentication outage. Every customer is blocked for the duration, and support volume and reputational damage usually exceed the fraud avoided. - **Fail-open at 100%** is an open window, and a sophisticated attacker will notice the window is open — the degraded path is itself a target, which is why a scorer that can be pushed into timing out is a risk in its own right. - **Challenge at 100%** is capacity-bound. If the step-up channel is sized for the few percent of traffic normally routed to it, sending every login there melts it. Either the channel is sized for a degraded day or the degraded action is tiered by exposure. ## Who owns the call This is a signed-off policy, not an engineering default. The action taken when no score is available accepts either loss or lost customers, and both belong to the risk owner and the product owner together. What engineering owes them is: - the behaviour written down per exposure tier **before** the incident, and rehearsed; - a **reason code** on every degraded decision, so the number and shape of unscored logins is measurable rather than inferred; - an alert on the **rate of degraded decisions**, which is the earliest visible symptom of a scoring problem; - the reserve in the latency budget that makes the degraded path actually run on time. ## Failure modes to avoid - **A degraded path that also needs the thing that broke.** Rules that read the same online store the timeout came from are not a fallback. - **A stale cached prior with no staleness bound.** A score from three days ago is a fabrication about the current attempt. - **No reason code.** The decisions are then indistinguishable from scored ones after the fact and cannot be re-examined. - **Untested path.** The fallback runs only during incidents, so it is the code most likely to be broken when it is finally needed; exercise it deliberately. - **Flipping the policy during the incident.** Choosing between loss and a login outage under pressure, with no pre-agreed line, is how both get chosen badly.
- Why is a blanket fail-closed policy usually rejected on a consumer login path?Because the timeout is correlated: whatever broke the scorer broke it for every request, so fail-closed blocks the entire customer base for the duration rather than a handful of risky sessions. The avoided fraud is small next to an authentication outage, and the recovery is slower because support and reset volume spike at the same time.
- What has to be true of the rules-only degraded path for it to count as a fallback at all?It must not depend on anything that could be the cause of the timeout. If it reads the same online store, or calls the same model host, it fails in exactly the situation it exists for. A genuine degraded path decides from data carried in the request itself, plus at most an independently-failing cache.
- What should be recorded when a login is decided by the fallback?A reason code naming which rung of the ladder fired and why the score was missing, attached to the decision. Without it, degraded decisions are indistinguishable from scored ones afterwards, the rate cannot be alerted on, and the sessions allowed unscored cannot be pulled back for review once scoring recovers.
saying these in an interview costs you the question
- Treating fail-open versus fail-closed as the whole design
- Forgetting that a challenge is available and is neither open nor closed
- A degraded path that reads the same store whose slowness caused the timeout
- Reusing a cached prior score with no staleness bound
- Assuming the step-up channel can absorb one hundred percent of logins
- Leaving the policy to whoever is on call during the incident