skip to content

A transit gate denies all travel when the entitlement service is down — how do you weigh fail-open against fail-closed?

level: principalimportance: should knowfreq 38%

answer

  1. one dependency, everyone denied
  2. denying is also a security failure
  3. not one global switch
  4. different actions, different degraded answers
  5. time-limited, recorded, reconciled later

basics

~20 s

Fail-closed turns one dependency's outage into total denial of service; fail-open opens a window of fraudulent travel. Remove the single point first, decide per action in advance, and bound the fallback in time, scope and audit.

solid answer

~50 s

This is a single point of failure on the critical path, so the availability threat exists with no attacker at all, and an attacker who reaches that one service buys a full outage cheaply. Fail-closed is not automatically the secure choice: at a gate it strands every traveller and loses the fare revenue, trading availability away wholesale. Fail-open trades the other way, opening a spoofing and elevation window where expired or cloned entitlements pass. Refuse the binary. First shrink the dependency with short-lived signed entitlements the gate verifies offline. Then split by consequence - allow ordinary entries, still deny refunds and high-value actions. Then bound the fallback: a hard time limit, a tamper-evident record of everything allowed, reconciliation afterwards, and an alert on entering degraded mode. The default itself is a business risk call made in advance, not by whoever is on call.

go deeper

for a junior

Be ready to explain that a component every request depends on is a single point of failure, and that denying everyone when it is down is itself a loss of availability.

for a middle

Explain both directions of the trade: fail-closed sacrifices availability, fail-open opens a window where invalid or expired entitlements pass, and cached signed entitlements shrink how often the choice arises.

for a senior

Show you would split the behaviour by action rather than flip one switch, distinguish unreachable from denied, and bound the fallback with time limits, tamper-evident records, reconciliation and alerting.

for a principal

Own the decision itself: who sets the default before the incident, what residual loss the business accepts, how it is written down and rehearsed, and how the accepted risk is recorded in the model rather than left implicit.

## Reading the design One service sits on the path of every gate decision. If it is unreachable, nobody travels. That is a denial-of-service entry in the model on its own merits, before any adversary appears: threat modeling asks what could go wrong, and *the entitlement service is unavailable* is a complete answer. Adding an adversary only makes it cheaper — a single component whose loss denies service to every user is the most attractive target in the architecture, so it concentrates risk rather than spreading it. The asset at stake is availability of a public service and the fare revenue attached to it, with a safety dimension: a station whose gates all deny entry at rush hour is a crowding problem, not merely an inconvenience. That framing is what makes the naive answer wrong. ## Why fail-closed is not automatically correct The reflex — *when in doubt, deny* — is right when the cost of a wrong allow dominates. Denying a wire transfer, a privilege grant or a key release while the authorisation service is down is defensible, because a wrongly allowed action is unrecoverable and the denied user can wait. At a transit gate the arithmetic inverts. A wrongly allowed passenger costs one fare, which is small, sometimes recoverable, and statistically bounded by how long the outage lasts. A denied station costs every fare for the duration plus a crowd-management incident. So *fail-closed is secure* is only true if availability is not a security property — and in STRIDE it plainly is. Choosing fail-closed is choosing to accept a denial-of-service threat in order to close a spoofing or elevation one. That is a legitimate trade, but it must be made as a trade, with both sides named. ## Do the structural work first Before arguing about the fallback, attack the single point: - **Push the decision local.** Issue short-lived, signed entitlements that the gate can verify offline against a public key it already holds. Now the common case involves no network call at all, and the central service becomes an issuer and a revocation feed rather than a synchronous dependency. - **Cache with a defined staleness.** A gate holding recent entitlement state can serve most travellers correctly through a short outage. Say explicitly how stale is acceptable, because that number *is* the residual fraud window. - **Remove correlated failure.** Redundant instances that share one database, one network path or one certificate authority are one component wearing three hats. Ask what a single expiring credential or one bad configuration push would take down at once. - **Distinguish reachability from health.** A gate that cannot reach the service is a different case from a service that answers *deny* — treating a timeout as a denial is exactly how one network blip becomes a system-wide refusal. Fail-open is only what you do when all of that has already failed. ## Split the decision, do not flip a global switch One binary for the whole system is a design smell. Different actions have different costs when wrong: | Action | Sensible degraded behaviour | |---|---| | Ordinary single entry | Allow, record, reconcile later | | Entry on a card flagged as lost or blocked | Deny — the local revocation list is small and cacheable | | Refund, balance transfer, high-value product | Deny — a wrong allow here is real money and hard to unwind | | Staff or maintenance access | Deny, with an out-of-band manual procedure | This is the move that separates a principal answer from a senior one: the question is not *open or closed* but *which decisions may degrade, by how much, and for how long*. ## Bound the degraded mode An unbounded fallback is a permanent hole that nobody remembers. Constrain it: - **Time**: offline tokens with short validity, and a hard limit after which the mode escalates rather than continuing silently. - **Record**: every allowed action written locally, signed or tamper-evident, and reconciled when connectivity returns. Allowing without recording is the version of fail-open that gets an organisation into trouble, because the loss is then unmeasurable. - **Visibility**: an alert when degraded mode is entered, and a continuing signal while it persists. Degraded mode must never be a quiet steady state. - **Rehearsal**: exercise the fallback on a schedule. A fallback path that has never run is a second untested system that will be discovered during the incident. ## Who decides Not the engineer at three in the morning. The choice is a business risk decision — expected fare loss and fraud exposure versus denied travel, crowd safety and reputational cost — so the product and risk owners set the default in advance, the design encodes it, and on-call may move only within pre-agreed bounds. Writing that default down, with the reasoning, is what turns a design-time argument into an operational policy that survives staff turnover. It also gives the threat model something to point at: the D threat is not *unmitigated*, it is *accepted, bounded and owned*, with a named residual.

  • What makes this a Denial of Service finding rather than a reliability concern for another team?
    STRIDE's D covers loss of availability whatever the cause, and the model asks what could go wrong rather than only who is attacking. A component whose loss denies service to every user is both an outage risk and the cheapest target in the design. Either way the mitigation family is the same: remove the single point, cache locally, degrade deliberately.
  • How do you stop the degraded mode becoming a permanent security hole?
    Bound it on every axis: short-lived offline tokens, a hard time limit after which the mode escalates, a tamper-evident local record of everything allowed, reconciliation when connectivity returns, an alert on entry and while it persists, and a scheduled rehearsal so the path is exercised and reviewed rather than discovered mid-incident.
  • Who should own the fail-open default, and when is it chosen?
    The product and risk owners, in advance. It is a comparison of expected fare loss and fraud exposure against denied travel, crowd safety and reputation — a business judgement, not an engineering preference. The design encodes the chosen default and on-call may deviate only inside pre-agreed bounds, so nobody improvises the risk appetite during an outage.
  • Would you give the same answer for an authorisation check on a money transfer?
    No. There a wrongly allowed action is unrecoverable and the denied user can retry in minutes, so fail-closed is the defensible default. The reasoning is identical — compare the cost of a wrong allow with the cost of a wrong deny — but the arithmetic points the other way. That is why the decision belongs per action, not per system.

It is the choice a bridge operator faces when the signalling fails: keep the barrier down and strand everyone, or wave traffic through under a written rule, with a log and a time limit. Both are decisions; only one of them is usually made in advance.

saying these in an interview costs you the question

  • Says fail-closed is always the secure option
  • Treats availability as outside the security remit
  • Adds retries and calls the single point mitigated
  • Leaves degraded mode unbounded and unlogged
  • Lets the on-call engineer set risk appetite mid-incident
  • Counts redundant instances sharing one database as redundancy

context