If your moderation classifier times out, should the request fail open or fail closed?
answer
- unknown is not the same as safe
- the answer differs by audience
- a third option exists besides block and allow
- log what bypassed the check
- availability multiplies when you block
basics
~20 sIt depends on the context, and it must be written down in advance. Fail closed where unchecked content is unacceptable, fail open where a classifier outage taking down the product is worse than a rare miss — and log every bypassed request for retrospective review either way.
solid answer
~50 sFail closed means blocking or holding content the classifier could not score; fail open means delivering it unchecked. Neither is universally right, so the answer an interviewer wants is a policy, decided per surface and per harm category, not a global default. A workable shape: fail closed in high-risk contexts — for a game, under-13 lobbies — and fail open in adult lobbies where a classifier outage would otherwise turn one dependency's downtime into a full product outage. The mechanics matter as much as the choice. Set an aggressive timeout, since a classifier that answers after the response has shipped is useless; put a circuit breaker in front so you fail fast during an outage instead of adding the timeout to every request; and stamp every bypassed item so it can be scored and actioned once the classifier recovers. A fail-open path with no audit log is not a degraded mode, it is a silent hole.
code
python · 15 linesFAIL_CLOSED_CONTEXTS = {"under_13_lobby"}
def classify(text, timeout_s):
raise TimeoutError("classifier unavailable")
def moderate(text, context, timeout_s=0.3):
try:
return classify(text, timeout_s)
except TimeoutError:
if context in FAIL_CLOSED_CONTEXTS:
return {"action": "block", "reason": "classifier_unavailable"}
return {"action": "allow", "reason": "classifier_unavailable", "audit": True}
print(moderate("gg wp", "under_13_lobby"))
print(moderate("gg wp", "adult_lobby"))go deeper
Know the two terms: fail closed blocks content that could not be checked, fail open lets it through, and the choice has to be deliberate.
Explain the tradeoff in terms of which failure you are converting the outage into, and describe the timeout and retry behaviour that makes either choice workable.
Show the operational depth: circuit breaking so outages stay fast, hold-and-score as a third option, an audit trail with retroactive action, and different policies per surface.
Own the written policy and its sign-off — which audiences get fail-closed treatment, what availability commitment that implies, what redundancy it funds, and how the unchecked window is reported after an incident.
## The two options, precisely When the moderation classifier is slow, erroring, or unreachable, the calling system must decide what to do with content it could not score. - **Fail closed**: treat unknown as unsafe. Block the message, or hold it pending a later verdict. - **Fail open**: treat unknown as safe. Deliver it, unchecked. This is a real decision that exists in every deployment, and the most common defect is that nobody made it — the behaviour is whatever the try/except in the client happened to do. ## Why it is a policy decision, not a default Each option converts a dependency failure into a different product failure. Fail closed converts a classifier outage into an availability outage: during a 20-minute classifier incident, nobody can talk. Fail open converts it into a safety gap: for 20 minutes, everything ships unchecked. Which is worse is a function of the surface, the audience and the regulatory exposure — exactly the kind of question engineering cannot answer alone and should not answer implicitly. ## Differentiating by context The strongest version of this answer is not a single choice but a documented matrix. For a competitive shooter's chat, a defensible written policy is: **fail closed in under-13 lobbies, fail open in adult lobbies.** In the first, unchecked chat between minors is a risk the business will not accept, and going quiet for a few minutes is a tolerable degradation. In the second, the population is adults, the harm ceiling is lower, and shutting down all communication because one internal service is unhealthy is a worse product outcome than a short unchecked window. The same differentiation applies per category and per placement. You might fail closed on the output side (your product's own voice) and fail open on the input side (a user's message to the model), or fail closed only for the categories where a miss is catastrophic. ## Mechanics that make either choice work **Aggressive timeouts.** A moderation verdict that arrives after the content has shipped is worthless. Budget the classifier tens to low hundreds of milliseconds, and treat exceeding it as a failure rather than waiting. **Circuit breaking.** Without a breaker, an outage means every request pays the full timeout before falling through. With one, after N consecutive failures you stop calling for a cool-off period and apply the fail policy immediately, so latency stays flat during the incident. **A third option: hold.** Fail open and fail closed are not the only choices. For asynchronous surfaces you can *hold* — accept the content, do not deliver it yet, queue it for scoring when the classifier returns. This preserves safety without a hard block, at the cost of delivery latency, and it is often the best answer for anything that is not a live conversation. **Degraded fallback.** A cheap local heuristic — a blocklist, a lightweight local model — can cover the outage window. It is far worse than the real classifier, but far better than nothing, and it is a legitimate middle setting. ## Paying back the debt Every item that bypassed moderation must be tagged with the reason (`classifier_unavailable`) and persisted. When the classifier recovers, replay that set and take retroactive action: remove content, apply enforcement, escalate what a reviewer needs to see. Without this, fail-open is not a degraded mode, it is a permanent gap that appears in no metric. The audit log is also what lets you quantify the incident afterwards — how many messages went unchecked, how many were harmful — which is the input to whether the fail-open policy was correct. ## Availability arithmetic Fail closed makes your product's availability the *product* of your availability and the classifier's. A 99.9% product in front of a 99.5% classifier is a 99.4% product. Teams are routinely surprised by this. If you choose fail closed on a critical path, the classifier's availability target becomes your availability target, which means redundancy: a second provider or a self-hosted fallback, with the failover itself tested. ## Failure modes - No decision at all — the behaviour is an accident of exception handling. - Fail open with no audit trail, so the gap is invisible. - Retry loops with no cap that turn a slow classifier into cascading timeouts. - One global policy applied to an under-13 surface and an adult surface alike. - A fail-closed choice on a critical path with no redundancy behind the classifier. ## What interviewers listen for That you frame it as a per-context policy rather than picking a side, that you name hold-and-score as a third option, that you insist on an audit log and retrospective action, and that you understand fail closed puts the classifier's availability directly into your own.
- Besides blocking and allowing, what third option exists when the classifier is unavailable?Hold: accept the content, withhold delivery, and queue it for scoring once the classifier recovers. It keeps the safety property of fail-closed without a hard rejection, at the cost of delivery latency. It works well for asynchronous surfaces such as posted messages or generated documents, and poorly for live conversation where delayed delivery is itself a broken experience.
- What must you record when a request bypasses moderation under a fail-open policy?The content, the context and audience, the timestamp, and an explicit reason code such as classifier_unavailable. That record is what lets you replay the backlog through the classifier after recovery, take retroactive enforcement, and quantify the size of the gap. Fail-open without this trail is an invisible hole rather than a degraded mode.
- How does choosing fail-closed on a critical path change your availability target?It multiplies your availability by the classifier's. A 99.9% service gated on a 99.5% classifier delivers about 99.4%. That makes the classifier's uptime your uptime, so a fail-closed choice obliges you to build redundancy behind it — a second provider or a self-hosted fallback — and to test the failover, not just declare it.
saying these in an interview costs you the question
- Picks one global policy without asking about the surface
- Fails open with no record of what bypassed the check
- Treats a timeout as if the content were scored safe
- Retries the classifier indefinitely, cascading the latency
- Chooses fail closed on a critical path with no fallback classifier