A feature-flag service's if/elif rule chain leaves some requests with no decision. How do you diagnose and prevent that?
answer
- The value that is not a value
- What runs when nothing matches?
- Chains without a final else
- Two causes look identical to the caller
- Make it total: log, default, or raise
basics
~20 sA chain with no final else falls through silently, so the function returns None instead of a decision. Reproduce the input, confirm nothing matched, then make the chain total: an else that logs and returns an explicit default, or raises.
solid answer
~50 sThe symptom is an implicit `None`: when no `if`/`elif` condition is true and there is no `else`, execution falls out of the statement and a function whose branches each `return` returns `None`. Diagnosis is to capture the exact inputs that produced it and separate two causes -- no branch matched, versus a branch matched and returned `None` itself -- then read the chain for the gap. Here the branches compared a timestamp against rule windows and each branch read the clock again; a clock-skew artefact put one read before every window start, so nothing matched. The fixes are structural: hoist the timestamp so every branch sees one consistent value, close the chain with an explicit `else` that logs the unmatched input and returns a safe **fail-closed** default -- flag off -- or raises where a silent default would be worse, and prove totality with branch coverage plus a test asserting a non-`None` result for generated inputs.
code
python · 8 linesdef rollout(percent):
if percent >= 100:
return "on"
elif percent > 0:
return "partial"
print(repr(rollout(0))) # Nonego deeper
Remember that a function falling past an if/elif chain with no else returns None implicitly. Ending value-producing chains with an explicit else is the habit that prevents it.
Explain the mechanism and the two causes that look identical to a caller -- nothing matched, versus a branch that returned None itself -- and know that an annotated return type lets a static checker flag the fall-through path.
Show the production instinct: capture the inputs before touching code, trace where the None originated rather than where it crashed, and separate the local fix from the structural one -- a single shared input value, a logged fail-closed default, a totality test.
Own the policy. Decide where the system fails closed versus fails loud, where inputs get validated so chains only see a known domain, and how gaps become monitoring signals instead of incidents nobody sees until a downstream service breaks.
### The mechanism behind the symptom An `if`/`elif` chain with no `else` is a legal statement that may do nothing at all. Execution simply falls past it. In a function whose branches each `return`, falling past the chain means reaching the end of the function body, and Python returns `None` implicitly. There is no error and no log line -- the caller receives a value that looks like a decision and is not one. ```python def rollout(percent): if percent >= 100: return "on" elif percent > 0: return "partial" # percent <= 0 falls through rollout(0) # None, not "off" ``` In a flag service consulted by a dependency graph of seventeen services, that `None` travels. It is stored, compared, serialised, and finally raises an attribute or type error several hops away from the chain that produced it, which is why the first instinct -- fix the service that crashed -- fixes nothing. ### Diagnosing it 1. **Capture the exact inputs.** Log or record the arguments at the call that produced the bad value, including the values the conditions actually compared. Without them you are guessing at which branch should have won. 2. **Separate the two causes.** `None` can mean *no branch matched* or *a branch matched and itself returned `None`*. They have different fixes, and you distinguish them by adding one temporary log line in the fall-through position, or by asserting a non-`None` result at the chain's exit while reproducing. 3. **Read the chain as a partition.** For each branch, ask which inputs it claims; then ask what is left over. The leftovers are the gap. Boundary values -- exactly zero, exactly equal to a window edge, empty collections -- are where gaps live. 4. **Check that the branches agree on their inputs.** If several branches recompute a value rather than sharing one, they can disagree. That was the cause here: each branch called the clock separately, and a clock-skew artefact between reads put one comparison before every rule's start time, so every window test failed and the chain fell through. Hoisting the timestamp into a single value computed once at entry removes the whole class of bug -- every branch then decides against the same snapshot. ### Preventing it **Make the chain total.** Every chain that produces a value should end in an `else`, and the `else` should say something. Two defensible endings: - **An explicit, logged default.** For a request path, silently failing the whole request is worse than serving the conservative answer. For a feature flag the conservative answer is *off* -- fail closed -- and the `else` should log the unmatched input at warning level so the gap is visible in monitoring rather than invisible in production. - **A raise.** For configuration validation, batch jobs, or anywhere a wrong answer is worse than no answer, `else: raise` with the offending value in the message. It converts a silent wrong decision into a loud, locatable failure. What the `else` must never be is `else: pass`, or an absent `else` justified as "it can't happen". Those are the same thing with different amounts of confidence. **Declare the return type.** A function annotated as returning a string but able to fall through returns `None` on that path; a static type checker reports the missing return. That turns a runtime incident into a check that runs before merge -- but only if the annotation is there. **Test for totality, not just for the happy branches.** Branch coverage tells you which branches ran; it does not tell you that some input reaches none of them. Add a test that feeds a spread of inputs -- including every boundary value and one input outside each window -- and asserts the result is never `None`. A randomised or property-style test over the input domain is stronger still, because the gap you did not think of is exactly the one that bit you. **Validate at the boundary.** Much fall-through comes from inputs the chain was never designed for. Parsing and validating at the edge of the system, so the chain only ever sees values from a known domain, shrinks the number of branches the chain has to cover and makes the `else` genuinely unreachable rather than merely unlikely. ### What the interviewer is listening for That you reach for the inputs before the code; that you name the implicit-`None` mechanism precisely; that you distinguish *no match* from *matched and returned `None`*; that you separate the local fix (add the `else`) from the structural one (one shared input value, validation at the boundary, a totality test); and that you have an opinion on fail-closed versus fail-loud rather than defaulting to whichever is easier to type.
- Why is returning None on fall-through worse than raising?Because `None` is a plausible value. It passes through assignments, comparisons and serialisation, and only fails several call frames later where the original inputs are gone. A raise fails at the chain, with the offending value in the traceback, so the diagnosis is the stack trace rather than an investigation. The cost of raising is availability, which is why the choice depends on the path.
- When should the final else return a safe default rather than raise?On a live request path, where failing the request is worse than serving the conservative answer -- for a feature flag that means off, fail-closed. Raise where a wrong answer costs more than no answer: config validation, batch computation, anything that writes. Either way the else must log the unmatched input, so a default never hides a growing gap.
- How do you prove in tests that such a chain is total?Branch coverage shows which branches ran, not that every input reaches one, so add a test that feeds boundary values and out-of-window inputs and asserts the result is never `None`. A property-style test over the generated input domain is stronger, because it explores the case you did not think of. An annotated return type also lets a static checker flag the fall-through path.
- Why did hoisting the timestamp out of the branches matter here?Because each branch read the clock independently, so the branches were deciding against different values; a skew between reads pushed one comparison outside every rule window and nothing matched. Computing the value once at entry and passing it in makes the chain a pure decision over a consistent snapshot -- which also makes it testable, since the test supplies the timestamp.
saying these in an interview costs you the question
- Assumes a chain without else always returns a value
- Adds a bare else: pass to silence the fall-through
- Reads the wall clock separately inside each branch
- Treats a None decision as equivalent to the flag being off
- Fixes only the downstream service that raised the error
- Says it cannot happen instead of making the chain total