The user-mapping feed stalls mid-shift and entries expire: what should enforcement do?
answer
- the catch-all rule is the failure posture
- invisible allow versus visible outage
- a disabled posture is no posture
- age of the newest mapping
- freeze with a ceiling, do not extend
basics
~10 sNeither blanket allow nor blanket deny. Detect the stall, freeze existing mappings instead of letting them expire, and send unattributed traffic to a designed restricted policy rather than to a catch-all nobody chose.
solid answer
~50 sThe unknown-user rule is your failure posture, so it has to be a designed third state rather than an accident. Blanket allow means identity policy silently degrades to address policy - an intruder's traffic and everyone else's is permitted, and nobody finds out until an audit. Blanket deny means an entire shift loses access mid-shift and whoever is on call disables the rule under pressure and never puts it back, leaving you with an undesigned fail-open anyway. The workable answer has three parts: detect the stall from the mapping table itself - the age of the newest entry and the total entry count - rather than from tickets; freeze the mappings you already hold with a bounded staleness ceiling instead of expiring them into nothing; and route unattributed traffic to a pre-agreed restricted policy that keeps core business flows on address-based rules while denying the paths you refuse to allow blind.
go deeper
Know that traffic with no user mapping lands on a catch-all rule, and that whether that rule permits or denies is a deliberate choice with a cost either way.
Explain why a stalled feed drains the table gradually over one timeout rather than failing all at once, and what each of the two simple postures costs.
Demonstrate the designed middle state: detect from the table's own health signals, freeze with a staleness ceiling, route unattributed traffic to a pre-agreed restricted policy, and stage the rollout by population.
Own the argument that a posture predictably disabled under pressure is not a posture, and get the restricted-policy split agreed with the service owners before the day it is needed.
## What actually fails The mapping table is fed by something off-path - an agent fleet, a subscription to logon activity, a portal. When that feed stops, no new mappings are learned and, worse, the existing ones keep counting down toward their timeout. Nothing breaks at the moment of failure. It breaks progressively, over the length of one timeout, as user after user drops out of the table and their traffic starts resolving to no user. That shape matters, because it means the failure presents as a slow rise in unattributed traffic rather than a clean outage, and it is often first reported by users who cannot reach one specific thing. ## The two postures people reach for, and why both are wrong on their own **Unknown means allow.** Traffic with no user resolves onto address-based rules and keeps flowing. Nobody complains. That is exactly the problem: the identity layer you deployed and paid for is gone, every user-scoped restriction is void, and an adversary's traffic is permitted alongside everybody else's. The failure is invisible by construction and will typically be found by an audit, or by nobody. **Unknown means deny.** Traffic with no user is dropped. The failure is now extremely visible - a shift of workers loses access while logged in and working. The real-world outcome is worth being honest about: the service desk escalates, and someone on call disables or bypasses the rule to restore service. You now have a fail-open that nobody designed, applied under pressure, with no plan for reinstating it. A posture that will predictably be turned off is not a posture. So the confident answer *fail closed, obviously* is the one to aim at. In this specific system, closed means denying real users their jobs on the basis of an infrastructure fault they did not cause, and it is precisely how you end up with the worse of the two states. ## The designed third state **Detect the stall from the table, not from tickets.** Two signals live inside the mapping table itself: the **age of the newest entry** and the **total entry count** against the expected shift pattern. If the newest mapping is older than a normal interval between logons for that population, the feed is stalled - and you know it before the first user drops out. A rising hit count on the unknown-user rule is the confirming signal. **Freeze rather than expire.** On detecting a stalled feed, stop ageing entries out instead of letting the timeout empty the table. This is deliberately different from raising the timeout: freezing is bounded, explicit, reversible and comes with a staleness ceiling you set in advance, after which you stop pretending. Raising the timeout globally is a permanent configuration change made in a panic that also widens the window in which a reassigned address carries a stale name - the very failure the timeout exists to bound. **Route unattributed traffic to a policy you wrote earlier.** Not allow-all, not deny-all: a restricted set that keeps the business running on address-based rules while refusing the flows whose exposure you will not accept blind - administrative paths, and the destinations you only ever permit by group. The work is deciding that split calmly, in advance, and getting it agreed by the people who own the affected services. On the day, you are just switching to it. ## The evidence problem underneath A stalled feed does not only affect enforcement. For its whole duration your logs carry no user names, or - if you froze the table - names whose freshness you cannot vouch for. Anyone later reconstructing what happened during that window needs to know it. Record the stall period as a first-class fact: from this time to that time, attribution was frozen or absent. Otherwise a future investigation will read a stale name as a live one, and the reasoning built on it will be wrong in a way nobody can see. ## Rolling it out without being the outage When you first introduce a non-permissive unknown-user posture, stage it. Put one department behind the restricted policy, watch the unknown-user hit rate for a full shift cycle including the awkward populations - shared hosts, service accounts, contractors on portal mappings - and only then widen it. The traffic that lands on the catch-all in week one is almost always legitimate machine traffic nobody had rules for, and finding that in a pilot is much cheaper than finding it across the estate at 09:05 on a Monday.
- How do you detect a stalled mapping feed before your users do?Watch the mapping table as a health signal. Alert when the age of the newest entry exceeds a normal interval between logons for that population, and when the total entry count falls away from the expected shift pattern. Add the unknown-user rule's hit counter as confirmation. All three move well before the first ticket, because the table drains over the length of one timeout rather than emptying at once.
- Why is freezing existing mappings safer than raising the timeout during an outage?Freezing is bounded, explicit and reversible: you keep what you already had, learn nothing new, and set a staleness ceiling after which you stop trusting it. Raising the timeout is a global configuration change made under pressure that outlives the incident, and it lengthens the window in which a reassigned address carries a departed user's name - the exact failure the timeout exists to prevent.
- How do you introduce a non-permissive unknown-user rule without causing the outage yourself?Stage it by population. Put one department behind the restricted policy and watch the unknown-user hit rate through a full shift cycle, deliberately including shared hosts, service accounts and contractors on portal mappings. Almost all of the early hits are legitimate machine traffic that never had rules of its own; write those rules, then widen. Finding that in a pilot costs a few tickets, not a shift.
saying these in an interview costs you the question
- Answers fail closed without costing the outage it causes
- Leaves unknown-user as permit and calls it graceful degradation
- Waits for user tickets to discover the feed stopped
- Raises the timeout globally during the incident and never reverts it
- Does not record the stall window, so later investigations trust stale names