Your EAP server certificate expired at 03:00 and every 802.1X site is dark: why won't dead-server fallback trigger, and what break-glass path won't also admit an intruder?
answer
- the servers were never unreachable
- one object, whole fleet, same instant
- an answer arrived, so no timers ran
- the fix path is behind the door
- off switches need scope and a timer
basics
~20 sBecause the server answered. Supplicants that validate the expired certificate abort the exchange and the server returns a reject, so the switch sees an authentication failure, not an unreachable server, and the unreachable-server fallback never applies.
solid answer
~60 sIn EAP-TLS and the tunnelled methods, the supplicant validates the authentication server's certificate before it will continue. When that certificate expires, every supplicant configured to check it refuses at the same instant — one object, one expiry, the whole fleet. Critically the servers are healthy and reachable: they answer, the exchange fails, and a reject comes back. Dead-server timers never run, so the critical VLAN you built for the unreachable case does not engage. Ports either land in an auth-fail VLAN if you configured one, or stay unauthorized. Recovery therefore depends on a break-glass path that does not itself run through 802.1X: out-of-band console or cellular access to the site switches, and a pre-approved, pre-written change that relaxes enforcement on a named scope. Make it site-scoped rather than estate-wide, time-boxed with a verified re-enable, logged loudly, and invocable by two named people — because a documented way to turn the control off is the most valuable thing you own, and it is exactly what an intruder inside the estate would use.
code
text · 10 linesPort 12 - supplicant starts EAP; switch relays to RADIUS
no reply, retransmits exhausted, all servers tried
-> servers marked dead -> critical VLAN applied, port forwards
Port 14 - supplicant starts EAP; switch relays to RADIUS
supplicant validates server certificate -> expired -> aborts the TLS handshake
RADIUS answers: Access-Reject
-> an answer arrived, so dead-server logic never runs
-> port stays unauthorized (or auth-fail VLAN, if one is configured)
...go deeper
Know that in certificate-based 802.1X the device checks the server's certificate too, so an expiry there can stop authentication everywhere at once even though the servers are healthy.
Explain that a returned reject and an unanswered request take different paths in the switch, and that dead-server fallback keys on the absence of an answer rather than on failure in general.
Show the recovery thinking: an out-of-band path to the switches, a pre-approved scoped relaxation, and awareness that the servers, jump host and patch source all sit behind the control that is denying.
Own the position that a shared trust anchor is a single point of estate-wide denial, and that the break-glass capability itself needs an owner, an approval path, a time box and evidence that enforcement came back.
## Why this failure is a class of its own Most 802.1X outages are reachability outages: a circuit, a server, a route. Estates design for those, and the design is the critical VLAN. A server certificate expiry is a different animal, and it defeats that design in a way that surprises competent senior engineers. In EAP-TLS and the tunnelled EAP methods, the supplicant establishes a TLS session with the authentication server and validates the server's certificate before proceeding — that validation is the entire reason the method resists a rogue authentication server. When the certificate expires, every supplicant that actually checks it stops at the same point. The failure is therefore **simultaneous across the fleet**, because it is one object expiring, not thousands of independent events. There is no gradual signal, no canary site, no partial degradation to notice at 20% adoption. And the servers are fine. They are up, reachable, responsive; the exchange simply fails and a reject is returned. From the switch's point of view an answer arrived. Retransmit and timeout logic never fires, the server is never marked dead, and the fallback built for unreachability sits there unused. Ports go to the auth-fail treatment if one exists and to unauthorized if it does not. The interview value is precisely this: **the fallback you designed is scoped to one failure mode, and the fleet-wide failure mode is a different one.** A candidate who says 'the critical VLAN covers us' has revealed that they think of the fallback as a general safety net. ## The circular dependency Now count what you cannot reach. The authentication servers are behind switch ports. So is the jump host. So is the patch source, the configuration management system, and the laptop of the engineer who would fix it. At a national estate of unstaffed sites, the nearest human may be hours away, and there is no seat at the site to sit in anyway. A recovery plan that begins 'connect to the management network' has already failed. So the break-glass path must be built on something that does not depend on the control: - **Out-of-band reach to the switch itself** — a console server or cellular link, on a path that is not 802.1X-enforced, at least at aggregation points. - **A management path for the authentication tier** that never sits behind its own decision. - **A pre-written, pre-approved change** that relaxes enforcement on a named scope. Writing this at 03:00, under pressure, on a bridge, is how estates end up disabling enforcement globally and leaving it that way for a fortnight. ## Designing the break-glass so it is not the gift A documented, tested method for turning admission control off is enormously valuable — to you, and to anyone already inside your estate. Treat it as a privileged capability in its own right: | Property | Why it matters | |---|---| | Scoped to named sites or switches | Estate-wide relaxation converts a regional outage into a total one and is rarely walked back quickly | | Time-boxed with automatic revert | Prevents the fortnight of silently disabled enforcement that follows every incident | | Two named invokers, recorded | Makes misuse attributable and stops a single stolen account from being enough | | Loudly alarmed, not just logged | The signal must reach somebody outside the bridge, so an unauthorised invocation is visible | | Verified re-enable | Someone must prove enforcement is back on, per site, before the incident is closed | Notice that the last row is where estates actually get hurt. The outage ends, everyone goes to bed, and nobody confirms that all four hundred sites came back into enforcement. Months later a routine audit finds a region running open — and no adversary had to do anything at all. ## Prevention, stated honestly The durable fix is not clever fallback engineering. It is refusing to let a single object's lifetime become a fleet-wide event: know the expiry, hold a second trusted issuer or a staged replacement so supplicants will accept a successor before the incumbent dies, and rehearse the swap. The mechanics of issuing and rotating that certificate belong to your PKI practice; what belongs to network defence is recognising that a single shared trust anchor is a **single point of estate-wide denial**, and that the failure it produces is invisible to every timer you tuned. ## What a strong answer covers Name why the fallback does not engage (an answer arrived), name the simultaneity (one object, whole fleet), name the circular dependency (the fix path is behind the control), and then describe a break-glass that is scoped, time-boxed, attributable and verified on the way back. If you also say who is allowed to invoke it and how you prove enforcement was restored, you have answered at the level the question is asked.
- How would you tell this apart from a circuit outage in the first five minutes of the bridge?Ask whether the servers are answering. In a reachability failure the switches report exhausted retries and unreachable servers, and the sites that lost a shared circuit fail together geographically. In this failure the servers show a flood of rejects with the exchange failing at the same stage, and the affected devices are the ones that validate the server certificate — spread across every region at once.
- Would devices that skip server-certificate validation stay online?They would, and that is not a comfort. A supplicant configured not to validate the server certificate is one that will also accept a rogue authentication server, which is the attack the method exists to stop. Finding that half your estate rode through the outage tells you half your estate is not really protected.
- Why is estate-wide relaxation the wrong shape for break-glass?Because it turns a bounded incident into total exposure and it is almost never reverted at the same speed it was applied. Scope the relaxation to named sites, expire it automatically, and require someone to verify enforcement is back on site by site before the incident is closed.
saying these in an interview costs you the question
- Expects the critical VLAN to cover a certificate expiry
- Treats an expiry as a gradual failure rather than a simultaneous one
- Plans recovery through a management network that is itself enforced
- Disables enforcement estate-wide with no time box
- Never verifies that enforcement was restored per site