An 802.1X port loses its RADIUS servers: what decides how long it waits, and what happens to devices the fallback admitted?
answer
- retransmits times timeout times servers
- the wait is the outage
- short timers make blips into bypasses
- coming back does not undo it
- everything re-authenticates at once
basics
~20 sThe wait is the per-server retransmit count times the timeout, repeated across every configured server, so the port sits dark for that whole window before any fallback applies. Ports the fallback admitted keep that access until something forces re-authentication, and often nothing does.
solid answer
~50 sDead-server detection is arithmetic you configure: each authentication request is retransmitted a set number of times at a set timeout, and the switch tries the servers in order before declaring them all unreachable. The total is the time a device at an unstaffed site is simply offline, and it is paid on every transient WAN blip. Tune the timers down and a two-second circuit wobble drops the whole estate into the fallback; tune them up and a real outage keeps the site dark for minutes longer. Once the fallback applies — a critical VLAN, sometimes called inaccessible-authentication-bypass — the port forwards. The part teams get wrong is recovery: those sessions are not re-evaluated just because the servers came back. Unless you configure re-authentication on recovery, a port that was admitted during the outage stays admitted indefinitely, including the one someone plugged a laptop into. And if you do configure it, the whole estate re-authenticates at once and can flatten the servers you just restored, so stagger it.
go deeper
Know that the switch waits through retransmits and timeouts before deciding a server is unreachable, and that whatever happens next was configured by someone rather than chosen by the protocol.
Be able to compute the wait from retransmit count, timeout and server list, explain what a critical VLAN does, and state that recovery does not automatically re-evaluate ports already admitted.
Demonstrate the trade in both directions — short timers turn circuit wobble into an attacker-triggerable bypass, long timers extend a dark site — and show how you audit the port list an outage produces.
Frame timer and fallback settings as a risk position the business is buying, with the frequency of triggering, the reach of the fallback and the audit obligation stated together rather than tuned by an engineer in isolation.
## What the wait is made of A switch does not know a server is dead; it infers it. Each authentication request is sent, and if nothing comes back it is retransmitted a configured number of times at a configured timeout. When those attempts are exhausted the switch moves to the next configured server and repeats. Only when every server in the list has failed does it declare the authentication service unreachable and apply whatever fallback exists. Many platforms then hold that server in a dead state for a fixed dead-time before probing again, so a recovered server is not used the instant it returns. The practical number is roughly *retransmits times timeout times number of servers*, plus the supplicant's own patience. That number is not a detail — **it is the amount of time a device at a site with nobody in it is offline**, and you pay it on every authentication attempt during the outage, not once. ## The tuning trade you are actually making | Timers tuned short | Timers tuned long | |---|---| | A brief WAN wobble drops the estate into the fallback | Transient blips ride through with no fallback | | Genuine outages degrade quickly and predictably | A real outage keeps sites dark for minutes before anything opens | | The fallback triggers often, so it must be safe to trigger | The fallback triggers rarely, so it is rarely tested | There is no correct setting, only a defended one. If your fallback is a tightly scoped VLAN that is alarmed and audited, short timers are cheap. If the fallback is generous, every blip is a small breach, and you should make it hard to reach. A second-order effect: an adversary who wants the fallback does not need a long outage if your timers are short. Inducing a few seconds of loss between a site and the central servers is a low-skill act against a rural circuit, and it produces an open port on demand. Short timers convert a fragile circuit into an attacker-controlled bypass. ## The fallback itself When the servers are unreachable, the two honest options are: leave the port unauthorized, or place it in a critical VLAN. The second is only defensible if you decided in advance what that VLAN can reach. A fallback that lands devices in the production data VLAN has not degraded gracefully — it has disabled the control. A fallback that reaches only the local operational systems the site needs to keep running, with no route to the management network, has kept the site alive at a bounded cost. Also decide whether the fallback applies to *new* authentications only, or whether devices that fail re-authentication during the outage are also caught. A re-authentication that cannot reach the server should not tear down a working session — that turns a server outage into a site outage even where you meant to fail open. ## Recovery is the half that gets forgotten When the servers come back, ports already sitting in the critical VLAN are **not** automatically re-examined. They were authorised by a decision that never happened, and nothing revisits it. Two consequences follow. First, the security consequence: any device admitted during the window keeps that access for as long as its session lasts, which at an unstaffed site could be months. If an intruder used the outage, the fallback is now their persistent access, and the log entry that would have told you sits in the middle of a flood of identical entries from every legitimate device. That is why the fallback must be logged and alarmed *per port*, and why the list must be audited after every event — you need the small set of ports that took the fallback and never appeared in normal authentication records before or after. Second, the operational consequence: if you do configure re-initialisation on recovery, every port across the estate re-authenticates at nearly the same moment. The servers you just brought back now face the entire fleet at once, fail again, and the estate flaps back into the fallback. Stagger recovery — by site, by switch, or with a randomised delay — and bring the servers up with headroom before you release the fleet at them. ## What to say in an interview Give the arithmetic, name the trade in both directions, and then move to recovery unprompted: the fallback is a state that persists, the return of the server does not undo it, re-evaluation must be configured and staggered, and the port list from the outage is the artefact you go and audit. Candidates who stop at 'it falls into the critical VLAN' have described half a mechanism and none of its price.
- Should a re-authentication that cannot reach the server tear down a working session?Generally no. A session that is already up was authorised by a real decision; killing it because the server is unreachable converts a server outage into a total site outage even in a design meant to degrade gracefully. Let existing sessions persist, apply the fallback to new authentications, and record which sessions outlived their intended re-authentication.
- How do you find the ports that were admitted by the fallback once the outage is over?Log the fallback placement per port at the switch and forward it, so the outage produces a bounded list rather than a gap. Then subtract the ports that authenticate normally before and after the window; what remains is the set that only ever appeared during the failure, which is exactly where an opportunistic device would be.
- Why stagger re-authentication when the servers return?Because a synchronised estate-wide re-authentication is a self-inflicted flood against servers that have just come up cold with empty caches. They fail again, the fleet drops back into the fallback, and you have turned one outage into a loop. Randomised delays by site or switch spread the load into something the recovered tier can absorb.
saying these in an interview costs you the question
- Cannot say what determines the wait before a fallback applies
- Believes the fallback is revoked when the server returns
- Sets timers aggressively short without scoping the fallback
- Never logs which ports took the fallback
- Releases the whole estate at freshly recovered servers at once