Always-on VPN lockdown leaves a traveller with no network at all. What recovery path do you sign for?
answer
- The fix needs the network the lockdown denies
- Pick a posture before it is needed
- The unlock is a supported bypass
- Verify out of band, expire in minutes
- Track the rate and fix the root cause
basics
~10 sChoose one and name its owner: hard lockdown plus an out-of-band, time-limited, per-machine unlock the helpdesk logs, or a bounded automatic fail-open. The unlock is what a caller claiming to be stranded will target.
solid answer
~40 sThe trap is circular: the client will not come up, and the fix needs the network the lockdown denies - including the path to the helpdesk. So pick a posture in advance. Fail-closed with an out-of-band unlock keeps the property you bought, at the price of staffed recovery across time zones and a stranded user until someone answers. Bounded fail-open keeps people working, at the price of a documented untunnelled window on an unknown network. Neither is free, and it is a risk acceptance with a named signatory, because it gets tested at 23:00 by an executive. Whichever you pick, harden the recovery: the unlock is per-machine, single-use, expires in minutes, requires callback verification, and every issuance is logged. Then track the unlock rate - a rising one usually means a fixable root cause.
go deeper
Know why a locked-down machine is a support problem at all: with no network there is no way to reach help or download a fix, so recovery has to be arranged in advance.
Explain the two postures and the mechanics of a safe unlock — bound to the machine, single-use, short-lived, verified out of band and logged.
Show that you reduce the need for recovery as well as designing it: transport choice, multiple gateways, certificate renewal windows and staged client updates, backed by a measured lockout rate.
Own the acceptance itself: state each option's cost in the signatory's currency, get a named owner on the posture with a review date, and be ready for the meeting an angry stranded executive triggers.
## Why this is a decision rather than a setting An always-on client fails for ordinary reasons: an expired device certificate, a client update that did not land cleanly, a gateway or path outage, a network that will not carry the tunnel's port or protocol. In lockdown, all of these have the same symptom — the machine has no network at all. It cannot reach the helpdesk, cannot download the fix, and cannot even complete the portal sign-in that would have given it a path to either. The recovery has to be designed before it is needed, because at the moment it is needed there is no channel to design it over. ## The two postures and what each actually costs | Posture | What you keep | What you pay | | --- | --- | --- | | Hard fail-closed with out-of-band unlock | the single-path property, and an evidence trail | staffed recovery across time zones, a stranded user until it answers, and a bypass mechanism that now exists | | Bounded automatic fail-open | users keep working, and the helpdesk queue stays small | a documented untunnelled window on an unknown network, with nothing inspecting it | The second is not a failure of nerve; it is a defensible choice if it is bounded, alerted, and paired with a reduction in what the machine can do while it is open. What is not defensible is an unbounded fail-open nobody wrote down, which is what you get by default when a client is configured permissively and never tested. ## Hardening the recovery path, because it is the new target An unlock mechanism is, by construction, a supported way to switch the control off. Treat it as such: - **Bind it to the machine.** A code derived from a device identifier is useless on any other endpoint. - **Expire it in minutes**, not for the trip, and make it single-use. - **Verify the human out of band.** A callback to the number of record, a manager confirmation, or an authenticator prompt to a registered device — the phone has its own connectivity, which is precisely why it works here. `The caller sounded senior and was in a hurry` is the pressure the process exists to withstand, and social engineering against a stranded-traveller story is the obvious play against this control. - **Reduce the blast radius while it is open.** Some estates pair an unlock with a session that has fewer rights, or with a forced re-authentication when the tunnel returns. - **Log every issuance** with requester, approver, machine, reason and duration. This is the evidence you will be asked for, and it is also your input data. ## Reduce how often you need it at all Most of the argument disappears if the client comes up in places it currently does not: - offer the tunnel over a port and transport that hostile-to-VPN networks are least likely to block, and confirm it works from real venues rather than from the office; - run more than one gateway so a single outage is not an estate-wide lockout; - set device certificate lifetimes and renewal windows so renewal never depends on being connected at the moment of expiry; - stage client updates, because a bad build here does not degrade the fleet, it disconnects it. ## The organisational part, which is the actual question Who signs? The engineer configuring the client cannot accept, on the business's behalf, the risk of a documented untunnelled window, nor can they alone commit the budget for overnight helpdesk cover. The output of this decision should be a short written risk acceptance with a named owner, a stated posture, a bounded exception mechanism, and a review date. Bring three things to that conversation: 1. the measured lockout rate per thousand devices per week, with causes attributed; 2. the cost of each option stated in the currency the signatory cares about — staffed hours for fail-closed, exposed device-minutes on unknown networks for fail-open; 3. a commitment to revisit at a threshold, so the decision is not permanent by default. And expect the conversation to be triggered rather than scheduled. The realistic sequence is that an executive is stranded once, demands the control be removed entirely, and you have roughly one meeting to convert that into a bounded mechanism instead of a blanket exemption. Having the numbers and the drafted posture ready before that meeting is the whole job at this level. ## The number that ends the argument Track unlocks and fail-open events per thousand devices per week, with a cause on each. A high and rising rate is rarely evidence that users are undisciplined; it usually points at one fixable root cause — a certificate lifetime, a transport that venues block, a client build. Fixing that costs less than either posture and removes the pressure on the bypass, which is the outcome you actually wanted.
- What verification do you require before a helpdesk agent issues an unlock code?Out-of-band proof that is hard to arrange under time pressure: a callback to the number of record, an authenticator prompt to the registered device, or a manager confirmation for anything unusual. The code should be bound to that machine, single-use and expiring in minutes, and every issuance logged with requester, approver and reason. Assume a caller with a convincing stranded-traveller story is exactly the case the process must survive.
- Why not simply fail open after five minutes of failed connection attempts?You can, but only bounded, alerted and written down. Unbounded fail-open means the estate quietly loses the property you bought whenever a gateway has a bad day, and the same conditions that trigger it can often be produced by the network the machine is on. If you choose it, cap the duration, raise an event, reduce what the machine may do meanwhile, and get the acceptance signed by a named owner.
- What signal tells you the posture is wrong rather than the users being careless?The unlock and fail-open rate per thousand devices per week, with an attributed cause on each event. A cluster on one cause — a certificate lifetime, a transport blocked by venue networks, a bad client build — means the fix is engineering, not process. Fixing it is cheaper than either posture and removes the pressure on the bypass path, which is where the real risk sits.
saying these in an interview costs you the question
- Leaves the failure posture undecided until an incident forces it
- Issues unlock codes on the caller's word alone
- Makes the unlock reusable or valid for the whole trip
- Treats an unbounded fail-open as an acceptable default
- Accepts the residual risk personally instead of getting an owner to sign
- Never measures how often the recovery path is used or why