skip to content

Your break-glass route is opened for the first time during an outage and fails - what should a rehearsal have proven about its dependencies?

level: seniorimportance: should knowfreq 42%

answer

  1. the route must not need what failed
  2. circular dependency, found at the worst time
  3. identity, holder, alert path, network
  4. never opened is never proven
  5. rehearse, then replace what was read

basics

~20 s

A rehearsal has to prove the route opens while the failing system is unavailable: the identity used to open it, the place the emergency value is held, the alert path and the operator's own access must not depend on what is down.

solid answer

~40 s

A break-glass route is a system, not a document, and it has dependencies like any other. The classic failure is circular: the emergency credential is kept in the store the outage took out, or the route is opened by signing in to the identity provider that is the reason nobody can sign in, or the alert goes to a channel hosted on the platform that is down. A rehearsal is what finds these, and only a rehearsal that **actually opens the route** finds them — reading the runbook aloud proves the runbook is readable. Rehearse on a schedule, with whoever would really be on call rather than the person who wrote it, with the dependency genuinely removed, and finish the way a real use finishes: replace what was read and file the same review.

code

pseudocode · 20 lines
pseudocode
open_break_glass(engineer, reason, now):

    # dependency 1: proving who is asking
    identity = identity_provider.verify(engineer.second_factor)
    if identity == none:
        abort("the provider that proves identity is part of this outage")

    # dependency 2: where the emergency value is held
    value = emergency_holder.read("payments/emergency/database-owner")
    if value == none:
        abort("the emergency value lives in the store that is down")

    # dependency 3: somebody other than the opener seeing this
    delivered = alert.send(channel = "security",
                           who = identity, reason = reason,
                           expiresAt = now + 60m)
    if not delivered:
        abort("this use would be invisible")   # a design choice, not a law

    return grant(value, expiresAt = now + 60m)

go deeper

for a junior

Remember the core point: a route for emergencies has to be tried before the emergency, and trying it means actually taking the credential, not reading the instructions for taking it.

for a middle

Explain the dependency chain: identity, second factor, where the value is held, the authorising rule, the alert channel, the network path. Name one way each can be part of the same outage.

for a senior

Show that you have run one. Describe rehearsing with the dependency removed and the likely on-call person rather than the author, the time-to-open you measured, and the replacement and review you did afterwards.

for a principal

Decide what the estate stakes on this route: which failure domains it must survive, whether an offline copy is worth its custody cost, and whether the route fails closed when its use cannot be observed.

## The route is a system, not a document Every break-glass route has a chain of things that must all work at the moment it is used, and that chain is almost never written down next to the steps. The chain typically includes the identity the opener authenticates with, the second factor that identity needs, the place the emergency value is held, the rule that says this identity may open this route, the channel the use alert is published to, the network path from wherever the engineer happens to be at 02:00, and the document that says which route to open at all. The failure mode this leaf exists for is that **one or more links in that chain is the very system the route was built to work around**. It passes every desk check, because on a normal day every link is up. ## The circular dependencies that surface at 02:00 | Dependency | How it fails during a broad outage | What independence looks like | |---|---|---| | Where the emergency value is held | It is held in the same store the outage took out | A holder on a separate failure domain, or a sealed offline copy under two-person custody | | The identity used to open it | The route is opened by signing in to the provider that is down | A local account on the emergency path, exercised often enough to still work | | The rule authorising the open | The rule is evaluated by the service that is unreachable | The authorisation for the route is pre-resolved and cached where the route lives | | The alert channel | The channel runs on the platform that is failing | A second, boring channel that has no dependency on the estate | | The document | The runbook lives behind the sign-in that is broken | An offline copy, printed or synced, with the date it was last proven | | The operator's network path | The only path in is through the thing that is down | A second path, exercised as part of the rehearsal, not discovered during it | The honest framing for a candidate: you cannot make a route depend on nothing. You can make it depend on a **different** set of things than the estate it rescues, and you can write that set down so the next person can check it. ## What a rehearsal actually has to do 1. **Schedule it.** A route proven once at build time decays: people leave, identities change, the holder moves, the alert channel is renamed. 2. **Open it for real.** Take the credential, all the way to the point of using it. The steps that fail are almost always past the point a dry run stops. 3. **Use the right person.** The author knows the unwritten steps. Run it with someone who would plausibly be on call and has not done it before. 4. **Remove the dependency.** Rehearse with the system the route rescues genuinely unavailable, or as close to that as you can stage. A rehearsal performed while everything is healthy proves only that the route works when it is not needed. 5. **Measure the time to open.** How long from "we need it" to "we have it" is the number that decides whether people use the route or keep a private copy instead. 6. **Finish the way a real use finishes.** Replace what was read, let the window close, file the same review. A rehearsal that skips the ending never proves the ending works either. ## The offline copy, and what it costs Sometimes the only genuinely independent holder is an offline one: the value written down, sealed, and stored physically. It is defensible precisely where every online path can be part of the same failure. It is not free. A sealed copy needs known custody, two people to open it, a tamper-evident seal, a place in the inventory so nobody forgets it exists, and replacement of the value after **every** opening including rehearsals — otherwise you have created a long-lived shared secret whose holders are whoever has been in the room since it was written. ## Failing to open, on purpose Designs differ on what happens when part of the chain is unavailable. Some routes refuse to hand over the credential if the use alert cannot be delivered, on the grounds that unobserved emergency access is the thing you were trying to avoid. Others hand it over and record locally, accepting that observation is best-effort so that the route never becomes the outage. Either is defensible; what is not defensible is discovering which one you built while the payments service is down. ## What an interviewer is listening for That you treat the route as a system with a dependency list, that you name at least one genuine circular dependency, and that your definition of rehearsal includes actually opening it and cleaning up afterwards.

  • What must a rehearsal of a break-glass route do afterwards that an ordinary drill does not?
    Finish like a real use. The value that was read is now known to whoever rehearsed, so it is replaced; the window is allowed to close on its own clock so the expiry path is proven too; and the same post-use review is filed, which is how you find out whether the review process itself works. A rehearsal that reads a credential and leaves it in place has quietly widened the set of people holding it.
  • How often should a route be rehearsed, and what else should trigger one?
    On a fixed cadence slow enough to be sustainable and fast enough to catch decay - many teams settle on quarterly - plus an event-driven rehearsal whenever a link in its chain changes: the holder moves, the identity provider changes, the alert channel is replaced, the pre-authorized list changes, or the estate the route rescues is re-platformed. The trigger list matters more than the cadence.

saying these in an interview costs you the question

  • The emergency credential lives in the store, which is always available
  • We wrote the runbook, so the route counts as tested
  • We rehearse by reading the document, not by opening the route
  • Opening it for practice is risky, so we never actually do it
  • It worked at the last rehearsal, so it still works now