skip to content

Your production cloud account keeps a break-glass login for emergencies — what must be true of it before it is safe to keep?

level: seniorimportance: should knowfreq 45%

answer

  1. for the day the normal login fails
  2. must not depend on the directory
  3. approval before, alarm during
  4. expires on its own, rotated after
  5. rehearsed, or it is only a rumour

basics

~20 s

A break-glass login is safe to keep only when it cannot be used quietly: a named trigger, an approval by a second person, a session that expires on its own, an alarm nobody can suppress, rotation afterwards, and a rehearsal recent enough to prove it still works.

solid answer

~50 s

It has to be impossible to use quietly and impossible to use for convenience. Concretely: a written trigger that says which failures justify it; a second person approving the opening; a session that **expires by itself** rather than relying on someone to sign out; an alarm on any use at all, routed to a human who is on the hook right now; rotation of the secret and its second factor afterwards; and a review of everything the session did. One property is easy to miss and is the one that decides whether it works: the break-glass path **must not depend on the company directory**, because a directory or federation failure is one of the emergencies it exists to survive. And it must be rehearsed — an untested emergency path is a rumour, and the outage is a bad time to discover the second factor expired.

code

pseudocode · 15 lines
pseudocode
on breakGlassRequest(engineer, incidentRef):
    if directory.isReachable() and normalLogin.worksFor(engineer):
        deny('no emergency: use the federated login')

    approval = secondPerson.approve(engineer, incidentRef)
    alarm.page(channel = 'security-oncall', actor = engineer, incident = incidentRef)

    secret  = offlineSafe.open(emergencyIdentity, approvals = [approval])
    session = platform.signIn(secret, expiresInMinutes = 60)
    audit.record(engineer, incidentRef, session.id, session.expiresAt)
    return session

on sessionExpired(session):
    offlineSafe.reseal(emergencyIdentity, rotateSecret = true, rotateSecondFactor = true)
    review.open(session.id, evidence = audit.callsMadeBy(session.id), dueWithinDays = 1)

go deeper

for a junior

Know what break-glass means: a rarely used emergency way into an account, deliberately made hard to use quietly, for the case where the normal login route is unavailable.

for a middle

Be able to list the properties and say why each exists — trigger, approval, self-expiring session, unsuppressible alarm, rotation afterwards — rather than describing a spare administrator password.

for a senior

Demonstrate the design instinct: the path must not depend on the directory it exists to survive without, and it must be rehearsed, because an untested emergency path fails silently until it is needed.

for a principal

You own the trade-off between friction and reachability. Every control here slows an incident, so decide deliberately which are non-negotiable and make the drill the mechanism that keeps the rest honest.

## What break-glass is for Break-glass is the way in when the normal way in is gone. The normal way in for a workforce is a federated session from the company directory; the emergencies that justify a second path are the ones where that fails — the directory is unreachable, the federation trust is broken or misconfigured, a change has removed everyone's administrator access, or the one person who could act is unreachable and something is on fire. It is not a second administrator login kept around because permissions are slow to request. That version is simply a standing credential with an exciting name, and it carries all of the risk of the real thing and none of the control. ## The properties that make it safe to keep 1. **A written trigger.** Which failures justify opening it, and which do not. Without this the path drifts into 'the fast way to get administrator' and stops being exceptional. 2. **Approval by a second person.** One human should not be able to become the account alone. Two-person opening turns a bad decision or a compromised individual into a conversation. 3. **A session that expires by itself.** Time-boxing must be a property of the credential, not a promise from the person using it. Verify this rather than assume it — a path that mints a session someone has to remember to end will eventually leave one open. 4. **An alarm nobody can suppress.** Not a log line; a page to a human who is accountable in the moment. The alarm is *detective* — it does not stop the use, it guarantees the use is seen — and describing it as if it prevented anything is a misunderstanding worth avoiding in an interview. 5. **Rotation afterwards.** The secret and its second factor are both replaced and resealed, so whoever saw them during the incident cannot reuse them. 6. **A review of what the session did.** Read the record of management calls made under it and ask one question: what did this do that a normal grant could not have done? If the answer is 'nothing', the trigger needs tightening. ## The property that is usually missed **The break-glass path must not depend on the directory it exists to survive without.** A break-glass identity that signs in through the company directory is fine on every day when it is not needed and useless on the day it is. This is the single most useful thing to say in an interview about it, because it is the one that separates people who designed a path from people who inherited one. The same reasoning extends outward: if the approval step runs through a chat tool that shares a dependency with the thing that is down, if the sealed secret is behind a system that needs the same login, or if the alarm route depends on the failed component, the path has a hole in the same shape. | | A standing second administrator login | A governed break-glass path | |---|---|---| | Who can use it | anyone who knows it | one person, after a second approves | | When | whenever it is convenient | only under a written trigger | | How long | until somebody signs out | until the session's own expiry | | Who finds out | the record, if anyone reads it | a human, immediately | | After use | nothing | rotate, reseal, review | ## Why the rehearsal is the part that fails Everything above is a design, and designs rot silently. The second factor expires. The person holding half the secret changes team. The recovery contact points at a mailbox nobody reads. The alarm route was renamed during a migration and now goes nowhere. None of these produce a symptom until the one moment when the path is needed, which is the worst possible time to find out. So rehearse it: on a schedule, against an account of the same shape, with the alarm live and the session actually used, and at least sometimes against production itself with everyone told in advance. Record how long it took, what blocked, and rotate afterwards exactly as a real use would. A drill that produced no friction usually means the drill skipped a step. ## What an interviewer listens for The weak answer stops at 'we have a break-glass account with a strong password in a safe'. The strong answer names the properties, gets the alarm's role right (it detects, it does not prevent), and says the thing about the directory dependency unprompted. The very strong answer adds what it costs: every property here is friction during an incident, and a team that has never rehearsed will quietly route around all of it at the worst moment.

  • Why must the break-glass path avoid the company directory?
    Because a directory or federation failure is one of the emergencies it exists to survive. A path that authenticates through the component that is down is not a path. The same test applies to every step around it: the approval channel, the place the sealed secret is kept and the alarm route must not share the dependency that failed.
  • How do you rehearse it without weakening it?
    Drill it on a schedule against an account of the same shape, with the alarm live and the session genuinely used, and occasionally against production with the team told in advance. Treat the drill exactly like a real use: approval, alarm, expiry, rotation, review. Record the elapsed time and whatever blocked, because that is the list of things that would have blocked you for real.
  • The alarm fires every time, and during a long incident people find it noisy. What do you do?
    Do not let it be suppressed — the alarm is the whole reason the path is safe to keep. Fix the noise instead: one page per opening rather than per call made, the incident reference carried in the alarm so it is obviously correlated, and an acknowledgement that stops repeats for the life of that session only.

The sealed handle behind glass by a fire door: anyone may reach it in a real emergency, nobody can reach it quietly, and somebody walks past regularly to check the glass is still there and the handle still moves.

saying these in an interview costs you the question

  • Break-glass is just an administrator login kept around for emergencies
  • The alarm can be muted during a long incident because it is noisy
  • It is fine for break-glass to sign in through the company directory
  • Nobody has used it in two years, so it clearly still works
  • Time-boxing the session is impractical when a real incident is running
  • Any on-call engineer may open it without telling anyone, since it is urgent