skip to content

Finance wants to cut broker headroom that is only used when a stolen-credential incident forces a mass revocation — what do you argue, and what disconnect time do you sign for?

level: principalimportance: nice to knowfreq 26%

answer

  1. not capacity — the ability to use the control
  2. everyone reconnects in the same minute
  3. the failure mode is a revocation nobody runs
  4. offer stagger, burst, scoped predicates
  5. commit a share and a window, exclude offline

basics

~20 s

Argue that the headroom buys the ability to actually fire a mass revocation: without it, the reconnect wave takes the broker down and the response becomes one nobody dares run. Then commit only to a measured, scoped number, with offline clients excluded.

solid answer

~60 s

Do not defend the headroom as capacity; defend it as the thing that makes a containment action usable. When a credential compromise forces you to revoke broadly across a fully remote workforce, every disconnected client re-authenticates within the same minute — against the broker and the identity provider at once — and the service desk absorbs the calls. If that wave exceeds what the platform can serve, the outcome is not a slow reconnect; it is that the incident commander is told the revocation would cause an outage and chooses a narrower, slower response. That is the real cost line, and it is the one finance can price. Then offer the cheaper alternatives honestly: staggered re-admission so the peak is spread, burst capacity that is only paid for when triggered, and revocation predicates scoped so you can cut a cohort rather than everyone. And commit to a number you can evidence — a stated share of sessions dead within a stated window, measured by a scheduled exercise, with unreachable clients excluded in writing.

go deeper

for a junior

Understand that cutting everyone off at once creates a reconnect wave, and that the systems people log back in through have to be able to serve it.

for a middle

Explain why simultaneous re-authentication concentrates load on the broker and the identity provider at the same moment, and what staggering re-admission changes.

for a senior

Show you would size and test the reconnect path deliberately, and produce a measured disconnect figure rather than asserting one.

for a principal

Own the argument that this budget buys a usable containment action, name the cheaper alternatives, commit only to an evidenced and scoped number, and get the service desk and incident owner to sign alongside you.

## Reframe the line item The request on the table is "why are we paying for capacity that sits idle". Answering with utilisation graphs loses, because the graphs prove finance's point: on an ordinary Tuesday the headroom is unused. The honest framing is different — this capacity is not for serving traffic, it is for **surviving the response**. The scenario it exists for: a credential compromise forces a broad revocation. Across a workforce with no office at all, every one of those users is on home broadband, hotel wifi or a tethered phone, and every application they use is reached through the same broker. Cut them all, and they all come back inside the same minute — a simultaneous re-authentication wave against the broker's front end and the identity provider behind it, plus a support call from every person whose work was interrupted. If the platform cannot serve that minute, re-admission stalls, the wave retries, and you have converted a security action into an outage. ## The cost finance has not been shown What gets bought with the headroom is not comfort — it is the willingness to use the control. The failure mode of an under-provisioned broker is not a slow morning; it is a conversation during an incident where the security lead says "we should revoke everyone" and the platform owner says "that will take us down", and the organisation settles for a narrower revocation that leaves the adversary a session. The cost of that outcome is the cost of the incident continuing, and that is a number the business already knows how to reason about. So the argument is: *this line item is what makes our stated containment action executable. Remove it and we should also remove the containment action from the incident plan, because we will not run it.* Whoever cuts the budget is then explicitly accepting that. ## Offer the cheaper options, because they exist An argument that only says "pay more" is weak. Put alternatives on the table and let the owner choose: - **Spread the peak.** Re-admission with jitter turns one impossible minute into several manageable ones. It costs recovery time, not money — but that recovery time is the thing you would otherwise be committing to. - **Burst rather than standing capacity.** Front-end capacity that is provisioned when the revocation fires costs far less than permanent N+1, at the price of a warm-up delay and a dependency on that expansion working under stress — which means it must be exercised, or it is a plan rather than a capability. - **Scope the predicate.** If revocation can cut a cohort — one region, one device class, one application tier — you rarely need the whole-population wave at all. This is the cheapest option and the one that needs the most design work up front. - **Prioritise re-admission.** Bring back the flows the business cannot run without before the rest. This requires somebody outside security to say which ones those are. ## Signing for a number The second half of the question is the harder one, because a committed number is held against you. Commit only to what you can evidence, and state the exclusions plainly: - A **stated share** of live sessions terminated within a **stated window** of the revocation being issued — a share and a window, never "all sessions immediately". - Measured by a **scheduled exercise**, not by design intent: a synthetic revocation through the real path, timed to last byte. - **Devices that are offline or asleep are excluded**, explicitly, because you cannot disconnect what you cannot reach; the guarantee for those is that they cannot re-admit, not that their existing flow is dead. - A separate **re-admission** commitment, because the business cares more about being back than about being cut off, and it is a different number with different owners. ## Who actually signs This is what makes it a leadership question rather than a capacity question. The broker owner commits the disconnect window. The service-desk owner commits the staffing to absorb the calls, or the window is fiction. The incident owner accepts the residual — the offline tail and the traffic no enforcement point reaches. And finance owns whichever capacity model was chosen. If any of those four has not signed, you do not have a commitment; you have an intention that will be re-litigated in the middle of an incident, which is the worst possible time to discover it.

  • How would you evidence the disconnect number you signed for?
    A scheduled exercise, not a design document: a synthetic principal admitted through the real path, revoked, and timed to last byte at a point the endpoint does not control, repeated across awkward conditions — mobile tether, recently woken device. Report the distribution and its tail. If the exercise has not been run this quarter, the commitment is stale and should be described as such.
  • What do you tell an executive who asks for a guarantee that every session is dead within a minute?
    That you can guarantee no further admission almost immediately, and that a stated share of live sessions ends within a measured window — but not all of them, because a device that is offline or asleep cannot be reached, and some traffic never traverses an enforcement point at all. Offering a guarantee you cannot evidence costs you more the first time it is tested.
  • If the budget is refused outright, what changes in the incident plan?
    The plan must stop naming whole-population revocation as an available action, and say what replaces it: scoped cohort revocation, a longer staged rollout, or accepting a longer exposure window. Write that down and have the incident owner accept it. The unacceptable outcome is a plan that still lists an action the platform cannot survive.

saying these in an interview costs you the question

  • Defends the headroom with average utilisation figures
  • Commits to a disconnect time never measured in an exercise
  • Promises every session dies immediately, offline devices included
  • Ignores the support-desk load in the same minute
  • Offers no cheaper alternative to permanent standing capacity

context