skip to content

One credential is held by hundreds of collectors you cannot update at once — how do you sequence the emergency replacement?

level: principalimportance: should knowfreq 33%

answer

  1. population against update rate
  2. the window is the attacker's too
  3. decide who breaks in advance
  4. no way back to the old value
  5. finished means refused and reporting

basics

~20 s

Compute the cutover from the population and the update rate, then decide in advance who breaks. When the fleet is larger than the rate, no sequence gives both a dead old value and an unbroken fleet, and the window you buy is also the attacker's.

solid answer

~40 s

Start with arithmetic: holders divided by the rate you can actually update gives the window. Eight hundred collectors at fifty an hour is sixteen hours, so a graceful cutover is a proposal to leave the leaked value live until tomorrow afternoon. That framing forces the real choice. If the accepting system honours two values at once you get a window, but shrink it to what the fleet genuinely needs and measure it rather than guessing. If it does not, the cutover is hard and you decide who loses access first — ideally splitting by what the credential reaches, withdrawing immediately for the high-reach subset and running a window for the rest. Decide which holders will simply fail before you start, not while it is happening.

code

pseudocode · 18 lines
pseudocode
holders             = 800
holdersUpdatedPerHour = 50
cutoverHours        = holders / holdersUpdatedPerHour        // 16

// split only where the accepting system can refuse per holder; often it cannot
if acceptingSystem.canRefusePerHolder:
    highReach = holders.where(reach == HIGH)
    withdraw(oldValue, forHolders = highReach)               // breakage accepted here, now
    remaining = holders.without(highReach)
else:
    remaining = holders                                      // one value, one decision

for each h in remaining:
    give(h, freshValueFor(h))                                // one value per holder while you are here

withdraw(oldValue, at = acceptingSystems)                    // at completion OR at the deadline

finished = provenRefused(oldValue) and every(holders, reportingOn = itsOwnNewValue)

go deeper

for a junior

Recall that a fleet cannot be updated instantly, so an emergency replacement takes a window, and during that window the leaked value still works for everyone who has it.

for a middle

Explain the arithmetic — population divided by the achievable update rate — and what that number means for both the outage and the attacker's remaining access.

for a senior

Show that you decided who breaks before starting rather than during: which holders are withdrawn immediately, which are moved inside the window, and which are simply going to fail.

for a principal

The call is what you stake: hours of degraded service against hours of a live leaked credential, with no third option, plus a standing decision about which way the estate leans when nobody has time to debate it.

## Start with the arithmetic The conversation is undecidable until two numbers are on the table: how many holders present this credential, and how fast you can genuinely move them. Divide the first by the second and you have the cutover window. Eight hundred collectors at fifty an hour is **sixteen hours**. Say it that way out loud, because "we will roll it out gracefully" and "the leaked credential keeps working until tomorrow afternoon" are the same sentence and only one of them gets challenged. Both numbers need honesty. The update rate is not the theoretical one; it is the rate that includes devices asleep, on poor links, or behind someone who has to physically touch them. A window computed from the optimistic rate is a window you will blow through with the old value still live. ## Three shapes, and what each stakes | shape | possible when | what it stakes | |---|---|---| | hard cutover — withdraw, then move | the accepting system honours one value at a time, or reach is too wide to wait | the whole fleet is down for the window | | windowed cutover — move, then withdraw | two values can be honoured at once | the leaked value works for the whole window | | split — withdraw for the high-reach subset, window the rest | the accepting system can refuse per holder or per right | a partial outage, and a shorter window on what matters | The split is the one worth reaching for and the one most often unavailable — designs differ on whether acceptance can be refused for one holder rather than for the value as a whole. Check it early, because the answer changes the plan rather than decorating it, and do not assume the granularity exists because a previous estate had it. ## Deciding who breaks, before you start In every shape except the free-window fantasy, someone loses access. The decision is which holders, and it is a judgment call a lead owns rather than an operational detail: - **Order by the cost of failure.** Holders whose silence is tolerable for a few hours move last and are allowed to break; holders whose silence is not go first, or are exempted and carried by another route. - **Order by reach.** If some holders can do more with the credential than others, they are the ones the attacker most wants, and they are the first withdrawal regardless of convenience. - **Name the unreachable up front.** Every fleet has holders that will not be updated tonight. Whether they fail is a decision you record beforehand, not a surprise you discover at hour nine. - **Prove the new value on a small group first.** A handful of easy, noisy holders that report clearly are a cheap check that the replacement actually works before you commit six hundred more to it. ## The back-out is forward, not back Ordinary change management assumes a back-out to the previous state. There is none here. The old value is leaked, so restoring it hands the attacker their access back — and no amount of urgency makes that acceptable. If the replacement turns out to be broken, the back-out is to issue *another* new value and roll that, with the leaked one staying withdrawn throughout. Pausing the rollout and leaving both values honoured is the other tempting non-answer: it extends the exposure indefinitely, which is the one thing the whole exercise exists to end. ## What finished means Two conditions, both required: 1. an attempt with the old value is refused, at a recorded time, at every system that honours it; 2. every holder is reporting on the new value — reporting, not merely "not complaining", since a silent collector is indistinguishable from a broken one. A rollout script exiting zero satisfies neither. ## What the emergency should leave behind The replacement is the cheapest moment you will ever get to end the shared value: as each holder is touched anyway, give it one of its own. The next emergency then costs one holder rather than the fleet, and the arithmetic that dominated this night stops applying. That is the one lasting output of a night like this, and it is lost if the fleet is simply re-pointed at a new shared value and everyone goes back to bed.

  • How do you decide the order within the fleet?
    By two axes. First, what each holder can do with the credential: the highest-reach holders are withdrawn earliest, because they are what the attacker wants. Second, what each holder's failure costs: the tolerable ones absorb the outage, and the intolerable ones are exempted and carried another way. Prove the new value on a small, noisy group before committing the rest.
  • What tells you the replacement is finished, as opposed to the rollout being finished?
    Two things together: an attempt with the old value is refused at every system that honours it, at a recorded time, and every holder is actively reporting on the new value. A silent collector is indistinguishable from a broken one, so absence of complaints is not evidence. A rollout script exiting cleanly says the change was attempted, not that it took.

saying these in an interview costs you the question

  • Promises both no outage and an immediately dead value
  • Guesses the cutover window instead of computing it
  • Discovers which holders break during the withdrawal
  • Offers restoring the old value as the back-out
  • Calls it finished when the rollout script completes