skip to content

One shared credential is held by several hundred upload collectors and one host is compromised - what does withdrawing it cost?

level: middleimportance: must knowfreq 66%

answer

  1. count the holders, not the hosts
  2. one value, one switch
  3. withdrawal stops every holder at once
  4. containment equals fleet-wide outage
  5. per-holder identity makes it one host

basics

~20 s

Withdrawing a fleet-shared credential stops every holder at once, not just the compromised host: several hundred collectors fail together. Sharing traded per-host containment for one estate-wide switch, so containment and outage became the same action.

solid answer

~50 s

Blast radius is the set of holders and rights that one value's compromise reaches. With one shared upload credential that set is the whole fleet: the attacker on the breached host can do everything any collector can do, and the only control that acts on the value itself - withdrawing it - applies to all holders simultaneously. So containment costs a fleet-wide outage that lasts until every host has been reconfigured with a replacement, and that duration is set by how fast configuration reaches several hundred machines, not by how fast the destination can revoke. Had each collector held its own credential, the same withdrawal would have been one host's outage. Note the direction: withdrawing is what makes the old value stop working, putting a new value in place is a separate step, and doing only the second leaves the leaked one live.

go deeper

for a junior

Recall that a value many holders share can only be taken away from all of them at once. There is no per-host switch when there is no per-host value.

for a middle

Explain the count behind the answer: how many holders hold the value, what rights it carries, and why the cost of containment is the time to reconfigure every holder rather than the time to revoke.

for a senior

Show you have run it. Name the long pole - pushing the replacement - say which rights you would cut first because that needs no host change, and state the order you would withdraw and replace in.

for a principal

Argue where the fleet's identity boundary belongs and what the estate will pay in per-identity operations to make containment local instead of total. The trade is a wide radius against hundreds of enrollments.

## What blast radius actually counts **Blast radius** is the answer to one question: if this exact value gets out, what can its holder do, and to how many systems? Both factors are properties of the *credential*, not of the machine that was breached. - **Reach** - the set of operations and objects the destination accepts that value for. A value that may write, read back and delete has a wider reach than one that may only write, even though both are one string. - **Holders** - every place a copy currently rests: each of the several hundred collectors, the configuration that renders it onto them, any image or archive it was baked into, and the ticket or chat message where someone once pasted it. A compromise of one host is therefore not a one-host incident. The attacker inherits the fleet's reach, because the value they now hold is the value every collector holds. When you are asked for the blast radius of a credential, count in this order: 1. Every holder of the value, including the copies nobody is operating. 2. Every right the destination accepts it for. 3. Every system that accepts it at all - a value reused at a second destination doubles the count. 4. Everything downstream that trusts what those systems received, because poisoned uploads travel further than the credential does. ## Why withdrawal is all-or-nothing when the value is shared Withdrawal is an action against the *value*: the destination stops accepting it. There is no per-host handle to act on because there is no per-host value to act on. Everything that would have been three separable decisions under per-workload identity collapses into one switch. | | one shared credential | one credential per collector | |---|---|---| | Who stops on withdrawal | every holder | the one holder | | Cost of containment | fleet-wide outage until reconfiguration | one host offline | | Who the access record names | the shared value | the collector | | Time to contain | time to push new configuration everywhere | immediate | | Identities to operate | one | several hundred | The last row is the honest price of the right-hand column, and it is why fleets end up sharing in the first place. Nobody chooses a wide blast radius; they choose one credential to issue, one to configure and one to remember, and the radius is what that choice cost. ## The numbers that decide it Assume 400 collectors and a configuration push that reaches an individual host in about ten minutes but is staged across the fleet over roughly a day, because you will not restart four hundred data-collecting processes simultaneously. Then revocation takes a second and recovery takes a day. That asymmetry - not any property of the store - is what makes the decision to withdraw feel impossible during a live exposure, and it is the thing scoping is bought to avoid. State the assumption when you answer: if your fleet can be reconfigured in fifteen minutes, the same incident is a fifteen-minute outage and the argument changes. ## Four things that do not shrink this radius - **Rotating the shared value on a schedule.** The replacement is shared by the same four hundred holders, so the next incident reaches exactly as far. Rotation bounds how long a leaked value stays useful; it does not bound how many holders one withdrawal stops. - **Moving the value into a secret store.** Custody improves - fewer copies rest in configuration repositories, and reads can be recorded. The reach and the holder count are unchanged, because the same one value is still handed to every collector. - **Encrypting the store's contents at rest.** That protects the copy the store holds against a stolen disk or a stolen backup. It does nothing against a caller the store will answer, and nothing at all against a holder that already has the value. - **Asking the destination to block the bad host.** With one value it can only block by something other than the credential - a source address, for instance - which is a coarser and more easily evaded cut than refusing an identity. ## Withdrawing, replacing and the order between them These are two different actions and interviewers listen for whether you keep them apart. **Replacing** puts a new value in place; the old one keeps working until something stops it. **Withdrawing** makes the old one stop working; nothing works afterwards until the replacement has landed. Replace-then-withdraw avoids the outage but leaves the exposed value live for the length of the push. Withdraw-then-replace stops the attacker immediately and buys that with the outage. Neither is the default answer - the condition decides, and a candidate who states one as correct without naming the condition has skipped the whole question.

  • The fleet is on one shared credential and you cannot reconfigure all of it today - what narrows the radius without replacing the value?
    Narrow what the value is accepted for: drop read-back, delete and list rights at the destination so it may only write. Narrow where it is accepted from, if the destination can filter by source. Narrow how long it stays valid. Each cuts a dimension of the radius with a change at the destination and no host touched, which is why it lands in minutes. None of them restores attribution.
  • Does per-workload identity reduce blast radius if every workload identity carries the same rights?
    It reduces the containment radius and restores attribution: you withdraw one holder, and the record names which one. It does not reduce the reach of any single compromise - each collector can still do everything any collector can do. Scoping by holder and scoping by right are separate axes, and cutting one leaves the other at full width.
  • Your fleet is split across twenty sites. Does a credential per site help if you cannot manage one per host?
    Yes, proportionally. The radius drops from four hundred holders to about twenty, withdrawal costs one site's outage rather than the estate's, and the access record can at least name the site. Partial scoping is not failed scoping - it is a smaller number, bought for a twentieth of the identities a per-host cut would have required.

One master key cut for every contractor on site: when one goes missing the only move is to change the lock, and that shuts out every contractor at once - and then a locksmith has to visit every door before anyone works again. The door log records that the master key opened it, never whose hand held it.

saying these in an interview costs you the question

  • Counts the blast radius as one machine because one machine was breached
  • Assumes only the compromised host loses access when the value is withdrawn
  • Thinks rotating the shared value contains the breach without withdrawing the old one
  • Expects the destination to disable one host while the others keep the same value
  • Treats moving the same shared value into a store as a reduction in radius