skip to content

The fleet's shared upload credential is publicly exposed and reconfiguring several hundred collectors takes a day - do you withdraw it now?

level: principalimportance: should knowfreq 37%

answer

  1. price both sides, then choose
  2. narrow the rights first
  3. buffering decides what an outage costs
  4. withdrawal is the last step
  5. set the deadline before you need it

basics

~20 s

Not as the first move. Cut what the exposed value may do at the destination, which takes minutes and stops no collector, then push the replacement, then withdraw the old value once nobody is presenting it - with a deadline after which you accept the outage anyway.

solid answer

~50 s

The question is a trade, and the answer is the trade you chose plus the price you accepted. Withdrawing immediately stops the attacker now and stops several hundred collectors with them; replacing first keeps the fleet running and leaves the exposed value live for the length of the push. Neither is a default - evidence of active misuse inverts it toward withdrawing first. Before choosing, establish two things: what the value can actually do, because narrowing its rights removes the danger in minutes without touching a host; and whether the collectors buffer and replay, because if they hold a day of data locally an outage costs backlog, while if they drop it, it costs a day of telemetry. The usual answer is staged: narrow rights immediately, push the replacement, confirm from the access record that nobody still presents the old value, withdraw it - and set a hard deadline, because waiting for the last holder has no natural end.

code

pseudocode · 23 lines
pseudocode
# t0 - exposure confirmed; ~400 collectors still hold oldValue
narrowRights(oldValue, allow = ['write'], deny = ['read', 'list', 'delete'])
restrictSources(oldValue, to = knownCollectorRanges)
# fleet is still uploading; the dangerous operations are gone

# t0 + 20m - start the long pole
newValue = issueCredential(destination, rights = ['write'])
pushConfig(allCollectors, newValue)        # hours to a day

# rolling check - the record distinguishes VALUES even when it cannot name holders
zeroRuns = 0
every 30 minutes until withdrawn:
    stillOnOld = countRequestsPresenting(oldValue, sinceLastCheck)

    if stillOnOld == 0:
        zeroRuns = zeroRuns + 1
        if zeroRuns == 2:                  # two consecutive empty checks
            withdraw(oldValue)             # old value stops working here
    else:
        zeroRuns = 0                       # someone is still on it; keep pushing

    if now > t0 + maxExposureWindow:
        withdraw(oldValue)                 # deadline wins; stragglers fail closed

go deeper

for a junior

Recall that replacing a credential and withdrawing the old one are two separate steps, and that only the second makes the exposed value stop working.

for a middle

Explain why withdrawal is fleet-wide under sharing and why narrowing what the value may do is the cut that can be made without touching any host.

for a senior

Run the incident: establish the rights, check for use, measure the push, choose an order, name the cost you accepted, and confirm disuse before withdrawing.

for a principal

Own the trade and the deadline. Decide how much exposure the estate tolerates against how much data an outage loses, and turn the finding into a scoping change rather than a better runbook.

## The question behind the question An interviewer asking this is not looking for *yes* or *no*. They are looking for whether you know that **containment and availability are the same lever when a credential is shared**, and whether you can price both sides rather than picking the one that sounds more responsible. The honest shape of the answer is: here is what I would establish first, here is the order I chose, here is the cost I accepted, and here is the deadline at which I stop waiting. ## What to establish before choosing - **What the value can do.** A write-only credential in public is a different incident from one that can read back and delete everything the fleet ever uploaded. This is checkable in minutes, and narrowing it is the cheapest cut available because it changes what the destination accepts and touches no host. - **Whether it is being used.** The access record cannot tell you *which* holder is presenting the value - that is the price of sharing - but it can tell you whether requests are arriving that your fleet did not make: from unexpected sources, at hours the fleet is idle, or performing operations the fleet never performs. - **What an outage actually costs.** If collectors buffer locally and replay, a day of refusal costs a backlog and some disk. If they drop what they cannot send, it costs a day of data that never existed. These are different decisions, and which one you are in is a question with an answer. - **How long the push really takes.** A day is an estimate; the number that matters is how long until the last holder has the replacement, and someone should be measuring it rather than asserting it. ## Two orders, two prices | | withdraw, then replace | replace, then withdraw | |---|---|---| | The exposed value | stops working immediately | stays live until the last holder moves | | The fleet | stops until reconfiguration completes | keeps running throughout | | You accept | an outage of known length | continued exposure of unknown use | | Chosen when | there is evidence of misuse, or the rights are dangerous | the value is narrow, and nothing suggests it is in use | A candidate who names one of these as correct without the condition has answered a different, easier question. ## The staged path that usually wins 1. **Narrow the rights on the exposed value** so it may only do what the fleet actually does. Minutes, no host touched, and it removes the outcomes you fear most. 2. **Restrict where it is accepted from**, if the destination can filter by source. Coarse, evadable, and still worth the five minutes. 3. **Issue and push the replacement.** This is the long pole and everything else is arranged around it. 4. **Watch for use of the old value.** Zero requests presenting it, sustained across two consecutive checks, is your evidence that holders have moved. 5. **Withdraw the old value.** This is the step that makes it stop working; steps 1 to 4 did not. 6. **Withdraw it at the deadline regardless.** Waiting for the last holder has no natural end, and a stranded collector is a smaller loss than an indefinitely live exposed credential. ## What makes this a judgement call rather than a procedure The deadline in step 6 is the whole decision compressed into one number, and nobody can compute it for you. It is set by how sensitive the destination is, what the narrowed value can still do, what the data is worth, and how much evidence of misuse you have - and you will be setting it with incomplete information, because the shared credential has already cost you the ability to say who is using it. ## What you owe afterwards The incident's real finding is not that a value leaked. It is that **containment cost an outage**, and that is a property of the fleet's scoping rather than of this leak. The follow-up work is on the axes: narrow what the credential may do permanently, split it per environment, and decide what granularity of holder identity the fleet will pay for - so that next time, withdrawal is one host's problem and this conversation does not happen again.

  • What changes your answer from staged replacement to withdrawing the shared value immediately?
    Evidence that it is being used by someone else - requests from sources the fleet does not use, at hours it is idle, or performing operations it never performs - or rights you cannot narrow far enough, such as a value that can delete what has already been uploaded. Either inverts the default: you take the fleet-wide outage now and explain it afterwards.
  • How do you know when it is safe to withdraw the old value?
    From the destination's record: count the requests still presenting the old value. Under sharing that record cannot tell you which holder they came from, but it distinguishes values perfectly well, so sustained zero across consecutive checks is real evidence that holders have moved. It is not proof - a collector that uploads once a day may simply not have run yet, which is why the deadline exists.
  • Why not simply issue the replacement and leave the old value in place indefinitely?
    Because replacing is not withdrawing. Until the old value stops being accepted, the exposed credential still works and the incident is still open; the new value has only given your fleet somewhere to move to. An exposed value left live is a standing grant to whoever found it, and it will not show up as a failure anywhere.

saying these in an interview costs you the question

  • Answers yes or no without pricing the outage or the exposure
  • Treats issuing the replacement as having contained the incident
  • Assumes the fleet buffers and replays without checking
  • Waits for the last holder to move with no deadline at all
  • Skips narrowing rights, which needs no host change and lands in minutes
  • Claims the access record will show which host is still on the old value