A scheduled rotation changed a datastore account's password but crashed before recording the new value — why is re-running it not automatically safe?
answer
- two sides, no transaction
- which identity makes the change
- the live value is the lost one
- failed attempts can lock the account
- reconcile before you re-run
basics
~20 sThe live value is now lost and the two sides disagree. A re-run may be unable to authenticate to make a second change, may consume the account's only spare credential slot, or may trip lockout and reuse rules. Reconcile first.
solid answer
~40 sThe crash left the downstream accepting a value nobody holds, and the store holding the previous one. Whether a re-run helps depends entirely on how the job authenticates. If it changes the account's password while logged in **as that account**, using the value the store holds, it can no longer log in at all — and every attempt is now a failed authentication, which can lock the account and bury the real signal. If it changes the password through a separate administrative identity, a re-run does work, and the right first move is still to establish what is actually true on each side. Then check the downstream's reuse history, how many credentials it will hold, and whether consumers are failing right now. A rotation that cannot verify both sides cannot tell half-done from done.
code
pseudocode · 19 linesrun = { id: newId(), account: account, state: "planned" }
record(run) // no credential value ever goes in here
candidate = generate(account.valueRules) // length + character set the downstream accepts
if not admin.canAuthenticate(): // fires when the job's only identity is the account itself
fail(run, "no change path independent of the rotated value")
store.writePending(account, candidate) // candidate now survives a crash
run.state = "candidate-stored"; record(run)
admin.setCredential(account, candidate)
run.state = "downstream-changed"; record(run)
store.promote(account, candidate)
if verify(account, store.read(account)):
run.state = "verified"; record(run)
else:
alert("half-finished rotation", run) // do not retry: the live value may be the candidatego deeper
Know that changing a credential touches two systems that cannot be changed together atomically, so a failure part-way leaves them disagreeing about which value is correct.
Explain the two orderings and which failure each produces. Say why a downstream that holds only one credential gives you no ordering to choose at that end.
Reconcile before acting: establish what each side believes, use a path that does not depend on the lost value, and explain how retries can lock the account and hide the real signal.
Make the intermediate state a first-class part of the design. A rotation platform other teams depend on must record intent before acting, verify as part of the run, and alert on incompleteness rather than retry.
## Two writes, and no transaction across them Every rotation of a credential in use is at least two writes in different systems: a change at the **downstream** that accepts the credential, and a record in the **store** that hands it out. There is no transaction spanning them. Any crash, timeout, deploy or network partition between the two leaves the sides disagreeing, and the disagreement has a direction that matters. ## What this crash left behind | Side | State after the crash | |---|---| | Downstream account | accepts a value that exists nowhere outside the crashed process | | Store | holds the previous value, which no longer authenticates | | Consumers | failing, or about to fail the next time they authenticate | | Run record | says the change started and nothing more | This is the worse of the two orderings. The other one — record first, then apply downstream — leaves the store holding a value that does not work *yet* while the old value still does, which is recoverable by either re-applying or discarding the candidate. Applying downstream first can destroy the only working value, which is what happened here. ## The order you thought you chose With a downstream that holds **one** credential for the account there is no ordering to choose at the downstream itself: applying the new value is simultaneously the replacement and the withdrawal of the old one. The only ordering you control is when the candidate value is recorded relative to being applied. Recording a candidate before applying it is what turns this crash from a loss into a look-up. Where the downstream will hold **two** credentials, neither ordering is fatal, which is the practical reason to ask for that capability before automating anything. ## Why a blind re-run is a gamble - **The job may no longer be able to authenticate.** If it logs in as the account it rotates, using the value from the store, that value is stale and every attempt fails. - **Repeated failures can lock the account**, converting a recoverable half-finished rotation into an outage plus a support call. - **A second slot may be consumed.** Where the downstream holds a limited number of credentials, a re-run can spend the spare you needed for recovery. - **Reuse and history rules may reject the next value**, producing a failure that looks like the first one and is not. - **You still do not know what is true.** The crash may have happened before, during or after the downstream change actually applied; acting before establishing that is a second uncontrolled change on top of the first. When the job authenticates through an administrative identity independent of the rotated credential, a re-run genuinely is the fix — which is exactly why that independence is the property to design for rather than to discover. ## Reconciling before acting 1. **Establish truth on each side.** Does the value in the store authenticate? Do the downstream's own change records show a change at the time of the crash? 2. **Use a path that does not depend on the lost value** — the administrative identity, or the human break-glass route the supplier offers. 3. **Record the candidate before applying it**, so the recovery does not repeat the original defect. 4. **Verify by authenticating** with exactly the value the store now holds. 5. **Only then look at consumers**, whose failures are a symptom of the disagreement and will clear when it does. ## Designing a run that is safe to re-enter - Give the job an identity **independent** of the credential it rotates. - Write the run's intended state **before** either side is touched, so an incomplete record is itself the alert — and keep the credential value out of that record. - Make verification part of the run rather than a separate job, so "changed" and "working" are never confused. - Bound retries and alert on the incomplete state instead of looping. - Ensure only one rotation for an account can be in flight; two jobs racing produces the same disagreement twice, with no way to tell which value won. The general lesson is not about crashes. It is that a rotation has an **intermediate state that is legal, reachable and invisible**, and a job that cannot describe that state cannot recover from it.
- Which order would you build the job in, given that both orders can fail?The one whose failure leaves a working value somewhere you control. Recording a candidate before applying it means a crash leaves the old value still live and the candidate on record — recoverable either way. Applying downstream first can destroy the only working value. Where the downstream will hold two credentials, neither failure is fatal, which is the real reason to ask for that slot.
- How would you detect the half-finished state without waiting for a consumer to fail?End every run with a verification that authenticates using exactly the value the store now holds, and alert on any run that changed the downstream without completing that step. Write the run's state before each side is touched, so an incomplete record is the alert rather than something you reconstruct afterwards.
- The account is now locked by repeated failed attempts. What did the job do wrong?It retried an authentication it could no longer satisfy. A rotation job should authenticate through an identity independent of the credential it rotates, stop after a bounded number of failures, and raise an incomplete-run alert rather than loop. Blind retries turn a recoverable disagreement into a locked account and a support queue.
Changing the only lock on a door and then losing the new key before it was copied. The door works perfectly — for nobody, and you cannot cut a copy from a key you do not have. You go back through the locksmith, who was never holding your key in the first place.
saying these in an interview costs you the question
- Just run it again, rotation jobs are idempotent
- The store still has the old value, so nothing was lost
- Retrying until it succeeds is the safe default
- The downstream will fall back to the previous password
- A crash mid-run means nothing was changed
- Ordering does not matter since both steps happen anyway