skip to content

A restored secret store hands out a credential the downstream system now rejects — what does that tell you about the restore point, and what is the fix?

level: seniorimportance: should knowfreq 37%

answer

  1. the store went back, the credential did not
  2. rotation is two writes, restore undoes one
  3. healthy store, rejecting downstream
  4. rotate forward, never reinstate
  5. an overlap window hides it for a while

basics

~20 s

The credential was rotated after the restore point, and the restore rewound only the store's copy while the downstream system kept the new value. Repair forward: set a fresh value downstream, write it into the store, never reinstate the old one.

solid answer

~40 s

Rotation is two writes — a new value into the store and the same change into the system that accepts it. A restore undoes one of them, so the store is left confidently serving a value the far end has stopped accepting, with nothing on the store's side reporting a fault. The repair direction matters: pushing the old value back into the downstream system reinstates a credential someone deliberately replaced, and if that rotation answered an exposure it reopens it. Instead set a new value at the downstream system, write it into the restored store, and let consumers re-read. Finding the affected set means probing, not asking the store — its own change history was rewound with everything else.

code

pseudocode · 12 lines
pseudocode
# the store is back at restorePoint; every other system is still at now
for each name in restoredStore.names():
    value = restoredStore.read(name)
    if downstream(name).accepts(value):
        ok(name)                                  # unchanged since restorePoint
    else:
        newValue = downstream(name).setNewCredential()   # repair forward
        restoredStore.write(name, newValue)
        markForReRead(consumersOf(name))

# never: downstream(name).setCredential(value)
# that reinstates a credential someone deliberately replaced

go deeper

for a junior

Recall that a restored store can serve an out-of-date credential, and that the system on the other side is the one deciding whether a credential still works.

for a middle

Explain rotation as two writes in two systems, and a restore as something that undoes only the store's half — which is why the store looks healthy while authentication fails.

for a senior

Show the operating judgment: repair forward, never reinstate, and describe how you would find the stale set by probing when the store's own change history was rewound with it.

for a principal

Decide what must survive a restore for this to be recoverable at all — where rotation records live — and put a named reconciliation step with an owner into the recovery procedure.

## The failure: a store that is confidently wrong After a restore, a consumer authenticates with the value the store just served and is refused. The store is healthy: it opened, it serves reads, its own checks pass. The refusal comes from the far end. This shape — healthy store, rejecting downstream — almost always means the **restore point sits behind a rotation**. ## Why a rotation survives a restore only halfway Rotating a credential is two writes in two systems: 1. The new value is written into the store, so consumers will fetch it. 2. The same change is made at the system that accepts the credential — a new password set on the account, a new key registered, the old one withdrawn. The restore rewinds the first. It cannot rewind the second, because the downstream system was never part of the backup. So the two halves disagree, and the store is the half that is wrong. It is worth separating two words that get used interchangeably here. **Rotation** put a new value in place. **Withdrawal** is what made the old value stop working at the far end. If a rotation replaced the value but never withdrew the old one, a restore behind it causes no visible failure at all — and that silence is not good news, it means an old credential is still live. ## Which direction to repair | Situation | The move | |---|---| | The rotation was routine and the current value is recoverable from where the rotating process recorded it | Write that current value into the restored store; nothing downstream changes | | The current value is not recoverable anywhere | Set a new value at the downstream system, write it into the store, let consumers re-read | | The rotation answered a suspected exposure | Always rotate forward; reinstating the old value would reopen the exposure you closed | The wrong move is the tempting one: make the downstream system accept what the store already holds. It is one command, it fixes the symptom instantly, and it reinstates a credential that was deliberately replaced — with no record of why it was replaced, because that record was in the store and was rewound too. ## Finding the affected set You cannot ask the restored store which of its values are stale; its change history went back with it. The sources that do know are all outside it: - The rotation schedule or runbook: which credentials were due to change in the window between the restore point and now. - Whatever the rotating process recorded elsewhere — a log line, a ticket, a record in a sink outside the store. - The downstream systems' own last-changed timestamps, where they keep one. - A direct probe: attempt an authentication with each value the store serves and treat a refusal as the signal. The probe is the only one that is complete, and it is also the one that needs care — a run of failed attempts against many accounts can look like an attack, and on some systems it can lock the account you are testing. Say that out loud in an interview; it is the detail that separates someone who has done it from someone reasoning about it. ## Why it can look fine for hours and then fail If the rotation ran with an **overlap window** — a period in which the old and new values are both accepted, so consumers can move without an outage — the restored store's old value keeps working until that window closes. The restore looks clean, monitoring is green, and the failure arrives when the old value is finally withdrawn, possibly days later and by then apparently unrelated to the recovery. This is the strongest argument for reconciling deliberately right after the restore rather than waiting to see what breaks. ## What the store cannot tell you, and what that implies A restored store's confidence is structural: it serves the newest value *it* holds, and it has no way to know that the world moved on. Nothing in the read path can flag staleness, because staleness is defined at the far end. Two design consequences follow, and both are fair to raise: - Rotation records are worth keeping somewhere the store's own restore cannot rewind, because they are exactly what you need at the moment the store has been rewound. - A post-restore reconciliation sweep belongs in the recovery procedure as a named step with an owner, not as an instinct someone has at the time. ## The sentence that answers the question "The restore point is older than the last rotation of that credential, so the store is serving the pre-rotation value while the downstream system only accepts the post-rotation one. I repair forward — new value downstream, written into the store — and I never push the old value back, because that reinstates a credential we deliberately replaced."

  • How do you know which credentials were rotated after the restore point?
    Not from the restored store — its history was rewound with it. From the rotation schedule, from whatever the rotating process recorded outside the store, from the downstream systems' own last-changed timestamps, and finally from probing each value against the system that accepts it. Only the probe is complete, and it needs care around lockouts.
  • Why might a restore behind a rotation look fine for hours and fail later?
    Because the rotation may have run with an overlap window in which both values are accepted. The restored store's old value keeps working until that window closes, so the failure lands when the old value is withdrawn — long after the recovery, and looking unrelated to it.

saying these in an interview costs you the question

  • Pushes the old value back into the downstream system to make the two match
  • Assumes the store is authoritative because the restore reported no errors
  • Thinks the mismatch resolves once consumers re-read from the store
  • Calls the restored value current because it is the newest one the store holds
  • Says the rotation failed, rather than that the restore rewound one side of it